Compare commits

...
Author SHA1 Message Date
omarandClaude Opus 5 bbb493ea91 fix(ci): the package-count assertion lives in two scripts and only one was updated
test / go + panel tests (push) Successful in 1m40s
release / test gate (push) Successful in 1m40s
release / apk aarch64_cortex-a53 (push) Successful in 8m55s
release / apk x86_64 (push) Successful in 2m51s
release / release apk (push) Successful in 8s
D29 removed byedpi, so the feed carries three packages. sdk-build-apk.sh was
changed to >=3; build-feed-apk.sh still demanded >=4 and killed both arch lanes
of v0.2.22 with `expected >=4 .apk … found 3`. Nothing was published from that
run. The comment now says the count is duplicated, because reading one script
was what made this look done.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 17:31:35 +03:00
omarandClaude Opus 5 4869d62e02 feat(egress)!: remove byedpi — what it replaced was not weak, it was broken (D29)
test / go + panel tests (push) Successful in 1m39s
release / test gate (push) Successful in 1m39s
release / apk aarch64_cortex-a53 (push) Failing after 2m54s
release / apk x86_64 (push) Failing after 2m54s
release / release apk (push) Failing after 1m35s
The `byedpi` egress kind, the `openwrt/byedpi` package (`ciadpi`), the readiness
endpoint and the panel plate are gone. D13 is not deleted from DECISIONS.md; it
is REVERSED there, with the reason, because the reason is the whole point.

D13 adopted an external desync process on an observation: the engine's own
`tls_fragment`/`tls_record_fragment` were tried against a live ISP and did not
get through, so the method was judged too weak for anything past "just fragment
the ClientHello". The method was never tried. `common/tlsfragment` dropped a
number of labels equal to the number of DOTS in the name, and a name always has
one more label than it has dots — so the cut always landed inside the FIRST
label. `www.youtube.com` was split inside `www` and `youtube` went to the wire
in one piece, which is the word the DPI matches on. Of six blocked names exactly
one got through: `youtube.com`, the one whose first label IS the blocked word.
That defect is fixed (815011dfb, efb2177f4). With it fixed the built-in presets
do the job the external process was brought in to do, and the process is 100 KB
of binary, a second procd service, a second UCI file, a port that agreed with
our egress by hand-written comment only, a readiness prober, a five-state
service model and a panel plate — all to work around fifteen lines of ours.

So this is not "ByeDPI turned out to be bad". It is a good tool that turned out
not to be needed, and the reason we thought it was needed was ours.

A CONFIG THAT STILL SAYS `type 'byedpi'` IS THE PART THAT NEEDED WORK. Nothing
is migrated and nothing is rewritten: the kind stays unbuildable, therefore
fail-closed — no outbound, no mark, no `ip rule`, no routing table, so every
node, group and rule bound to it is blocked rather than released onto the plain
WAN. A migration to `direct` was considered and rejected: it is the only rewrite
that leaves the egress routing at all, and it would silently turn a blocked
egress into a live plain-WAN path with the router's real address — by an
upgrade, on a config nobody touched. `CurrentSchemaVersion` is therefore not
bumped either: no stored field changes meaning, and a bump would only make this
build's configs unreadable to an older daemon for no gain.

What changes is what the operator is TOLD. `model.RetiredEgressTypes` is a
closed, positive table read by BOTH `ValidateEgresses` and the generator (one
copy of the sentence, because two copies drift). It names the removal, denies
that it is a typo, says nothing is built and that the traffic is blocked rather
than leaked, names the replacement (`direct`/`interface` with `dpi 'record'`),
refuses to promise which preset defeats a given ISP, and says `apk del byedpi`.
The generic "unknown type" is still there and still says something different, on
purpose: "we took this kind away" and "you mistyped something" send an operator
to different places, and a value that was correct on the day it was written must
not be reported as a spelling mistake. The type list stays closed and positive —
`interface`, `direct`, the alias `tunnel` — and `EgressTypeKnown` does NOT admit
the retired kind: being told it was removed and having it work anyway is worse
than either alone.

`Egress.Port` goes with the kind: no surviving egress dials anything, so the
option is no longer parsed and drains out of /etc/config/shater on the next
render, the same way the deleted per-group probe_url/probe_interval did.

Tests, verified by mutation, each failing by name:
  - drop the retired branch in `ValidateEgresses` -> the retired kind is
    reported as "is not one of interface/direct" and
    TestRetiredEgressTypeIsReportedByTheValidator fails on both spellings;
  - drop it in the generator -> "unknown type \"byedpi\"" and
    TestRetiredEgressTypeIsReportedByTheGenerator fails;
  - the FAIL-OPEN mutation, which is the one that matters: let `byedpi` fall
    into the `direct` arm and be a known type -> four tests fail, including the
    two that check no outbound is emitted. A removal that quietly starts routing
    the traffic it used to block, under a reassuring message, is the failure with
    the worst consequence;
  - the panel half: empty RETIRED_EGRESS_TYPES -> two egressEdit tests fail.
Controls beside the claims: `interface`, `direct`, the `tunnel` alias and the
empty synonym must still resolve, warn about nothing and emit an outbound
(TestSupportedEgressTypesAreUntouched), and never-supported values — `proxy`,
`block`, `wireguard`, `byedpi2`, `bye dpi`, `sorcery` — must NOT draw the
removal sentence, which names a replacement for something that never existed.

CI and docs: the feed loses its fourth package everywhere the four were named —
`apk upgrade shaterd shater-core luci-app-shater`, in CLAUDE.md, both READMEs,
INSTALL.md, the release body and `shaterd`'s own diag bundle. The version
exception (byedpi carried upstream's version, ours come from the git tag) is
gone with it, so ci/version.sh and ci/sdk-build-apk.sh no longer have an
exception to remember and the "expected >=4 of OUR .apk" collect check is now 3.
INSTALL.md §5.3 gains the half a feed cannot do: dropping the package from the
feed does not take it off a router it is already on, so `apk del byedpi` is
written down, with what it removes and why it is safe.

Panel: 368 tests -> 339. Deleted with the mechanism they covered:
byedpiReady.test.ts, byedpiAge.test.ts, byedpiRefusal.test.ts (34 tests);
egressEdit.test.ts gains 5 for the retired-type sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 17:13:50 +03:00
omarandClaude Opus 5 efb2177f43 fix(tlsfragment): one cut, in the label a blocklist keys on — and a budget for it
Follow-up to 815011dfb, which fixed WHICH label is cut but left "a cut in every
candidate label" as an unconditional rule. Measured on this tree, loopback peer,
product default fallbackDelay, one ClientHello per row:

    cuts   tls_fragment (*net.TCPConn)   tls_fragment (proxy conn)   tls_record_fragment
       1                        502 ms                      500 ms                 <1 ms
       2                       1.004 s                     1.001 s                 <1 ms
       4                       2.008 s                     2.002 s                 <1 ms
       8                       4.015 s                     4.003 s                  539 us
      21                      10.540 s                    10.509 s                  525 us

So a cut in the PACKET modes costs half a second of connection setup, and it
costs that on BOTH branches — not only on the sleep path. writeAndWaitAck sleeps
the whole fallbackDelay whenever the ACK returns inside 20 ms (its "under
transparent proxy" case), and N.UnwrapReader reaches the *net.TCPConn only when
nothing in the chain transforms the stream, which a proxy protocol conn always
does. A proxied egress — every subscription node — therefore takes the flat
500 ms branch regardless of RTT. The number of labels is chosen by whoever picked
the hostname, and a 253-byte SNI is 85 of them: ~42 s of one connection's setup,
bought from the LAN.

In tls_record_fragment nothing waits: the ClientHello leaves in ONE write, split
into more records. 21 cuts cost 525 us and 105 bytes of record headers, and
1.1.1.1 completed the handshake with the ClientHello in 22 records in the same
77 ms it took with 2. That is the mode the field measurement was taken in, and
the mode where cutting every label was always affordable.

Hence two budgets rather than one rule: 1 cut for the packet modes, 4 for
record-only — the latter not a cost limit but a shape limit, since real names
carry one to three labels outside the public suffix and a hostile one must not
turn a ClientHello into 85 records no ordinary client emits.

One cut is enough because of WHERE it goes. Candidates are now ordered, most
worth cutting first, and first is the REGISTRABLE label — the one immediately
left of the public suffix. That is what a name-based blocklist keys on
("youtube" of youtube.com, www.youtube.com and studio.youtube.com alike,
"ytimg" of i9.ytimg.com, "example" of a.b.example.co.uk), and severing it also
breaks any match on the whole FQDN, so one cut covers both matchers. It is
chosen by STRUCTURE, from the public suffix list — not by length, which is the
same trap from the other side: in cdn-static-assets.youtube.com the longest
label is not the blocked one. The rest follow longest-first, on the argument
that among labels with no structural ranking a long one is likelier to be a
distinctive token than "www", "m" or "tv"; they are reached only when the budget
allows more, or when the registrable label is too short to cut.

The offset now comes from the label's MIDDLE THIRD. Every interior offset severs
the label, but one byte in leaves "outube" of "youtube" and a matcher keyed on a
substring still reads it. The draw stays random inside that third: a fixed point
would be a constant a middlebox vendor can special-case in one line, and this
whole family of tricks lives on making reassembly the only counter.

Also in this commit, and the reason it is not merely a tuning change: the panic
that shipped in v0.2.21 now has an instrument of its own.
TestWriteDoesNotPanicOnAServerNameChosenFromTheLAN drives real ClientHellos
carrying ".youtube.com", "youtube.com." (a legitimate FQDN with the root dot,
which curl and every browser will send), "..", an IP literal and non-ASCII bytes
through all three modes, and FuzzCutOffsets does the open half — 25.7 million
executions found nothing, and the fuzzer is shown able to find a planted defect
its seed corpus cannot reach, in one second. A hand-built ClientHello reaches
the shapes crypto/tls refuses to emit: a zero-length name, a 253-byte name, and
a server_name_list with a SECOND entry, which is why planning runs on
MyServerName.Length rather than on everything left in the extension.

Nine mutations, each failing by name with the numbers: the old dot arithmetic,
the old rand.Intn offset, the exact original expression (panic: invalid argument
to Intn, conn.go:208 <- Write conn.go:67), the empty-plan guard, the budget, the
priority order, the sort back into wire order, the first-entry truncation, the
middle third, and a one-byte corruption of a segment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 17:02:06 +03:00
omarandClaude Opus 5 815011dfb0 fix(tlsfragment): the SNI was cut in exactly one label — always the first one
`splits[:len(splits)-strings.Count(serverName.ServerName, ".")]` is identically
`splits[:1]`: labels are always one more than dots, so the subtraction cancels
for EVERY name in existence. One label was ever cut, and it was the leftmost
one. On the provider measured from this router — which blocks by the name in
the handshake, proved by the same address answering for SNI www.google.com and
going silent for www.youtube.com — that is the whole observed table:

    youtube.com     cut inside "youtube"  -> 301
    m.youtube.com   cut inside "m"        -> blocked
    tv.youtube.com  cut inside "tv"       -> blocked
    www.youtube.com cut inside "www"      -> blocked
    music/studio.*  cut inside the label in front -> blocked

The one name that worked is the one whose first label IS the blocked word. The
count subtracted must be the labels of the PUBLIC SUFFIX, not the dots of the
whole name: "com" is one, "co.uk" and "com.br" and "pp.ru" are two.

Second half of the same defect, and the reason the table above shows a cut
"inside m" at all: the offset was `rand.Intn(len(label))`, whose 0 is the
label's own boundary — the label goes out whole in the next segment, which is
not a cut, it is a segment boundary that happens to touch a label. For a
one-byte label 0 is the ONLY value it can take. Offsets are now drawn from
[1, len-1], so a cut always leaves a non-empty piece of the label on both
sides, and a label too short to have an interior offset carries no cut instead
of a fake one. That also closes the 1-in-7 hole in the case that WAS working:
youtube.com drew offset 0 once every seven connections and handed the name over
intact.

Two panics went with it, both reachable from the LAN, because route/conn.go
wraps the outbound with this and the ClientHello it fragments is the client's:
an empty label (SNI ".youtube.com" or the perfectly ordinary FQDN
"youtube.com.", where the suffix list declines to answer and the trailing empty
label survives) reached rand.Intn(0) — "panic: invalid argument to Intn", the
daemon and with it the router's proxying. And a plan with no cuts at all would
have indexed b[:splitIndexes[0]] on an empty slice; Write now writes the
ClientHello unchanged in that case, which is the only honest thing to do for a
name of one byte.

The classification is closed and errs toward MORE cutting: narrowing the label
set needs proof (a public suffix that really is a tail of the name), widening
needs none, so a trailing dot, an unmanaged TLD, a name that IS a public suffix
("com", "co.uk", "localhost") and an IP literal all keep every label rather
than fall silently into "cut nothing". When no label is long enough to cut, the
name itself is cut once — a matcher looking for the whole FQDN still fails
across that split.

Dropped with it: `splits[0] == "..."`, unreachable since strings.Split on "."
cannot produce a token containing a dot. And the plan now runs over the FIRST
entry of the server_name_list (MyServerName.Length) instead of everything left
in the extension, so a second entry cannot be fed to the public suffix list as
if it were part of the name.

Tests (cutplan_test.go, package-internal so the plan itself is visible) are
verified by mutation five ways: the old dot arithmetic, the old rand.Intn
offset, the removed empty-label guard, the removed empty-plan guard, and a
one-byte corruption of a segment. Each fails by name and with the numbers. The
controls: youtube.com — the case that already worked — must still be severed;
the reassembled segments must be byte-identical to the ClientHello in all three
modes (tls_fragment, tls_record_fragment, both), with the record framing
re-parsed rather than assumed; and Write must report len(b).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 16:33:04 +03:00
omarandClaude Opus 5 0a34e64c2c fix(panel): byedpi off the status poll, and disabled stops speaking for two situations
test / go + panel tests (push) Successful in 1m39s
release / test gate (push) Successful in 1m36s
release / apk aarch64_cortex-a53 (push) Successful in 6m0s
release / apk x86_64 (push) Successful in 2m52s
release / release apk (push) Successful in 8s
GET /api/status no longer carries the readiness report — the daemon dropped it
with the cache behind it, after one probe was measured at 6.4 s on 16 enabled
instances behind a black hole while the panel polled that endpoint every 5 s
from every open tab and read the field NOWHERE. The Status type, the mock
fixture and every comment describing a cache, a background refresh or a 20 s
staleness rule now say what the daemon does: one endpoint, and it connects when
a human asks.

`disabled` covers two situations with opposite next actions: no instance is
enabled — how the package ships — and an instance that IS written and looks
enabled while /etc/init.d/byedpi refuses it (`port 'auto'`, `port '99999'`,
`enabled ' 1'`, `enabled 'TRUE'` — all four measured on the 25.12.1 testbed
against validate_data). The editor's fixed sentence said "that is how the
package ships" about a section the operator had typed themselves. The daemon
keeps its `problems` list off the wire, so `detail` is the ONLY carrier: the
refusal now shows that sentence verbatim plus a tail that says only what is
true of both — the consequence, never the fix.

byedpiRefusal moves to byedpiReady.ts beside the gate it explains, and its
table now EXCLUDES `disabled` from the type, so re-adding a fixed sentence for
it does not compile. `?mock&byedpi=rejected` reaches the second case in a
browser; `?mock&byedpi=noanswer` reaches "nothing has been measured", which is
now only a failed fetch — the fabricated cold-cache body is gone.

Also: two comments about `config_applied` that the daemon's pointer+omitempty
change made false — the removed "positively phrased so a naive client falls the
alarming way" rationale, and "absent means a daemon too old", which now also
means the offline `shaterd status` stub.

Tests (byedpiRefusal.test.ts, +10) verified by mutation both ways: a fixed
"that is how it ships" and a fixed "your typo" each fail, and the control
asserts the factory state still reads as the factory state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:54:21 +03:00
omarandClaude Opus 5 1708159ecf fix(apply,shaterd): the offline stub alarmed about an apply nobody attempted
`config_applied: false` means "/etc/config/shater was read and REFUSED — what is
running is the PREVIOUS configuration, your edit is not in effect", and the panel
draws a critical band saying exactly that. The field was a plain bool, so that
alarm was the ZERO VALUE OF THE TYPE — and `shaterd status`'s offline stub, built
by a process that never applied anything, over a data plane that may have been
installed and enforcing for weeks, published it by simply never mentioning the
field. It is the config_readable defect returning in a new field, with the one
difference that decides the fix: config_readable can be MEASURED by the stub and
now is, while this one cannot be measured at all without a daemon.

So the field says nothing when nobody measured it. ConfigApplied becomes a *bool
with omitempty; the live Applier.Status() assigns a verdict on BOTH arms, so an
absent key can only come from something that is not a live status. That is the
same closed-set-plus-unknown shape `plane`, `traffic` and `daemon_answered`
already have, and the one panel/src/appliedConfig.ts already implements
(=== true / === false / else unknown). The Go doc claiming absence should read as
false is gone: it contradicted the only consumer, and the consumer was right.

The four fields around it (apply_error, apply_error_stage, apply_attempts,
apply_failed_since_unix) stay plain: they are qualified by config_applied the way
enabled/kill_switch/panel_port are qualified by config_readable, and their zero
values point at "nothing was refused" — the quiet side, not the alarm.

Also: the stub shipped `warnings: null` on its happy path while apply.Status
documents Warnings as always non-nil so a consumer can map over it
unconditionally.

Three states, distinguishable ON THE WIRE through one `shaterd status`, with the
control that would catch the opposite break (a build that omitted the key for a
real refusal, deleting the alarm from the product).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:43:49 +03:00
omarandClaude Opus 5 9d7f0dc92f fix(panel): the byedpi report had a five-second timer and no reader, and its parser could forge "listening"
Two defects found by looking at both sides of the byedpi readiness check at once.

1. GET /api/status carried the whole readiness report from a cache that a poll
   refreshed in the background once the copy passed byedpiRefreshAfter = 3 s.
   The panel shell polls that endpoint every 5 s, so EVERY poll started a
   refresh: a PATH lookup, a read of /etc/config/byedpi, and one connect per
   enabled instance, forever, per open tab, hidden ones included. The design
   note rejected a background ticker because "a closed panel costs nothing" —
   true, and silent about the open one it had become.

   Measured, one enabled instance, twelve polls five seconds apart:
     before  12 connects, 13 ciadpi PATH lookups per minute per tab
     after    0 connects, 12 PATH lookups (one per poll, for byedpi_installed)

   And nothing read it: `grep -rn '\.byedpi\b' panel/src` finds no consumer —
   the readiness plate, the per-egress cross-check and the egress-type gate all
   come from GET /api/byedpi. So the field is gone from the status response, and
   with its only cached reader gone the cache went too, together with the
   background goroutine, the staleness rules, the negative-age contract and
   Server.Close's duty to wait for a probe. GET /api/byedpi still connects, on
   the goroutine of the request that asked.

2. readByeDPIInstances claimed to mirror /etc/init.d/byedpi "exactly" and did
   not. The init script validates each section with
   'enabled:bool:0' 'port:port:1080' and refuses to start one whose validation
   failed. Go read the port with strconv.Atoi and, on failure, KEPT the 1080
   default — so `option port 'auto'` on an enabled instance became "an enabled
   instance on 1080", and anything else accepting there produced state
   "listening": the one state that unlocks the byedpi egress type, handed out
   for a proxy that does not exist. `port '99999'` produced the second half:
   "unknown" with a sentence asserting a connection attempt that never happened.

   The same shape lived in `enabled`: strings.ToLower+TrimSpace read ' 1' and
   'TRUE' as on, while the router starts neither (measured — the first is
   refused by validation, the second normalises to an empty value so
   `[ "$enabled" -eq 1 ]` never fires).

   The parse is now a closed positive list, and its expectations were MEASURED
   on the 25.12.1 testbed against /sbin/validate_data with the init script's own
   spec rather than inferred from libvalidate's source:

     enabled: absent/"" -> off; exactly 1|on|true|yes|enabled -> starts;
              exactly 0|off|false|no|disabled -> off; anything else -> does not
              start, and is REPORTED by section, option and value.
     port:    absent/"" -> 1080; plain decimal digits 1..65535 -> that port;
              anything else -> NO port is assumed, the section is not counted as
              a listener and nothing is dialled for it.

   Deliberately narrower than libvalidate's `port` (which also takes a sign,
   leading whitespace and, through an overflow, twenty digits): narrow declines
   to call a working instance a listener and prints why, wide hands out a green
   apply onto a port nothing is on.

Two further sentences that asserted actions that never happened, found while
fixing the above and not reported by the review: instances past
byedpiMaxInstances were never dialled yet fell into the "the connection attempt
neither succeeded nor was refused" clause, and that clause listed their ports
alongside genuinely inconclusive ones. "Not dialled" is now its own tally with
its own sentence, and each sentence names only the ports its own claim covers.

Every test here was checked by mutation, and each carries its control:
byedpi_initparity_test.go proves the instrument BOTH accepts a valid section
(state listening, against a real socket, in a world where every connect is
accepted) AND refuses every value the init script would not start, dialling
nothing for them; byedpi_pollcost_test.go measures the poll cost with a meter
shown counting a real probe in the same test, and keeps the probe-cost control
(16 black-holed ports = 6.4 s) that explains why it is off the poll path.

Gate: bash scripts/run-tests.sh green, including -race; ok shater/panel by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:32:26 +03:00
omarandClaude Opus 5 c407771cf2 fix(panel): four cards, one log row and a status the daemon knew and nobody saw
The board is not one row per name. carryForward supersedes by (name, KIND) and
appends carried rows LAST, so a chain `x` and a node `x` both live on it — and
`new Map(results.map(r => [r.group, r]))` kept the last. The chain card showed
the node's milliseconds, exit address and verdict as its own end-to-end
measurement, unmarked. Attribution is now by kind (targetResult.ts), with
kind:'' and a missing kind as ordered last resorts.

A connection routed to the engine's `block` outbound was drawn as plain mono
text, indistinguishable from `nl-reality-1` — on the page where a DNS row about
the same host gets a crit rail and a BLOCK mark. It is the kill-switch's own
Final and a legitimate rule target, so the connection log now carries the same
outcome axis the DNS log has: killed / carried / no exit recorded, a crit rail
and a mark that survives the width where the exit column is dropped.

Insights.tsx held a raw NUL at byte 36359 — a template separator written as the
byte instead of the escape. `file` called the source binary and ripgrep, git grep
and every tree-wide search skipped it in silence. It is the escape now, and the
whole of panel/src is free of control bytes.

Three contract texts had drifted from the daemon: the searched-field list did not
mention `error` (fixed on the Go side, and there were two copies), the connection
hint named neither `proto` nor the chain hops, and rowMatches folded case with
toLowerCase() — Unicode-aware, where the daemon folds ASCII only, so a needle
could find rows in the panel that the router would never return.

And the four status fields the daemon started publishing: config_applied,
apply_error, apply_error_stage, apply_attempts, apply_failed_since_unix. A
refused configuration retried on a widening interval while `engine_running` was
true, the hash was the OLD config's and every warning described the OLD config.
engine_running is TRUE there and is not contradicted — the band says WHICH
configuration is running, and the hash row, the traffic default and the findings
list each say they are about that older one. Absent is not false: a daemon
without the field is `unknown` and raises nothing, because there is no evidence
its hash is stale.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:28:40 +03:00
omarandClaude Opus 5 26e1d38924 test(bridge): make the fragment sweep test assert the property it names
TestBridgeFragmentSweepIsPerCall claimed its probe used "an EXISTING key, not a
new one: the sweep must still run". It did not: the stale datagram carried IPv4
id 61 and the probe id 62, and fragKey includes the identification, so the probe
opened a NEW key — the one arrangement in which the sweep runs even when it runs
only on new keys. Moving r.sweep(now) inside the `entry == nil` branch left the
test green.

The probe is now the SECOND fragment of a datagram whose first fragment is
already cached, with the two entries opened half a fragTimeout apart so the
stale one is past its deadline and the live one is not (deadlines are set at
creation and never refreshed). Two assertions before the probe pin the setup:
the stale entry must still be there, and the live key must already exist — if a
later edit breaks either, the test says so instead of quietly proving nothing.
The released bytes are checked too, which is the half of the timeout this test
is about (the correctness half is already caught by TestBridgeFragmentTimeout).

Same sweep of TestBridgeFragmentMalformed, which had the same shape of hole: a
FIRST fragment carries MF=1 and can never complete a datagram, so `got != nil`
is unreachable whether the packet was refused or accepted, and "truncated
header" asserted only that. Every subtest now asserts on the cache, and a case
for the classic overread — a header claiming TotalLength 276 in a 28-byte
buffer — is added; its control is the aligned subtest already at the bottom.

Mutations (linux, -race): sweep moved into the new-key branch fails
SweepIsPerCall by name; clamping TotalLength to the buffer instead of refusing
fails the new malformed subtest — and, as predicted, leaves its `got != nil`
assertion silent. Control: moving the sweep after the entry lookup while keeping
it unconditional keeps every test green, so the test discriminates "per call",
not "the line moved". No production code changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:25:35 +03:00
omarandClaude Opus 5 42d84ac74c fix(model): a failed backup may stop a config write only when the filesystem is the reason
backupBeforeChange was added with "any failure aborts the write", justified by
"the uci commit that follows writes the same filesystem, so whatever stops one
stops the other". That holds for a full or read-only /overlay and for nothing
else — and the existence probe is a stat, which also returns ENOTDIR (something
dropped a file where /etc/shater should be), EACCES, ELOOP. In that state
PUT /api/config answered 500, `sub update` exited non-zero and the profile
watcher stopped saving, PERMANENTLY: none of those causes clears itself. A
convenience added this wave must not be able to take the product away.

Two changes, both about not inferring what can be measured:

- The probe is not evidence. stat(dest) answers "is this transition already
  captured?"; when it cannot answer, the copy is now ATTEMPTED and the attempt
  is the measurement. Only "the filesystem will not take bytes" short-circuits
  it.

- The failure is classified. filesystemRefusesWrites is a positive, CLOSED list
  — ENOSPC, EROFS, EDQUOT, EIO — each a condition under which the uci commit
  would fail too, so aborting only changes which error the operator reads and
  ours names the cause. Everything else is about the backup's PATH and falls to
  the recoverable side: the config is saved, and the missing undo is NAMED
  through reportBackupProblem (same shape as subCacheLogf; model cannot import
  logsink, which imports model) rather than skipped in silence.

TestWriteAbortsWhenTheBackupCannotBeWritten used a FILE where the backup
directory should be — that is ENOTDIR, the exact case that must no longer veto —
so it now injects ENOSPC at the copy, and the ENOTDIR case moved to
TestBackupPathFailureDoesNotVetoTheWrite. statBackup/writeBackupFile are seams
because the two deciding failures are the two a temp directory cannot produce.

Mutation-checked (linux, -race), each with the other half green: restoring "any
failure aborts" fails only the two carry-on tests; "nothing aborts" fails only
the two abort tests; restoring the old stat handling fails only the test that
pins "attempt the copy"; dropping ENOSPC from the list or adding ENOTDIR to it
fails the classifier test and the end-to-end tests that depend on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:25:20 +03:00
omarandClaude Opus 5 a05ad21b39 fix(apply,shaterd): four states the daemon was in and could not say
1. A SWITCHED-OFF SUBSCRIPTION CAN STILL GO OUT ON THE PLAIN WAN (blocker).
   Three places had to agree about `enabled=0` and did not: UpdateSubscription
   resolves by name and never reads it; cmdSubUpdate reads it only when no name
   was given; warnings.go skipped disabled subscriptions entirely on the stated
   premise that one "is never fetched". The premise was the false one, and the
   per-row Fetch-now button added this wave posts exactly the named request.

   Kept the behaviour, dropped the premise. Enabled means "include in the
   automatic refresh" everywhere else in the system — MergeSubCaches loads a
   disabled subscription's cached nodes unconditionally and they route traffic —
   and a refusal here is worked around by enable/fetch/disable, which enrols the
   sub in the 6-hourly sweep and is strictly worse. The automatic paths still
   honour it (the nameless sweep, and shater-cron's own `en = 1` check). The
   named path now says so on stderr and in the daemon log, and the leak finding
   fires for disabled subscriptions with the WHEN clause corrected — "every
   scheduled refresh" is false of a subscription no schedule touches.

2. refreshBootArmor DISARMED THE NEXT BOOT FROM A CONFIG THE DAEMON REFUSES.
   Every call site is gated on readErr == nil and nothing else; ParseUCIExport
   drops unknown options silently, so a config written by a newer build reads
   clean, and with the divert set emptied by the parse RenderHoldNft returns ""
   and the armor was REMOVED — with no log line at all, unlike the disarm one
   branch above it. Measured: with the new gate removed, the armor really is
   deleted. Now gated on the schema, and both removal paths are announced.

3. THE FIRST-BOOT DEADLOCK IS NAMED. Every subscription pulled through the
   tunnel, the tunnel built from nodes only a fetch supplies, the caches gone:
   the fetch waits for the tunnel and the tunnel waits for the fetch, forever,
   with the LAN dark. The CLI refusal goes to /dev/null (shater-cron) and the
   daemon line to a syslog `log_syslog='0'` switches off. It is now a critical
   finding in /api/status, which survives both, with the state named and two
   escapes — the free one first, the costly one priced.

4. A REJECTED CONFIGURATION WAS INVISIBLE, AND THE ENGINE CHURNED. Measured on
   the stand: with a config the engine cannot accept on disk, cron retries every
   60s and every attempt is a full engine swap, while status showed
   engine_running=true, the OLD hash, the OLD warnings, and `grep -ci` for the
   broken element returned 0. Invisible by construction: everything published
   about a config is published by a SUCCESSFUL apply, and engineDownCause is
   gated on the engine being down — here it is up.

   Status gains config_applied / apply_error / apply_error_stage /
   apply_attempts / apply_failed_since_unix, and a critical finding that says
   the running configuration is a DIFFERENT one and names the reason. Reconcile
   paces an identical retry (three free attempts, then doubling to a 15m cap);
   any change to the configuration cancels the wait, and POST /api/apply is
   deliberately not paced. The post-swap abort is deliberately NOT recorded —
   it is already loud and its retry costs no swap.

Also, from review-by-seams:

 - The netplane channel was graded critical wholesale over three distinguishable
   states. `udp '0'` + closed is the kill switch doing what it was told and may
   be exactly what was asked for; the leak and the total cut-off are not. The
   first is now `warning` (not `info`: attentionFindings drops info, and the
   blast radius is wider than the switch's name). Default stays critical, the
   exception is a closed list, and netplaneprotoseverity_test.go pins it against
   the REAL renderer so a rewording fails by name instead of drifting.

 - devicefilter_severity_test.go carried a FOURTH unlinked copy of
   DEVICE-FILTER-NOT-APPLIED and compared it with itself — the same shape as the
   noGatewayFinding fixture this wave removed. apply's copies are one constant
   now, and the real coupling is a test that runs generate and grades what comes
   back. Mutation: renaming the tag in generate fails it by name; the two old
   fixture tests survive that untouched, which is the whole point.

 - The history-write failure was logged ABOVE the deduplication gate its own
   call site documents eight lines below. At one cron reconcile a minute a
   standing cause (full /overlay, an entry over the 128 KiB ceiling) wrote 1440
   identical lines a day, and under log_persist=1 that many appends to flash —
   the exact wear the history ring's own dedup exists to prevent. Now gated on
   the message changing, cleared by a success. The Warning is still returned
   every time; only the log had a repetition problem.

Every fix mutation-checked with the failure text recorded, and every one has a
control showing the instrument can still give the opposite answer: an enabled
subscription still fetches and keeps the scheduled wording; a legitimate disarm
still happens and is still logged; a healthy box raises no rejected state; an
ordinary netplane finding is still critical; a DIFFERENT history failure still
prints. One mutation (the history-dedup latch) SURVIVED its first test — the
counter matched the success path's Info line too — and the test was fixed.

shater/apply is green. shater/cmd/shaterd was green when run 20 minutes ago and
now fails to BUILD on shater/panel/byedpi.go, a neighbour's in-flight refactor;
the full gate run for the same reason cannot be completed on this tree right now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 14:08:57 +03:00
omarandClaude Opus 5 43bb8913ea fix(diag): the bundle printed a DNS account id, and scrubbed the wrong file
Two holes, both in the direction the verb cannot afford: `shaterd diag` produces
the one text block a person SENDS somewhere.

1. resolver.address was on the allow-list, printed verbatim. For a DoH resolver
   that field is a URL, and generate/dns.go's parseDoHAddress keeps and USES
   u.Path — which is exactly where NextDNS, AdGuard and Control D carry the
   account identifier. Whoever holds it reads and rewrites this household's DNS,
   so it is a credential. The panel had always read it that way (DNS.tsx's
   resolverAddr shows u.host and flags the rest); the disagreement was resolved
   in favour of the side whose output goes to a stranger. Now: scheme and host
   survive, userinfo/path/query/fragment do not, and a BARE address
   ("1.1.1.1", "dns.adguard.com:853") is still printed in full because it is
   host and port and it is what the fault is read from.

   The fix could not be "delete the key from the list": TestDiagMasking-
   IsClosedOverTheWholeModel asserted the allow-listed fields come out
   UNMASKED, so it actively pinned the leak. The transform lives in a second
   closed table (diagMaskedForm), and the sweep now compares the masked render
   against the raw one line by line, expecting either the plain mask or exactly
   what that table declares.

2. The second layer collected its literals from `uci export shater` alone, and
   that file does not hold this router's credentials. model/render.go never
   writes a FromSub node; the several hundred subscription nodes live in
   /etc/shater/subs/*.json, which keep.d/shater-core describes in its own words
   as carrying "every node's credentials". The reachable path is not
   hypothetical: parse/sharelink.go quotes a rejected node's USERINFO into its
   error, generate/outbound.go warns it, apply/warnings.go logs it, and the last
   32 KiB of that log is section six of the bundle — with LogToFile on by
   default. diagSubCacheSecrets now reads those files by the same closed
   positive-list rule (unknown JSON key => collected, so a field added to
   model.Node tomorrow is covered), and a file it cannot read is NAMED in the
   bundle instead of silently reducing the scrub.

   Fixing the first half exposed the second: the log carried the userinfo, not
   the whole URI, so a literal scrub of the URI walked past it. diagSecretParts
   expands every refused value into its userinfo, username, password, query
   values (encoded and decoded) and path. Not the fragment — in a share link
   that is the node's display name, which is on the printable side.

The banner no longer says secrets are masked "throughout". It says what is
masked, and then names what is still in there: values under 8 characters (masked
in the config, not scrubbed elsewhere), list/ruleset URLs, and the limits of a
literal scrub.

Mutation-checked, each with the control that the instrument SEES the planted
secret in the unfixed output:
  resolver.address back on the allow-list      -> resolver test fails on the id
  diagMaskAddress made the identity function   -> transform test names the field
  sub-cache literals withheld from the scrub   -> log-scrub test fails
  diagSecretParts reduced to the whole value   -> log-scrub test fails
  sub-cache safe list turned into a blocklist  -> closure test fails
  unreadable cache file swallowed              -> honesty test fails
  nil collector seam read as "nothing to do"   -> honesty test fails
  transform applied per section, not per key   -> new-field test fails
  one section given an open default            -> sweep fails, by name
  an allow-listed value over-masked            -> sweep fails, by name
  masked lines dropped entirely                -> sweep's vacuity guard fires

scripts/run-tests.sh green (all 7 steps, -race included).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 13:49:08 +03:00
omarandClaude Opus 5 89571bcdb1 fix(netplane,generate): a network with option udp '0' had its UDP dropped in silence
Two defects of the same family: a state the plane produces and nobody names.

1. The divert is written per PROTOCOL, the fail-closed drop per INTERFACE.
   A tproxy inbound with `option udp '0'` puts its device in the drop scope
   (nftDivertRefs does not look at the flags, and must not: the drop is the
   backstop for ESP/GRE/SCTP too) while emitting no UDP TPROXY line for it.
   With kill_switch=closed every outbound UDP packet from that network is
   dropped; with kill_switch=open the same packets leave the WAN in the clear.
   TCP works, DNS works (dnsmasq answers it past the fib-local bypass), so it
   presents as "some sites do not load", not as a firewall. Verified by
   rendering: no second LAN is required, the shipped one-inbound shape does it.

   The drop is NOT narrowed to match the divert. Doing so would turn
   `option udp '0'` — which is how you kill QUIC so the engine can route by SNI
   — into "UDP now bypasses the proxy", i.e. it would convert a QUIC-blocking
   config into a QUIC-leaking one, and it would open a per-protocol hole in the
   kill switch through a knob whose name says nothing about leaking. The plane
   already takes the other decision one field over: with ipv6 off no v6 divert
   is emitted and closed mode drops v6 anyway, deliberately and in writing.
   So the state stays and is named instead, in three shapes (protocol dropped /
   protocol leaked / both flags off), each naming the network, the option, the
   kill-switch state and the concrete traffic that dies.

   coverage.go could not have caught this: it skips covered[i.Device], and the
   device IS covered. The new check is derived from the model alone and so runs
   outside that file's Interfaces() gate.

2. networkList's open `default:` sent (tcp=0, udp=0) to "" — which the engine
   reads as BOTH — so an inbound the plane feeds nothing acquired a listener for
   everything. The four cases are now named and closed, "neither" is a second
   return value rather than a synonym for "both", and a tproxy inbound that
   carries no protocol is refused with a warning that also names the netplane
   half: switching both flags off does not remove the network from the plane, it
   removes the way out of it, so closed mode cuts that network off completely.

generate_test.go: the three linux fixtures that built a tproxy inbound with
model.Inbound's zero-value flags now spell TCP/UDP out. UCI defaults both to
true; only a Go-built model gets false, and only that fixture relied on it.

Gate green (bash scripts/run-tests.sh, exit 0). Every new test mutation-checked
in both directions: suppressing the warning fails 6 tests by name, and making it
fire unconditionally fails the controls. Rendered ruleset text is byte-identical
for all nine shapes dumped before/after — the only diff is added warning lines.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 13:48:54 +03:00
omarandClaude Opus 5 6ced96fafa fix(panel): a switched-off subscription is not "never fetched", and an inline list is not "empty"
Two things the panel asserted that the code does not do.

1. THE OFF SWITCH IS NOT A GATE ON FETCHING. The row badge for a disabled
subscription with `fetch_via=proxy` and no detour was drawn quiet and said
"Nothing is disclosed yet — this subscription is switched off … so it is never
fetched". False in all three places that could have contradicted it:
Applier.UpdateSubscription resolves a subscription BY NAME and has never read
Enabled; `shaterd sub update` consults Enabled only when no name is given; and
this panel's own per-row "Fetch now" — new in this wave, previously buried in the
collapsed Options panel — is disabled on `busy || fetching` and nothing else. One
click sent the router's real address to the feed host under a badge saying
nothing was disclosed.

The two halves of the old condition are not alike, so they stopped being one
state. NO URL is real and refused at the bottom (subscribe/fetch.go rejects an
empty URL before it builds a request) — that branch keeps its quiet badge. OFF is
amber, and its sentence says what the switch actually does: it stops the
scheduled refresh, and the button on the row asks for a fetch whatever the switch
says.

The daemon reached the same conclusion from its side in this wave — the fetch is
deliberately allowed and logged, and its finding now varies on Enabled — so the
badge's own summary over that finding varies the same way. "On every scheduled
refresh" printed over a switched-off row is the same lie inverted: it sends the
reader hunting a cron job that is not running.

2. AN INLINE LIST HAS NO ENTRY COUNT, AND "NOT PUBLISHED" IS NOT "EMPTY".
engine.go fills RuleSetStat.RuleCount from (*rule.RemoteRuleSet).RuleCount(), and
LocalRuleSet.RuleCount does not exist in the tree at all, so an inline list always
arrives with rule_count 0. The chip called a working parental-control list
"empty — nothing matches". It reads "size unknown" now: unlit, never green and
never the amber that says something is wrong. A mixed group is counted as a floor
("1,284+") instead of presenting a partial sum as the whole.

BOTH INSTRUMENTS WERE HOLDING THE LIE UP. subFetch.test.ts pinned the sentence
verbatim, and deviceLists.test.ts fixed `{remote:false, rule_count:3}` — a record
no router can produce, so its green light was wired to nothing. The mock carried
the same impossible state on three local rule-sets and fabricated a count on
update. All of them now match what the daemon sends.

Verified in ?mock (new `?subleak=paused|pausedapplied|nourl`), 390 and 1280, both
themes. Each fix reverted in turn with the failure text; controls both ways — a
remote list that really is empty still says "empty", and the two leaking states
are still told apart from the one that is genuinely quiet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 13:35:53 +03:00
omarandClaude Opus 5 42536675e2 fix(panel): read the age, the scope and the hop — three fields the panel was ignoring
The daemon changed under the panel in three places, and in each one the panel
kept drawing a screen that was right only by accident.

byedpi readiness carries `age_seconds` now, because GET /api/status stopped
probing: sixteen enabled instances on ports that neither accept nor refuse cost
6.41 s per poll, measured, and the Apply page polls up to 27 times a minute. The
report is served from a cache and every sentence in it is present tense, so the
panel stamps it. Negative is not an age — zero is the common answer (a loopback
connect finishes in microseconds) — so "not a measurement" is its own reading,
and a daemon too old to send the field is a third one: the reading is real, its
age is not reported. The cold first poll after a start says "not measured yet"
rather than "not determined": nobody has looked is a normal state of a router
that booted ten seconds ago, and it calls for a different sentence than an
instrument that looked and failed. Both keep the unlit lamp and both keep the
egress type locked.

A test run no longer wipes the board, so a card can show a reading from twenty
minutes ago beside one from a second ago. Which is which comes from `scope`, not
from comparing timestamps — the router has no RTC and a computed "n minutes ago"
would be fiction. A carried row says "earlier run" and is drawn as a qualifier;
an empty scope is "cannot attribute", never "everything is carried", because a
real run always covers at least one target.

A chain blocked at a hop was kept red by matching a fragment of the daemon's
error sentence — the last place prose decided anything here. It arrives as
`blocked_by` now, so the match is gone and the row names the hop.

The mock carried the old contract: it emptied the board on every run while a
comment claimed the daemon did too. It carries forward now, by (name, kind),
capped at 64, and `?mock&board=carried` lands on a finished board holding both
kinds of row. `?mock&byedpi=cold` and `?mock&byedpiage=N` reach the two states
the freshness rendering exists for.

Verified in ?mock at 390 and 1280, both themes, no horizontal scroll. Each of the
three fixes was reverted in turn and the tests named the failure; the controls
run the other way too — a helper that marked every row carried, or reddened every
row, fails just as loudly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 13:35:25 +03:00
omarandClaude Opus 5 a4ea5dba44 fix(tests): the last two packages that wrote to the router's own /etc/shater
17843be5a measured six packages damaging the machine that runs the suite and
fixed four; model and generate were left because another agent held those trees.
Measured again on 2026-07-27 with the same instrument, scoped to the two
packages, and the diagnosis held EXACTLY:

    CREATED   /etc/shater/config.pre-unreadable.bak
    CREATED   /etc/shater/config.pre-v0.bak
    CREATED   /etc/shater/config.pre-v1.bak
    CREATED   /etc/shater/config.pre-v2.bak
    MODIFIED  /etc/shater/cache.db

model. Every test that reaches writeUCIWith or migrateWith goes through a
fakeUCI, and that seam is what makes the config they read and write a fake.
backupBeforeChange is the one part of the package that does NOT use it: it
os.ReadFile's liveConfigPath and writes into configBackupDir directly. On a dev
box neither exists and the function returns "nothing to copy"; on the testbed and
the router both exist, so the suite planted four bogus copies in the product's
state directory. Worse than litter: the function is create-ONCE per schema and
never overwrites, so a copy planted by a test SILENTLY PREVENTS the real
pre-migration copy that box was going to take.

generate. generate.go emits experimental.cache_file with Path: cacheFilePath(),
and every *_linux_test.go that hands a generated config to engine.Apply/box.New
opens that bbolt DB for writing. Per cache.go's own file comment that DB is a
SAFETY device, not an optimisation: with it, RemoteRuleSet.StartContext skips the
start-time fetch, so the daemon can come up before the WAN does. Rewriting it
from a test is rewriting the thing that keeps a reboot from taking the LAN down.

THE FIX is the one the other four packages already use, not a third one: a
TestMain per package pointing the product paths at a private os.MkdirTemp, plus a
test that still pins the SHIPPED value — because an isolation that leaves the
real decision untested has only moved the defect.

  model:    liveConfigPath/configBackupDir -> a private dir; liveConfigPath is
            pointed at a path that does NOT exist, which is exactly the dev-box
            case the function already documents, so every test that does not opt
            into backupSandbox behaves precisely as before.
            New TestConfigBackupPathsAreTheShippedOnes.
  generate: cacheDirPersistent/cacheFilePersistent/cacheFileFallback -> a private
            dir, and the persistent one is CREATED so the package keeps
            exercising the branch the ROUTER takes. The fallback had to move too:
            on a host without /etc/shater the decision lands on
            /tmp/shater-cache.db, which is just as hardcoded and just as much the
            product's. New TestCachePathsAreTheShippedOnes, which also pins that
            the DB lives inside the directory the free-space checks measure —
            path.Dir, not filepath.Dir, since the gate also runs on Windows.

No waiver was needed at shater/testguard: it follows
`cacheDirPersistent = filepath.Join(dir, ...)` back to os.MkdirTemp on its own.

Verified:
  - the sweep, scoped to the two packages: the five paths above BEFORE, "CLEAN"
    AFTER. Then the FULL scripts/check-test-fs-isolation.sh: 48 package
    verdicts, "CLEAN: the whole suite ran and not one path under /etc /var /usr
    /root /home /opt /srv /run /tmp changed."
  - positive control: a planted test in shater/model that restores the real
    paths and calls backupBeforeChange -> the sweep names
    "CREATED /etc/shater/config.pre-v9.bak", then bisects to "PACKAGE
    .../shater/model" and "TEST ....TestPlantedViolatorWritesTheRealBackup".
    shater/testguard stayed GREEN with the violator in the tree, which is the
    documented blind spot and the reason the dynamic half exists.
    NOTE, learned from the first attempt: a create-ONCE violator is named by the
    verdict but NOT by the bisect — seed_canaries only creates what is missing,
    so the file the whole-suite run left behind makes the per-package re-run a
    no-op ("no single package reproduced it"). The bisect can only name defects
    that repeat.
  - mutation, model: liveConfigPath -> /tmp/shater-live and configBackupDir ->
    /tmp each fail the new test by name; dropping the "keep the older copy"
    return fails TestBackupBeforeChangeKeepsTheFirstCopy ("the first copy was
    overwritten by a later write"); removing the ErrNotExist early return fails
    TestBackupBeforeChangeSkipsWhenThereIsNothingToCopy; removing the
    backupBeforeChange call from writeUCIWith fails
    TestWriteTakesTheBackupBeforeReplacingTheConfig ("the write took no backup").
  - mutation, generate: cacheDirPersistent -> /tmp/shater and cacheFilePersistent
    -> /etc/shater-cache/cache.db each fail the new test, the second one also on
    the dir/file mismatch; cacheFilePath forced to tmpfs fails
    TestCacheFallsBackWhenDirMissing's CONTROL, forced to persistent fails its
    first half; cache_file Enabled=false fails TestCacheFileEmittedAndEnabled.
    Green again after every revert.
  - counts, declared vs executed (go test -list vs top-level verdicts, shipped
    tags, linux): model 160/160, generate 395/395, 0 failures. generate's 3 skips
    are the pre-existing CAP_NET_ADMIN TestIntegrationL3* trio, which [5/7] runs
    and passes.
  - scripts/run-tests.sh: GREEN end to end, exit 0, including [4/7] under -race
    ("OK [race] in 43s") and [5/7] RAN all three privileged tests. The four
    TestByeDPI* races reported earlier no longer fire.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 12:56:30 +03:00
omarandClaude Opus 5 ea34e744bd fix(panel): the status poll stops dialling — a readiness cache that carries its age and has an owner
GET /api/status called byedpiProbe on every request. On a healthy loopback that
is nothing, but the probe's cost lives in exactly the state it was written to
report honestly: a port that neither accepts nor refuses burns the full
byedpiDialTimeout, and byedpiMaxInstances of them burn 6.4 s. Measured, on this
tree:

  1 enabled instance, live listener   0.45 ms
  1 enabled instance, refused         0.33 ms
  16 enabled, refused                 3.6  ms
  1 enabled, black-holed            400    ms
  16 enabled, black-holed             6.41 s

The panel shell polls /api/status every 5 s on every page and the Apply page
adds its own 4 s poll, so the pathological state hung the panel for seconds at a
time precisely while an operator was trying to find out what was wrong. A check
that gets slow exactly when it matters is worse than one that is always slow.

The poll now reads a cache (byedpiReadiness.cached), refreshed asynchronously off
the same path: 15 ns per call, primed, and 20 back-to-back polls against 16
black-holed ports cost less than one probe. A background ticker was rejected —
it would dial on a router whose panel nobody has open — and so was blocking the
first poll to fill a cold cache, since that is the same 6.4 s hang, just rarer.

The cache is not allowed to lie:

  - every served report carries age_seconds. Detail is written in the present
    tense, and a present-tense sentence about a measurement taken some seconds
    ago is a claim nobody checked;
  - a report older than byedpiCacheMaxAge is NOT SERVED. It is replaced by an
    explicit unknown with a negative age, so a panel that ignores the age fails
    to an unlit lamp rather than to a stale "listening" unlocking an egress type
    onto a port nothing is on;
  - GET /api/byedpi still really connects. A re-check button answered from a copy
    is a button that does nothing.

And it has an OWNER. The refresh runs a goroutine that dials; left as a package
variable it belonged to nobody, could not be awaited, and — as the race detector
showed — went on reading byedpiConfigPath / byedpiInstalled / byedpiDial after
whatever started it believed it was finished. The cache is now per-Server, with
stop() that forbids further refreshes and does not return while one is dialling,
called from Server.Close. The daemon already defers that Close, so the probe
cannot outlive the server.

byedpiDial became a seam alongside byedpiInstalled and byedpiConfigPath: the
timeout branch is the expensive one and the one a real loopback cannot be
provoked into, so without it neither the cost nor its removal could be shown.

Ten mutations, each killed by a named test: the probe back on the request path;
the cached copy claiming age 0; an over-age reading quoted anyway; a cold cache
returning a blank instead of an explicit unknown; a refresh that is not
single-flight; a late older probe overwriting a newer one; /api/byedpi answering
from the cache; the unmeasured report claiming the binary is absent; stop() not
waiting; Server.Close not stopping. The last one survived its first test — which
asserted a poll straight after Close did not dial, and passed with the stop
removed entirely because the reading was fresh and no poll was due — so the test
now advances the clock to make it due.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 12:55:34 +03:00
omarandClaude Opus 5 e5dfd74b61 fix(engine): testing one node stops wiping the group board; the blocked hop becomes a field
Two separate honesty defects in the shared test board, both surfaced while
closing the byedpi readiness fix.

1. startTestRun replaced the results slice outright, so pressing Test on one
   NODE blanked every group and chain card on the Targets screen, and testing a
   group blanked the nodes. Nothing on screen explained it, because nothing had
   happened to those targets — the daemon had thrown their readings away.

   Earlier results are now carried forward for every target the new run does not
   itself re-measure. The alternative, one board wiped per run, is simpler and
   has no staleness question at all; it was rejected because it destroys
   information the daemon still has. These are the OBSERVATORY's numbers, taken
   by a prober that never stopped, and a group's reading does not become false
   because somebody tested a node afterwards.

   The staleness question it does raise was already answered: every result
   carries tested_unix, the instant the OBSERVATION was taken, and GroupTestStatus
   publishes this run's scope — so a carried row is identifiable as carried
   without comparing timestamps, and drawn with its age. The board is capped at
   groupTestCarryMax, evicting the oldest first; that cap is the only way a row
   can leave without a newer one taking its place, and it is documented as such.
   done/total still describe this run's targets only.

2. A chain whose exit was never dialled, because an earlier hop was probed and
   did not answer, shipped that fact as prose only: source="" (correct — nothing
   measured the exit) plus a sentence naming the hop. A client reading source
   strictly filed it under "nobody looked", which is the wrong colour, so the
   panel kept the row loud by matching a fragment of our error message — the
   last place it read our prose to decide anything.

   GroupTestResult now carries blocked_by: the 1-based hop index, 0 everywhere
   else. It does NOT set source; nothing measured this target's own path, and
   stamping an instrument on a measurement that never happened is exactly the lie
   source was added to prevent. blocked_by>0 beside source="" is the complete
   statement. Field and sentence are produced together in chainBlockedResult so
   they cannot come to disagree.

Tests (grouptest_board_test.go), each verified by mutation:
  - carrying forward is asserted WITH its control, that a run does replace the
    rows it covers — "nothing disappeared" alone is also satisfied by a board
    that stopped updating;
  - target identity is (name, kind), so a group and a node of one name do not
    evict each other, with the empty-kind wildcard pinned both ways;
  - the cap drops the oldest end;
  - blocked_by carries the hop, keeps source empty, and every other result
    carries 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 12:55:11 +03:00
omarandClaude Opus 5 17843be5ad fix(tests): running the suite deleted the router's own state — four packages did it
shater/stats/store_test.go ended with

    _ = os.Remove(statsFilePath())

and statsFilePath() is not a test path. It is THE product path: /etc/shater/stats.db
on every host where that directory exists, which is the testbed and the router. So
`go test ./shater/...` deleted the accumulated query and connection log of whatever
machine ran it. The test passed. It had always passed — damage done by a test is a
side effect, not a wrong answer, and no instrument in this tree could see one.

A filesystem sweep (the new scripts/check-test-fs-isolation.sh: seed a router-shaped
canary tree in a container, run the whole gated suite, diff) found it was not alone.
Six packages, by measurement, not by reading:

  shater/stats    DELETED  /etc/shater/stats.db        (the line above; also
                           TestComboBackendSwitchSequence opened and pruned the
                           live DB, which the delete had been hiding)
  shater/logsink  DELETED  /etc/shater/shaterd.log and /var/log/shaterd.log —
                           New()/Reconfigure() purge BOTH product locations when
                           the file toggle is off, so Config.Path (which every test
                           here already set) never protected them. The daemon's own
                           log, the one an operator reads after an outage.
  shater/apply    DELETED  /var/run/shater.active — the ONE token hotplug and cron
                           check before touching the data plane. Clearing it on a
                           live router makes both stand down on a box that is up.
                           holdstate_test.go's `t.Cleanup(os.Remove(ActiveFlag))`
                           was not a cleanup; it was the delete.
  shater/panel    REWROTE  /etc/shater/stats.db — stats.NewStore("sqlite") from
                           TestStatsEndpointsAcrossBackends resolves the product
                           path too.
  shater/model    CREATED  /etc/shater/config.pre-v{0,1,2}.bak, config.pre-unreadable.bak
  shater/generate REWROTE  /etc/shater/cache.db

The last two are NOT fixed here — another agent is working in those trees. Both are
one TestMain away: model already has liveConfigPath/configBackupDir as vars, and
generate already has cacheFilePersistent; what leaks is product code (backupBeforeChange,
the engine's cache_file) called from tests that do not redirect them.

THE FIX is the seam generate/cache.go and generate/ruleset.go already use — the path
becomes a package-level var that only tests assign — plus, in each case, a test that
still pins the SHIPPED value, because an isolation that leaves the real decision
untested has only moved the defect:

  stats:   statsDirPersistent/statsFilePersistent/statsFileFallback + the exported
           SetPathsForTest (exported because shater/panel needs it from outside).
           New TestStatsFilePathPrefersPersistentDir covers both branches.
  logsink: PersistPath/TmpfsPath + a TestMain, since the hazard is in New(), which
           every test calls. New TestLogPathsAreTheShippedOnes.
  apply:   ActiveFlag + the existing TestMain. New TestActiveFlagIsTheShippedPath,
           which also records WHY /var/run: tmpfs, so a reboot clears it.

TestNewStoreSelection got stronger rather than weaker. Its "sqlite" case used to
accept "sqlite" OR "memory" because the real path might not open on this host — an
expected value that depended on the machine. At a private path there is no excuse:
a writable directory MUST report "sqlite", and a new control at an unopenable path
MUST report "memory" (the honest "persistence is not active" signal) without a crash.

TWO GUARDS, because one of them cannot see half of it:

  shater/testguard/fsisolation_test.go — parses every _test.go under shater/ and
  fails BY NAME when a filesystem-mutating call gets a path that is not PROVABLY
  temp-rooted. Positive and closed: what it cannot prove is a failure, not a
  default, which is the only rule that catches a path built by a function call.
  It follows local vars, closures, filepath.Join/Sprintf/+, helper parameters via
  their call sites, helper return values, and the save/override/restore idiom.
  Four waivers, each keyed on file+function+callee, each with the reason printed on
  every run, each a struct field traced by hand; a waiver that stops matching fails
  the test as STALE. Runs inside [2/7] and [4/7] — no new gate step, no new minute.
  Blind spot, stated: damage done by PRODUCT code a test merely calls (which is
  exactly logsink, model and generate above).

  scripts/check-test-fs-isolation.sh — the dynamic half, for that blind spot. It
  refuses to run outside a container unless told twice, because its method is to
  let the damage happen and then look, and it seeds/unseeds only what was missing.

Verified:
  - mutation, task 1: statsFilePath forced to the fallback -> the new path test
    fails ("with ... present = .../fallback-stats.db, want the persistent ...");
    newPersistent forced to memory -> "Backend = \"memory\", want \"sqlite\"";
    the fallback made to report "sqlite" -> "Backend = \"sqlite\", want \"memory\"".
    Green again after each revert.
  - mutation, the guard: the original os.Remove(statsFilePath()) put back -> named
    at store_test.go:154 with "the path comes out of statsFilePath(), which this
    check cannot follow"; a planted test writing "/etc/config/network" -> named as
    a literal path; the walk pointed at one package -> its own <150-file control
    fires ("reading a blank page"); a waiver matching nothing -> STALE WAIVER.
  - control, the sweep: with a planted violator it reports DELETED /etc/shater/stats.db
    and MODIFIED /etc/config/network; without it, those are gone and only the two
    foreign packages remain. Its bisect named shater/stats.TestComboBackendSwitchSequence
    on its own.
  - counts, declared vs executed (go test -list against top-level verdicts):
    stats 112/112, panel 121/121, apply 122/122, logsink 26/26, testguard 1/1,
    0 skips, 0 failures.
  - scripts/run-tests.sh: [1/7][2/7][3/7][5/7][6/7][7/7] green. [4/7] -race fails on
    four TestByeDPI* in shater/panel — a data race between byedpi.go's background
    probe and byedpi_test.go's forceByeDPIBinary cleanup, in another agent's
    uncommitted work (shater/panel/byedpi_cache_test.go is untracked). Proven not
    ours: a pristine HEAD tree carrying ONLY this commit's files passes -race over
    all 35 packages, and the same run with -skip ^TestByeDPI is green on the live
    tree too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 12:33:45 +03:00
omarandClaude Opus 5 dac2f85c84 fix(insights): a failed DNS lookup stops being the healthiest row in the log
The DNS log drew every row from `action`, which answers WHICH WAY the lookup
went — so a query that left through a detour and then timed out came back as an
accent-coloured `proxy` tag, blocked=false, nothing else said. The row that
describes the exact moment the tunnel broke was the most reassuring line on the
page. `actionTag()` also fell open (`return 'pass'`), so every value the panel
did not recognise — including every value a future daemon might add — rendered
green.

The daemon now carries the outcome as its own axis (LogEntry.Status/Error,
a794fbe37). This brings it to the screen.

Two axes, and the outcome leads. logRoute.dnsRowMark decides both in one place:

  status  → answered | failed | '' (not recorded), POSITIVE and CLOSED, with the
            fallback on the recoverable side. `blocked` refines a recorded answer
            into the fourth situation and is never allowed to invent one on a row
            whose outcome was never written.
  action  → block | proxy | pass | unknown, the same discipline. The path stays
            VISIBLE on a failed row and muted, because "it failed" and "it failed
            in the tunnel" are different reports and the second one closes tickets.

Four situations, four looks: a plain answer has no rail; a filter block keeps its
crit rail and BLOCK tag; a failure takes an amber rail, an amber wash, a filled
FAILED chip, and its cause verbatim beside the rcode reading (-1 renders "no
response", anything else the code the server really sent); a not-recorded outcome
is dashed and faint and claims nothing. A failure with no recorded cause says
"cause not recorded" rather than showing an empty cell that reads as fine.

`error` joins the searched fields (the daemon searches it — q=timeout works) and
the hint under the box now names it. `status` stays out: q=failed must not sweep
up every failure while somebody is looking for a domain by that name.

Verified in ?mock at 390 and 1280, both themes, no horizontal scroll: the four
states are pairwise distinct in computed border/background/colour, and the two
chips take their own line on a narrow screen so the domain keeps 92px instead of
being pinned at its 30px minimum. 19 new tests, each shown to fail under 16
mutations of the code it covers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 12:25:09 +03:00
omarandClaude Opus 5 e065786a32 feat(panel): unlock byedpi on a listener, test one node, and stop two badges arguing
Three things the panel was saying that it had not established.

BYEDPI. The egress type unlocked on `status.byedpi_installed`, which is
LookPath("ciadpi") — "is the package installed", while the operator is asking
"will traffic sent here go anywhere". They come apart on the SHIPPED config: the
packaged /etc/config/byedpi is inert, so installing the package unlocked the
type, the egress went on 127.0.0.1:1080, the apply was green and nobody was
listening. The gate is now `byedpi.state === 'listening'` and nothing else
(byedpiReady.ts, the only place that decides it). GET /api/byedpi also carries
the per-egress port cross-check, so a mismatch is drawn on the row that has it,
naming both ports, in crit — the state where every other signal reads healthy.

`unknown` is neither answer. It keeps the type locked (a control that opens on
nothing established is the same defect wearing a new word) and it is never
painted as a refusal: dashed border, unlit lamp, "not determined", plus a
Re-check button so a dropped request is not a dead end.

ONE NODE. A freshly pasted node had no instrument — the group test reads the
observatory's board and the observatory only probes what the rules route
through, so the first question anyone asks answered "not routed by any enabled
rule". Every node row now has Test, over the same singleton run and the same
GET poll the Targets page uses.

The reading is classified on `source`, not on prose: measured-and-failed is red,
`source:""` is an unlit lamp and the faintest text on the row, because a
negative result that cannot be told from a check that never ran answers nothing.
One escalation survives, documented and narrow: a chain whose exit was never
reached because a hop it runs through WAS probed and failed. Targets keeps its
exact previous appearance while its instrument changes underneath.

The three refusals stay three facts — 404 the node is not in the saved config,
503 the config could not be read (an unknown, never a verdict about the node),
400 no name — with three tones and three sentences.

SUBSCRIPTION FETCH ROUTE. The panel's draft predicate drew amber "proxy · no
route" while the daemon now grades the same fact critical on the same row, so a
saved leaking subscription wore both, at two severities, about one thing.
subFetch.ts reconciles them: where the daemon has spoken it outranks the
prediction, in BOTH directions — including the dangerous one, where the
predicate is content ("via group:auto") and the daemon reports the leak anyway.
The prediction still speaks for a draft nothing has applied yet, and a
subscription that is switched off is not accused of a disclosure the daemon
deliberately does not report for it.

Fixtures reach every state: ?byedpi=<five states>|mismatch, ?nodetest=ok|dead|
unmeasured|400|404|503, ?subleak=draft|applied|divergent. The group-test fixture
also stopped being kinder than the daemon — engine.startTestRun replaces the
whole board, so refreshing one target really does blank the others.

42 tests, 8 mutations each killed by name, and both controls: the gate is shown
to open on `listening` and to stay shut on the other four, and "not checked" is
shown to be drawn differently from "did not answer" — the assertions fail if
either pair is ever drawn alike. Verified in the browser at 390 and 1280, both
themes, no horizontal scroll.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:55:42 +03:00
omarandClaude Opus 5 46a2d4aaad test(model): pin the wire contract the panel's PUT actually sends
handleConfigPut decodes model.Model with DisallowUnknownFields, so this is the
layer the missing fields bit at: not "the attachment is ignored" but "the whole
save is rejected with json: unknown field \"Blocklists\"", losing every
unrelated edit batched into the same PUT. Mutating the JSON name reproduces
that message exactly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:40:32 +03:00
omarandClaude Opus 5 67290ec2b6 feat(devices): attach named block/allow lists to a device, without a dangling tag
The panel half already ships Device.Blocklists/Device.Allowlists and PUT
/api/config decodes with DisallowUnknownFields, so until now the first
attachment rejected the WHOLE save with `json: unknown field "Blocklists"`,
losing every other edit in it. This is the engine half.

A device now references `config blocklist` / `config allowlist` sections by
name (UCI: `list blocklist` / `list allowlist`, since `block`/`allow` already
mean the typed domains), which brings geosite categories and url-sourced lists
to parental control for free.

Two things here are constructions, not checks.

The tag a device's rule references comes from the accumulator that emitted the
rule-set, never from the list name. Rule-set tags resolve at engine START
(RuleSetItem.Start), so a name-derived tag passes box.New and fails box.Start —
and because both configs share one cache_file path, every apply on a live
engine takes the close-old-then-start-new branch, so the old box is already
gone when the new one refuses. That is no engine, a closed kill switch and a
dark LAN, from one mistyped list name. A reference that yields no tag emits no
rule at all; the emptiness is warned, tagged DEVICE-FILTER-NOT-APPLIED so the
panel grades it critical rather than guessing from prose.

Materialisation is a single memoised point shared by both consumers. Devices
are built before the network-wide filter, so materialising a shared list twice
would hand dedupeRuleSetTags an already-claimed tag — which it DROPS, silently
switching the network-wide filter off for that list. The mutation test for this
reproduces exactly that: DNS-FILTER-NOT-APPLIED, filtering nothing.

Order is the feature: typed allow, typed block, attached allow, attached block,
then the network filter. Otherwise a parent who types youtube.com into a
child's Block loses to whatever an attached geosite category permits, and the
panel draws a "Blocked" chip over a rule that does nothing. Typed and attached
matchers stay SEPARATE rules — rule_set AND-gates over the domain matchers, so
merging them would mean "the domain AND the list".

Attaching a list is itself the switch for that device: Enabled=0 means "does
not participate in the network-wide filter", not "dead", so the list still
loads and filters here — and generate says so instead of leaving it to be
discovered. A blocked name's reply comes from the LIST (Blocklist.Response),
so tier 4 is up to two rules; the typed tier keeps NXDOMAIN, having no owning
object to say otherwise.

Purely additive: an old config has neither list, parses to nil, and the
generated engine config is byte-for-byte what it was. No schema bump.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:37:33 +03:00
omarandClaude Opus 5 9339e8e70d fix(generate): the cache fallback test ran or skipped depending on its neighbours
TestCacheFallsBackWhenDirMissing guarded itself with

    if fi, err := os.Stat(cacheDirPersistent); err == nil && fi.IsDir() {
        t.Skipf("%s exists on this machine; ...")
    }

i.e. its subject was the machine it happened to run on. The gate caught it as an
UNDECLARED SKIP in one container run and not in the next, with no change to the
code — and BOTH outcomes were green. Only the undeclared-skip check saw it at
all; every other instrument here reports `ok shater/generate` either way.

An order-dependent test proves nothing on the runs where it does run either,
because nobody can tell afterwards which runs those were.

The three cache paths become vars (production never assigns them, same seam
generate/ruleset.go already uses for listsDirOverride) and the test points them
at a temp tree. It now covers BOTH branches with no skip: an absent dir must
choose tmpfs, and — the control — a present one must choose the persistent
path. Without that second half the test is satisfied by a cacheFilePath that
returns the fallback unconditionally, which is exactly the regression the
persistent branch exists to prevent (a cache that never survives a reboot, so a
reboot before the WAN is up fails to start the engine and takes the LAN with it).

Verified:
  - both halves killed by mutation (force persistent -> the first assertion
    fails; force fallback -> the control fails), green again after revert;
  - 5 x `go test -shuffle=on ./shater/generate/`: 474 verdicts and 3 skips
    every time, TestCacheFallsBackWhenDirMissing PASS on all five, never SKIP;
  - shater/generate: 385 declared func Test*, 385 top-level verdicts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:27:40 +03:00
omarandClaude Opus 5 a67f51c22c docs(panel): the q= field list did not mention error, which the filter searches
shater/stats/filter.go:156 searches ConnLogEntry.Error along with the six
fields the doc names, so `q=timeout` works and the contract said it did
not. Verified against the predicate, not against a report.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:27:36 +03:00
omarandClaude Opus 5 56c9ea56e6 fix(gate): [4/7] reported a failure that did not exist — 664 s of pure sleep
The -race step failed with

    FAIL shater/netplane 600.019s
    panic: test timed out after 10m0s
      running tests: TestApplyIfaceSysctlsCoversRuleDivertedIface

over code that was neither hung nor wrong. Measured (golang:1.26, 32 cores):
shater/netplane is 1.971 s without -race and 663.762 s with it. A 337x factor
is not "-race is slower".

Nine of netplane's test files intercept nft/ip/ubus/uci/sysctl by re-exec'ing
the test binary as a no-op helper — the standard os/exec trick. Under -race
that child is ThreadSanitizer-instrumented, and TSan's atexit_sleep_ms DEFAULTS
TO 1000: every -race process sleeps a flat second before exiting, on no CPU.
~660 intercepted commands, one second each. The per-test times said so out
loud — 12.17 / 13.18 / 14.17 / 129.62 s — they were counting, not measuring.

Isolated, five runs each, of a `func main() {}` with nothing in it:

    built plain                       0.0014 s/run
    built with -race                  1.010  s/run
    built with -race, sleep disabled   0.008  s/run

So the children now run with GORACE=atexit_sleep_ms=0, set once in a package
TestMain rather than in each of the nine fakes (they all build the child env as
append(os.Environ(), ...), so one assignment covers the ones written later too).
TSan reads GORACE at process init, long before TestMain, so the detector of the
test process itself is untouched; only the children see it, and they do nothing
but write a canned string and exit. Proven, not assumed: a deliberate data race
in netplane is still reported under -race with this in place.

    shater/netplane   663.762 s -> 10.625 s   (203 === RUN and 128 top-level
                                               verdicts on both sides)
    shater/devices     28.412 s ->  0.358 s   (same disease, same cure)
    gate [4/7] end to end: was a 600 s timeout, now 56 s

WHAT THE GATE ITSELF WAS MISSING. A deadline and a failed assertion both exit
non-zero, and this script printed the same "FAILED [race]: go test exited 1"
for both — so the reader could not tell "the product is wrong" from "nobody
knows yet". [2/7]/[4/7] now name a timeout as a TIMED OUT, list the tests that
were still running, print only the goroutine dump instead of a quarter megabyte
of PASS lines, and spell out the two opposite fixes (a block, or slowness that
must be MEASURED first). Verified both ways: a sleeping test reads TIMED OUT, a
t.Fatal still reads FAILED.

The deadline stays at go test's own 10m, now written down with the measurement
beside it, and stays there as the hang detector — the slowest package under
-race is 18.9 s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:27:23 +03:00
omarandClaude Opus 5 cae5655dbe fix(panel): the hazard band said DNS leaves and that nothing leaves
Measured on the stand, real LAN clients in netns behind veth, counters in a
separate nft table: in the state this band predicts, 0 packets left the WAN
across the whole run, against 27 in the control that differs only by one added
catch-all rule. Except for exactly 2 — both plaintext UDP/53. So the band's
detail ("nothing reaches the internet") was wrong by those two packets, and its
own DNS step, which calls that lookup the one thing that still leaves, was
right. One word: nothing ELSE reaches the internet.

The DNS step was also behind reality. It named only the lookups devices send to
the ROUTER, but the shipped dns_intercept='1' pulls a query aimed at a resolver
the device picked for itself into the engine too, answers it there, and it
leaves in the same clear UDP/53 — measured both ways, each producing its own
plaintext packet on the WAN. Encrypted DNS is not the way out either: :853 out
of the LAN measured connects=0, because the plan rejects it. The generator's own
critical warning (generate/dns.go) has said all of this for as long as it has
existed; only the panel had fallen behind it.

"takes ... and drops it" is untouched, and measured: the engine accepts on the
local tproxy socket in ~100 us even for an unreachable address and then closes,
so the client gets an immediate ECONNRESET rather than a hang. "Blocks" and
"ignores" would both be less accurate. Nothing is added about ping: the stand's
ICMP probe was 100% loss in BOTH states, so it proved nothing either way.

The test is the point. The two halves live fifteen lines apart and each reads
fine alone, so a wording fix does not survive the next editor. The new test
checks the INVARIANT instead: the band is flattened to clauses and no clause may
claim that nothing leaves while another names something that does. Its detector
is proved on a fabricated band first (a prior that cannot fire measures
nothing), and it asserts the no-resolver band really does contain a clause
admitting the leak, so the check cannot pass by finding neither half.

Mutations, all caught: detail back to "nothing reaches" -> the invariant fails
and prints both clauses verbatim; DNS step back to the router-only wording ->
the resolver test fails; DNS step stops admitting the leak -> two tests fail.
Control: with one resolver configured the DNS step is absent and no clause
claims anything leaves; emitting the step unconditionally fails that control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:23:52 +03:00
omarandClaude Opus 5 fbcf211d19 test(stats): pin that the log filter does NOT search status
matchLog's field list is documented as positive and closed, and the new
`status` is deliberately outside it for the same reason `outbound_kind` is: it
is a fixed vocabulary word, so q=failed would silently match every failed row
while the operator was looking for text. The test row now carries a Status, so
the assertion is not vacuous — a filter block IS an answer, which is also the
Status/Error invariant this row models.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:18:10 +03:00
omarandClaude Opus 5 419eaf7bbe docs(porting): the schema number on line 146 was v0.1's, read as v0.2's
PART A is the frozen v0.1 survey, so `CurrentSchemaVersion=1` was archaeology
that happened to be correct about the branch it describes — and directly
contradicted the live schema subsection thirty lines below, which says
`shaterd migrate` writes 2. Anyone skimming the file map for "what is the schema
version" got 1. Say whose number it is, name v0.2's (2, steps {0->1, 1->2}), and
name what migrate1to2 did, since that is what the reader is usually after.

Also documents the `shaterd migrate` reporting contract in PART B: the closed
classification, the two non-syslog channels a failure reaches the operator on
with globals.log_syslog=0, and why 30_shater-core still exits 0 after one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:13:13 +03:00
omarandClaude Opus 5 a794fbe374 fix(stats): a failed DNS lookup no longer reaches the log as a healthy row
dnstrack.QueryEvent carries Failed and Error; stats.LogEntry carried neither.
A SERVFAIL, a timeout, a loopback or a rejected-cached lookup was therefore
written into the query log with action "pass" — or, when the resolver that
timed out had a detour, with the flow-coloured "proxy" — blocked=false, and
nothing anywhere saying no answer was produced. The daemon already knew, one
event at a time: TotalStats.Failed is counted from that very fact in the same
function. The row threw it away, so the aggregate said "N failed" while every
row said everything was fine.

LogEntry gains two fields:

  Status — closed vocabulary, "answered" | "failed" | "" (NOT RECORDED), same
    discipline as OutboundKind/RuleKind. It is a separate axis rather than a
    fourth Action value because Action says WHICH PATH the lookup took: a query
    that went out through a detour and then timed out is action=proxy AND
    status=failed, and folding the two would erase the one fact that says
    whether the tunnel is what broke. It is also what an old panel would have
    silently mapped back onto "pass" through its own open fallback.
  Error — the producer's own cause text, verbatim, meaningful only when
    Status=="failed". No grading is invented on top: three of the four causes
    are fixed literals ("loopback", "rejected (cached)", "rejected") and the
    fourth is the transport's err.Error(), which cannot be classified without
    guessing. "failed" with an empty Error is honest and reachable — the lookup
    failed and the cause was not recorded. What IS derivable stays derivable:
    Rcode separates "no response at all" (-1) from "the server refused".

queryStatus is a closed POSITIVE list over the sources a producer emits; an
unlisted or zero Source falls to "" (not recorded), never to "answered". The
aggregate is untouched: blocked/failed are computed once in handleEvent and the
row is labelled from those same two values, so the counter and the row can
never disagree and nothing is counted twice.

Cost: LogEntry 152 -> 184 B on 64-bit (+6.4 KB at the default 200-row ring).
Status is a package constant, so its body costs nothing; Error is interned in
its OWN table (maxErrKeys=128, clamped to 160 B) rather than the rule table,
because the transport's error text embeds the queried name and a flood of
distinct causes would otherwise keep clearing the routing-text table.

Tests: every assertion mutation-checked, and the control is three-state — the
same instrument separates answered from blocked from failed, with the
aggregate pinned to identical totals across the change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:12:59 +03:00
omarandClaude Opus 5 735aa5428f fix(shater-core): name which of the four ways shaterd migrate ended
Both call sites swallowed the result. /etc/uci-defaults/30_shater-core ran
`shaterd migrate >/dev/null 2>&1` — stdout, stderr AND the exit status gone, so
a refusal was indistinguishable from a success on the one screen the operator
who caused it was reading. /etc/init.d/shater logged, but with a single sentence
that described only one of the outcomes: "routing rules that still carry the
removed dst_domain/dst_ip options stay DISABLED until this succeeds. Free space
on /overlay and re-run". On a DOWNGRADE every clause of that is false — nothing
is disabled, /overlay is not the problem, and re-running never helps, because
the fix is to put the newer package back. A confident wrong diagnosis costs more
than no diagnosis.

The outcome is now classified with a CLOSED positive list — ok / downgrade /
unreadable / failed — and the last rung is the point of it: an unrecognised
failure says it is unrecognised and quotes the binary verbatim instead of being
reported as one of the causes we can name. `downgrade` is recognised by the
substring "newer than this build", which both model.migrateWith's refusal and
model.ErrSchemaTooNew contain; that seam is a contract and is now pinned.

log_syslog=0 is honoured, not worked around. It is a statement about the syslog
stream, not a request to be left uninformed, so failures go to two channels that
are not syslog: the script's own stderr (the operator's terminal on a hand-typed
restart; the package manager's output inside `apk add`), and
/etc/shater/migrate-failed on flash — written on failure, REMOVED on the first
success, so its absence is the honest all-clear. syslog gets the same line when
log_syslog allows it. A migration that SUCCEEDED stays routine.

uci-defaults still exits 0, deliberately: a uci-defaults script that does not is
kept and re-run at every boot, and this one re-runs a detached enable+restart of
shater/shater-cron plus a firewall reload — one recoverable failure would become
permanent boot-time churn, to carry a status nothing reads. The retry that
matters already exists in start_service, which runs the migration every start.

Found by mutation while writing the tests: reverting start_service's call site
left every other test green, because they all call shater_migrate directly. The
reporter would have been perfect and unreachable. TestInitScriptStartServiceUses-
TheReporter closes that.

Verified: sh -n and busybox ash -n on the target (ImmortalWrt 25.12.1 r37978),
the classifier exercised there under busybox ash against the real uci; six
mutations rolled back one at a time, each caught by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:12:54 +03:00
omarandClaude Opus 5 d1f43dbbe6 fix(shaterd): diag printed constant.Version instead of the version it was handed
renderDiag took a version through its seam and then ignored it, reading
constant.Version directly — so the one field the dead-daemon test could have
pinned was the one field it could not see change. The bundle now prints what it
was given, and the test asserts the value and not just the heading.

Also names the cost the schema gate adds: model.readDiskState's own comment says
"this runs once per write", and it now also runs once per apply, i.e. once a
minute from shater-cron — one `uci export shater` fork and two parses of a few
kilobytes. It reads the DISK rather than m.Globals.SchemaVersion deliberately:
m need not have come from disk (rollbackTo hands in an in-memory snapshot), and
the question is about the file this build would have to live with.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:11:06 +03:00
omarandClaude Opus 5 67c4829f03 fix(apply): name the fetch_via=proxy that is not going through anything
`fetch_via=proxy` with an empty `fetch_detour` is not "proxy, details to
follow". Applier.HTTPClient hands "" to resolveVia, which passes it through
(it is not a `chain:` selector), engine.ViaToTag maps "" to the tag `direct`,
and the feed is dialled through the box's direct outbound — over the ordinary
WAN, with the router's real address, merely from inside the daemon process
rather than from the CLI. Nothing fails. The subscription provider, the party
`fetch_via=proxy` is chosen to hide from, sees that address on every
scheduled refresh.

The picker exists and defaults to Direct, so the state is not "unconfigured";
it is "configured, and silently equal to direct". The message opens on that.

Critical, by this file's own rule at the top — protection the operator
CONFIGURED is not in effect — and by consistency: criticalMarkers already
grades the identical disclosure critical when generate says it about DNS
("in the clear", "your provider sees", "leaves over the plain WAN with your
real IP address").

A detour that names nothing is a SEPARATE finding at `warning`, because it
has the opposite consequence: resolveVia or the engine refuse by name and
UpdateSubscription returns the error rather than falling back, so nothing is
disclosed — what breaks is the refresh, loudly. One sentence for both would
send the operator to fix the wrong thing. A bare name that is really an
egress or a chain gets its own text giving the spelling that resolves, rather
than a false "nothing answers to that name".

Filed under section `subscription` + the sub's own name, which the panel
already routes to that row (Nodes.tsx entityFindings/findingsByName) and to
Overview. The severity is part of that binding, not just the volume: `info`
is filtered out of entity routing on purpose, so it would never reach the
row — recorded at the constant.

Also completes the FetchDetour contract in model.go, which listed neither
`chain:X` — the form apply.resolveVia has a dedicated branch for — nor what
"" actually does.

Verified: 9 mutations, each reverting one part, each caught by a named test;
controls show the same instrument silent for a resolved detour, for
fetch_via=direct, and for a subscription that is disabled or has no URL.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:09:30 +03:00
omarandClaude Opus 5 b43f673ad2 fix(apply,shaterd): a downgraded build may not RUN a config it can only half-read
model.WriteUCI already refuses to write a config whose schema is newer than the
build, so a downgrade can no longer eat the file. What was still open was
RUNNING one. ParseUCIExport ignores options it does not recognise — silently —
so a v3 config read by a v2 build yields a Model with the v3 settings simply
absent. The engine starts perfectly happily and routes traffic by a policy
nobody wrote. Nothing said so: cmdRun never called Migrate(), /etc/init.d/shater
calls it, logs one daemon.err line on failure and starts us anyway, and that log
defaults to a tmpfs file globals.log_syslog='0' can switch off entirely.

REFUSE OR START — and why refuse. Both sides, weighed by "the default falls to
the recoverable side":

  * REFUSE. With kill_switch=closed the fail-closed plane goes up and LAN->WAN
    forwarding stops. Loud, immediate, impossible to miss. SSH, LuCI and the
    panel stay reachable, nothing on disk changes, and reinstalling the build
    the router ran ten minutes ago puts everything back exactly as it was. The
    damage is an outage the operator caused themselves and can undo.
  * START ANYWAY. Traffic the missing rules were meant to tunnel leaves through
    the plain WAN with the router's real address on it, and nothing announces
    it. That is not recoverable in the sense that matters — the disclosure has
    already happened. It is the same choice `sub update` made when it was given
    FAIL over a silent direct fetch.

So: refuse. But the daemon does NOT exit and does not crash-loop — a refusal
nobody can see would be the third bad option. It stays up, keeps serving the
panel and the control socket, and says why in three places:

  1. apply.schemaDowngradeGate refuses every apply (step 0 of applyLocked), with
     the engine-swap failure policy of step 2: a previous engine that IS running
     a config this build understood is left alone; with no engine, holdLocked
     installs the fail-closed plane — and honours kill_switch=open, which is the
     operator's documented fail-open choice and may not be quietly overridden.
     This is in applyLocked and not only in cmdRun on purpose: cron reconciles
     once a minute, so a gate that only ran at startup would be bypassed sixty
     seconds later.
  2. Status carries the PAIR: schema_version (disk) and schema_supported
     (model.CurrentSchemaVersion). Either alone is unreadable — the panel
     already showed the disk version, and "v3" next to a build that understands
     v2 looks entirely normal. The difference IS the fault. schema_supported is
     a compile-time constant and is therefore set even on the offline stub, i.e.
     on the daemon most likely not to be answering. A critical warning naming
     the downgrade is computed at READ time, because in this state no apply can
     succeed and "the warnings of the last successful apply" would be empty.
  3. cmdRun consults model.Migrate() before reading the config (so a bare
     `shaterd run` gets the gate too) and classifies the outcome with
     CheckConfigWritable: ErrSchemaTooNew is the downgrade, anything else is an
     ordinary migration failure and is NOT reported as one.

Only ErrSchemaTooNew blocks. ErrUnmigratedConfig — schema-v1 dst_domain/dst_ip
leftovers — must not: the init script documents starting anyway with those rules
disabled, and turning that into a blackout would be a regression.

Also in this pass, reported by the coordinator: standing_state_test.go's
noGatewayFinding claimed to be quoted from netplane "so the test breaks if that
warning is ever reworded". It cannot — the string never leaves this package and
netplane.noGatewayWarning is never called — and the claim was already false when
it was read: netplane's text has since gained "over IPv4" and an IPv6 clause
while every test here stayed green. A fixture that advertises a guarantee it
does not provide is worse than one that advertises nothing. The comment now says
what it is, and netplanechannel_test.go pins the two couplings that are real:
the severity comes from the CHANNEL (warningFromText(t, "interface",
SeverityCritical), no classify pass), so no rewording can demote it — with the
control that the same texts on the generate channel are NOT critical — while
Section/Name DO come from the `kind "name": ` prefix, asserted in both
directions. A genuine text link is one exported helper in netplane away and is
left to whoever owns that file.

Verified: 13 seeded mutations. Twelve killed by named assertions; the
thirteenth SURVIVED — the guard in schemaWriteVerdict could not be seen, because
on a build host model.CheckConfigWritable answers nil for everything, so the
test reported success whether the guard was there or not. The checker is now
injected and both directions of that guard are killed. Every schema assertion is
walked over all three relations (disk newer / equal / older), so nothing here
passes by always answering the same way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:05:10 +03:00
omarandClaude Opus 5 da7411e29a feat(panel): attach named block/allow lists to a device, and show whether they loaded
A device could only carry hand-typed domains. It can now also reference the
`config blocklist` / `config allowlist` sections by name, which brings geosite
categories and url lists to parental control for free (Device.Blocklists /
Device.Allowlists — the Go half is landing separately).

The composition problem was the order. The engine decides a name in five steps —
allow typed, block typed, allow attached, block attached, network filter — so the
typed lane and the attached lane of ONE control are two steps apart, with the
other control's lane in between. Two controls therefore cannot show the order by
position. The card draws it instead: a five-stop rail, lit per step where this
device actually has something, and the same step number stamped on each lane
inside the two pickers.

ListPicker is a new component rather than a generalised SrcPicker: that one is
welded to useSrcOptions(), to CIDR validation, and to an empty state reading
"everyone · all LAN clients", which on a block list means the opposite of the
truth. It reuses SrcPicker.css and its whole interaction language.

Honesty, in three places it would otherwise have lied:

  - a list chip reports what /api/ruleset/status says, not that someone attached
    it. Never fetched reads "not loaded" in crit, an empty one "empty", one the
    engine has not mentioned "load unknown" — dim, never green. A name the config
    no longer has reads "no such list".
  - attaching a list is itself the switch for that device, so a list with
    Enabled=0 is NOT drawn as dead, and the DNS page's "configured but inactive"
    is replaced by a sentence naming the devices still running it. A row for such
    a list now reads "N devices only" instead of "off"/"inactive".
  - an attached allow list is terminal, so it lifts the network blocklists off
    everything it covers. Said in the picker at the moment of choosing and again
    on the card.

cleanDomain demanded /^[a-z0-9.-]+$/, so a colon could not be typed and the
engine's own full: / suffix: / keyword: vocabulary was unreachable from the
panel. parseDomainEntry accepts them from a closed positive list and refuses, by
name, the three shapes the engine silently discards: an unknown `word:` prefix, a
marker with no value, and an IP. The keyword case gets its own message — an empty
keyword is strings.Contains(host, "") and would take the device off the internet.

The logic lives in src/deviceLists.ts with tests, since `node --test` cannot load
a .tsx. Each test was mutation-checked, and the load reading is shown giving both
a positive and a negative result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:01:56 +03:00
omarandClaude Opus 5 73bd02dd71 feat(byedpi,nodetest): check the listener, not the file; and let one node be tested
A2 — the panel unlocked the byedpi egress type on LookPath("ciadpi"), which
answers "is the package installed" while the operator is asking "will traffic
sent here go anywhere". Those come apart on the SHIPPED configuration: the
packaged /etc/config/byedpi is inert (enabled='0'), and the port is coordinated
between the two packages by comment only — nothing in the daemon had ever read
that file. Result: type unlocked, egress on 127.0.0.1:1080, apply green, nobody
listening.

shater/panel/byedpi.go now decides on three separate facts (binary, enabled
instances + their ports read from the conffile, a TCP connect to each) and
reports a CLOSED state: unknown | not_installed | disabled | not_listening |
listening. Only "listening" may gate the egress type. GET /api/byedpi adds the
per-egress port reconciliation, so a mismatch is NAMED with both numbers instead
of going quiet. Nothing overclaims: the check is a connect, not a SOCKS5
handshake, and every sentence says so. A connect that is neither accepted nor
refused is "unknown", never "no".

C4 — a just-added node had no instrument: the group test reads the observatory's
board, and the observatory only probes what the rules route through, so the one
question a fresh node exists to ask ("is it alive?") answered "no rule routes
through it". POST /api/groups/test now takes {"kind":"node"} and runs the SAME
instrument — same singleton, same runner, same result type, same status poll,
same exit_ip through the target's own outbound with the same refusal to answer
from `direct`. The only addition is one fallback: a node the observatory does not
cover is measured once, here, through probeOneInto (the observatory's own
dialler), recorded under its own tag alone. A node whose base tag is a plan STORE
ALIAS — the egress-bound-group case — is NOT dialled: the board already holds its
egress-path number, and a bare-WAN measurement filed there would be the same
poisoning one layer down.

Results gained kind (group|chain|node|"") and source (observatory|on-demand|""),
so "nobody measured this" is distinguishable from "measured and dead".

Also: PUT /api/config maps model.ErrSchemaTooNew to 409 beside ErrUnmigratedConfig.
A downgrade refusal is the guard working, fixed by the operator, not by us; 500
sent the reader to the daemon log.

Every new test was verified by mutation (14 mutations, each killed by name), and
each detector has a control: the byedpi probe is shown seeing a real loopback
listener AND its absence with nothing else changed, and the node test is shown
telling a live node from a dead one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:37:48 +03:00
omarandClaude Opus 5 ffe4d8726f feat(apply,shaterd): a configuration history on disk, and a support bundle that outlives the daemon
Two holes from the operations audit, and I can confirm both of its readings.

CONFIGURATION HISTORY. Applier.lastGood and Applier.snapshot are fields in
this process. A daemon restart or a reboot loses both, and Rollback with no
snapshot goes to rollbackEngineAndPlane, which re-reads the CURRENT
/etc/config/shater — that is, it re-asserts the config that broke. With
confirm_timeout at 0 by the owner's choice there is no auto-rollback either,
so "what did the working config look like?" had no answer at all once the
daemon had restarted. Nothing on this router kept one.

Every successful apply now files RenderUCIExport(m) — the existing pure
function, not a second serializer — into /etc/shater/history/<unix>-<version>.uci.

  * DEDUPLICATED against the newest entry. shater-cron reconciles once a
    minute and every reconcile runs applyLocked to completion, change or no
    change, so a file per apply would be ~1440 identical writes a day onto
    overlay flash and would fill the ring with twenty copies of one config
    twenty minutes after the last real edit. One file is now one change.
  * 20 files / 512 KiB total / 128 KiB per entry, hard ceilings, not defaults.
    The shipped /etc/config/shater is 12.6 KB of which 459 bytes are actual
    configuration; a loaded one renders to a few KiB up to low tens of KiB, so
    twenty entries normally cost 50-200 KiB and the byte cap binds only for
    inline entry lists. Against what this product already grants itself on the
    same overlay — 4 MiB of compiled lists, an 8 MiB rule-set cache, a stats.db
    defaulting to 64 MB — 512 KiB is a rounding error. An entry over the
    per-entry ceiling is REFUSED rather than allowed to evict the whole ring,
    and the refusal is reported.
  * 0700 dir / 0600 files. Checked, not assumed: the Makefile installs
    /etc/config/shater with INSTALL_CONF, i.e. 0600 root:root, and these files
    carry the same node credentials and subscription URLs.
  * A failure NEVER fails the apply, and is never swallowed: it becomes a
    Warning folded into the set Status publishes (gather + append + finalize,
    the seam abortAfterSwap already uses), so the panel says the history has
    stopped instead of the directory quietly going stale.
  * NOT kept across sysupgrade. The audit's premise that /etc/shater is in
    keep.d is wrong — keep.d/shater-core lists four specific paths, not the
    directory. Excluding it follows model.backupBeforeChange's existing
    precedent for config.pre-v*.bak: the archive is held in RAM across the
    flash and routinely ends up in cloud storage, and this is a local undo for
    changes made on THIS box.

`shaterd diag`. The only thing this product could hand over was
GET /api/log?range=, served by the daemon — so in a crash loop the one channel
that does not need ssh dies with the process. `shaterd diag` prints version,
our packages from `apk list -I`, status, `nft list table inet shater`,
`ip rule`, the log tail and the configuration, as one block, collected entirely
by the short-lived process.

  * It works with a DEAD daemon, which is the case it exists for. The status
    section falls back to the same offline stub `shaterd status` prints and
    LEADS with the fact that no daemon answered, so an empty-looking section
    can never read as a healthy one. No section is ever silently absent: a
    missing nft/ip/apk produces "NOT COLLECTED: <reason>", and `uci export`
    failing falls back to the raw file and says so.
  * Masking is a POSITIVE, CLOSED list of the fields that may be PRINTED
    (diagSafeUCI), keyed by section type. Everything it does not name is
    masked — an unknown option, an unknown section, and every field added to
    model.Model after this build. That is the direction the open `default:`
    lesson demands: the recoverable side is "hidden", not "shown".
    TestDiagMaskingIsClosedOverTheWholeModel proves it by reflection over every
    string the model can render, with the control that the same instrument sees
    those values in the unmasked text.
  * A second layer scrubs the refused literals from the WHOLE document, because
    masking the config alone would only move the leak: the daemon prints a
    subscription URL into its own log on a fetch failure.
  * node.uri keeps its scheme and nothing else — "is this node vless or
    wireguard" is most of the diagnosis and a protocol name is not a secret.

Verified: 14 seeded mutations, every one killed by a named assertion (dedupe
removed, prune removed, ceiling removed, 0600->0644, 0700->0755, failure
swallowed, warning not folded, version not sanitized; allow-list defaulting to
ALLOWED, uri scheme dropped, scrub removed, failed sections made absent, stub
banner removed, masking removed). Every check is paired with its control — the
ring tests assert the newest entry is present and correct, so "the old one is
gone" cannot be satisfied by a ring that silently stopped writing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:34:39 +03:00
omarandClaude Opus 5 7f84ea9496 feat(panel): the logs answer why it went there, where it went out, and where the search stopped
Four fixes on one path — the one a person actually walks when a site does not
open: Insights -> DNS log -> Connections. Three of them were fields the daemon
already put on the wire and the panel dropped on the floor.

1. Connections shows the routing record. ConnLogEntry gains rule_kind/rule/chain,
   with the discipline stats.go wrote them under: "" is NOT RECORDED and can
   never be drawn as "no rule matched". Three states, three readings, and a
   CONTROL test that fails if any two of them render alike. The outbound path is
   printed rule-named-tag first, dialling-outbound last (the wire order is the
   reverse).

2. The DNS log says where the lookup left. outbound_kind is a closed four:
   detour (tag named) / default (the resolver names no detour -> the query went
   out the plain WAN, past the tunnel; marked amber) / local (cache, optimistic
   answer, filter block: nothing egressed) / "" (not recorded). An unrecognised
   value falls to "not recorded", the recoverable side, not to one of the answers.

3. Both logs take q=. The daemon filters inside the store on the same walk as the
   cursor, so limit counts MATCHING rows. The searched fields are named under the
   box, because a POSITIVE CLOSED list is also a statement about what is NOT
   searched: no ports, no rule_kind, no outbound_kind — q=default matching every
   default-egress row would be a trap wearing a filter costume. logRoute mirrors
   filter.go exactly so the ?mock backend finds and misses what hardware does.

4. A TRUNCATED page is not the end of the log. A filtered walk is budgeted
   (MaxFilterScan); a page that ended on that budget is short for a reason that
   has nothing to do with how much data exists. X-Stats-Log-Truncated is now read
   and the state is NAMED — an amber "Scan stopped" plate, the empty text saying
   "not the end of the log" instead of "nothing found", the count line refusing
   to say "all loaded", and the daemon resume cursor behind a button. The cursor
   matters twice: a truncated page can have ZERO rows, so there is no row seq to
   page from, and the live tail now advances on rows EXAMINED rather than rows
   matched — a filtered after= poll that matched nothing used to rescan the same
   window every tick forever.

Also: .fp-select gets max-width:100% + min-width:0. A <select> shrink-wraps to
its widest option and, as a flex item, refuses to shrink below it: the geo
provider label measured 501px in a 375px viewport and gave the PAGE a horizontal
scrollbar (scrollWidth 559 vs clientWidth 375, measured). Settings.css and
Networks.css each carried a narrow copy of this fix; the component is the right
place. Verified on an isolated harness with no page-local CSS: bare select
overflows a 320px row at 438px, adding the class alone brings it to 320/320.

Tests: 34 new, every one mutation-verified — unrecorded folded into default /
into local, rowMatches returning true unconditionally, outbound_kind added to the
searched fields, historyExhausted ignoring truncated, logEndNote drawing both
situations with one sentence, logCountLabel saying "all loaded" on an incomplete
scan. Each revert reproduced its own failure text. Both search directions are
covered (finds / does not find), which is what catches a filter that matches
everything. Browser-checked at 390 and 1280 in both themes, no horizontal scroll;
the six-click resume walk from "scan stopped" to "Nothing in the log matches" was
exercised live in ?mock.

NOT verified: no hardware or VM run — the truncated state was exercised against
the mock backend, whose scan budget is 60 rows where the daemon uses 20000.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:33:20 +03:00
omarandClaude Opus 5 fd8b424d5a fix(netplane): a missing IPv6 gateway is not an outage, and it never was one for IPv4
Measured on the production BPI-R3: egress `ewan` was carrying the entire
household's traffic (chain default, plane full, verdict tunnel, 15 hours up)
while the panel showed, at CRITICAL, "this egress CANNOT REACH ANYTHING outside
its own subnet — every node, group and rule bound to it will fail to connect".

table 8208 held `default via 10.0.0.1 dev eth1`; the IPv4 half was perfect. eth1
holds one address, fe80::.../64, and the ISP publishes no IPv6, so
`ip -6 route show default` is empty router-wide. The -6 pass found no nexthop
and one family-agnostic text declared the whole egress dead.

Two defects in one line. A per-family fact was stated as an absolute, and the
absence of an optional ISP feature was graded as an outage — in the loudest
register this codebase has, on a channel apply grades critical wholesale. Red
that stands for fifteen hours over a healthy router is not a warning.

IPv4 stays loud and unchanged in substance: an uplink with no IPv4 nexthop
carries nothing. It now scopes its consequence to IPv4 and says outright that
it is not describing IPv6.

IPv6 splits on one piece of evidence — does the device hold a global IPv6
address? If it does not, IPv6 is simply not provisioned on this link: nothing
is broken, nothing leaks (the v6 mark keeps its own table and its unreachable
floor, so it cannot fall through to main), and there is nothing the operator
can do because the missing thing is upstream. Silent. If it does, IPv6 is
configured and the nexthop is missing anyway — a real fault, still critical,
now scoped to IPv6. Silence requires positive evidence: a failed address read
makes us louder, never quieter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:27:50 +03:00
omarandClaude Opus 5 a9ec36e053 fix(netplane): a second LAN's inbound options meant nothing, and rule counters were structurally blind
Three things the divert plane got wrong once more than one tproxy inbound
exists, and one it got wrong all along.

Per-rule diverts read `tcp`/`udp`/`tproxy_port` off the FIRST enabled tproxy
inbound and applied them to every device the plan touches. With one LAN — the
shipped shape — first and owner are the same section and nothing showed. With
two, `option udp '0'` on the second inbound was ignored (UDP diverted anyway,
into another section's listener), `option udp '1'` was ignored the other way
(no per-rule UDP line at all, so the rule's counter never ticked for UDP and
Insights showed a rule that appeared never to match), and `option tproxy_port`
pointed at the wrong listener. Same for the dns_intercept :53 lines, which sit
above the fib-local bypass. Each ingress device now resolves to the inbound
that OWNS it; a device no inbound claims still falls back to the primary,
because that is the only listener its traffic can reach. Verified
byte-identical output for every single-inbound shape against the pre-change
renderer.

Rule counters: a counter exists only for a rule the plane emitted a divert
line for, and it only emits them from SOURCE selectors — so a rule written by
domain or ruleset never appears in RuleTraffic at all, and absence there could
not be told apart from "carried nothing". It cannot be measured: which rule a
packet matches is decided inside the engine after the divert, where nftables
cannot see it. So no counter is invented. Instead the plane says which rules it
can measure (RuleMeasures) and what the numbers it does have actually mean —
an upper bound, not the rule's traffic — and the two are pinned to the rendered
ruleset in both directions. Counters are now declared BY the emitting line, so
a rule whose fragments were all dropped no longer leaves a counter attached to
nothing, reading a confident, permanent, false 0 B.

untunnelable_egress could resolve, pass validation and still mark nothing when
the plan has no LAN ingress device — while apply's note, gated on the same
binding succeeding, told the operator that IPsec/GRE/SCTP now leave through it.
The gate is right (the marking rule has no safe unscoped form), the silence was
not; bound-but-inert is now named.

stats/panel do NOT consult RuleMeasures yet — wiring it is a change outside
this package, and the gap is still visible to an operator today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:27:33 +03:00
omarandClaude Opus 5 d40eeede0c fix(model): refuse a config newer than the build, and never write the version back down
A downgrade ate the configuration silently, and permanently. Migrate() refuses a
newer schema, but nothing on the write paths calls it: the daemon starts
regardless, ParseUCIExport reads the options it knows and drops the rest, and
WriteUCI replaces the WHOLE package. So an older build rewrote /etc/config/shater
with only what it understood. The second half is what made it unrecoverable —
schema_version round-tripped through the Model, so the rewritten file claimed the
OLDER version, and a newer build put back afterwards saw cur == CurrentSchemaVersion
and migrated nothing. Nobody had to be at the keyboard for any of it: the profile
watcher looks every 25 s and `sub update` runs from cron every 6 h, and both
persist through WriteUCI.

The mechanism for refusing already existed and already worked in the other
direction (ErrUnmigratedConfig + CheckConfigWritable); this is its second caller,
not new machinery.

- guardSchemaDowngrade refuses the write and the panel's pre-flight when the
  config on disk is newer than this build, naming both versions and the way back
  (put the newer package on again — the config is untouched). ErrSchemaTooNew so
  a caller can answer 409 instead of 500.
- withDiskSchema takes schema_version from the DISK, never from the caller. A PUT
  body that omits it sends 0, and a rendered 0 is an OMITTED option: the version
  would have vanished and the next `shaterd migrate` would replay every step. A
  body claiming 99 would have locked the box out of its own panel.
- backupBeforeChange copies the live config to /etc/shater/config.pre-v<schema>.bak
  before the first migration and before the first write — once per schema version,
  write-then-rename. A failed backup aborts: the `uci commit` that follows writes
  to the same filesystem, so refusing costs nothing that was not already lost, and
  best-effort-and-carry-on is the silent skip we keep paying for.
- The reverse direction is fenced by a test: an unmigrated v1 config still refuses
  a rule-changing write as ErrUnmigratedConfig, still allows one that leaves the
  rules alone, and still keeps its `list dst_domain` and its v1 stamp.

The shipped /etc/config/shater now says what "conffile" actually buys — values
across a package upgrade, not comments across the first write, which happens
without an operator — and the annotated file is installed a second time as
/usr/share/shater/config.sample, where nothing rewrites it.

INSTALL.md gains the downgrade procedure. Measured on the testbed VM (ImmortalWrt
25.12.1 r37978, apk-tools 3.0.5) against the real apk-v0.2.9/v0.2.10 feeds in an
isolated --root sandbox: `apk upgrade <named>` does not downgrade at all;
`apk add <pkg>=<ver>` does, and leaves a pin in world that a later upgrade obeys;
`apk upgrade -a` downgrades too but took four unrelated packages with it.

Tests in shater/model/schemadowngrade_test.go; every assertion checked by mutation
(8 mutations, each killed a named test) and every refusal paired with a control
that accepts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:21:01 +03:00
omarandClaude Opus 5 7c93019e81 fix(panel): draw the settings that decide whether traffic leaks
Five things the panel knew and did not say, each one a state where the
screen read healthier than the router was.

Rule.Kill was typed, round-tripped and drawn nowhere. `open` sends a
rule's traffic out direct — around the kill-switch, with the real
address — when its target cannot be built, and such a rule looked
exactly like one that fails closed. It now has an editor beside Target
and an amber mark on the row; the fail-closed default draws nothing, so
the two states are not priced alike. An unreadable value is its own
state: it blocks, like the daemon, and the picker re-surfaces it
verbatim rather than rewriting a value it never showed.

Alert channels were write-once for Type/Token/ChatID/URL/Events, so
fixing a typo meant deleting the channel and going back to BotFather for
a token you already owned. Add and edit are now one form. The token box
starts empty and the caption says what empty means — keep, never clear —
because the panel refuses to show the secret and a save may only clear a
field the editor could show. Same rule covers a type switch: the other
kind's settings stay stored and unused.

The add-rule form pre-filled Target=direct. An untouched form is a rule
with no matchers, i.e. the default route, so one press put the whole LAN
on the plain WAN. `block` would only have swapped the leak for an
outage; the recoverable default here is no default, so the form refuses
and asks.

The empty state said "all traffic follows the default route" without
naming it. On a fresh install that route is `block` — the LAN has no
internet — and this is the page the kill-switch alarm sends people to.
Both it and the lead now name the route in force.

The interception board was computed from the config alone and lit `lan`
green over a stopped engine. Green now needs the engine up AND the full
plane; a hold plane blocks rather than carries, and unknown is an unlit
socket.

Also: four rungs of the untunnelable copy claimed traceroute works. It
prints `* * *` and no hops on every setting — the wording is now
apply/warnings.go's own udpTracerouteFacts, said once.

Tests: killPolicy / alertEdit / defaultRoute / intercept, 38 cases, each
mutation-checked (16 mutants, all caught). Browser-verified at 390 and
1280, no horizontal overflow.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:10:32 +03:00
omarandClaude Opus 5 bc7ea359c8 feat(panel): the settings that only /etc/config/shater could reach
Five groups of real daemon settings had no control in the panel, so the
only way to change them was to edit the config over SSH. Each one is now
editable where it belongs, and each editor is built so it cannot lose a
field it declines to display.

DNS list refresh intervals. Blocklist.UpdateInterval was hard-coded to
"24h" in two places and shown nowhere, while the row displayed the
interval the ENGINE reported — a readout dressed as a knob.
Allowlist.UpdateInterval did not exist in the panel at all. It matters
because an allowlist is how a blocklist false positive gets corrected: one
pinned to a day delivers the fix up to a day after the site broke.

Geo data. GeoProvider and the four URL fields are consumed for real
(generate.SetGeoProvider, /api/ruleset/categories) and the panel USES the
data they pick, while offering no way to choose it. New Settings group
with the closed five-provider list, the custom {category} templates and
the two category indexes. Only `custom` reads the templates, so only
`custom` renders them; an unknown provider is preserved and marked rather
than silently rewritten to auto on page load.

Subscription filters. Include/Exclude/FilterProto/FilterCountry/Dedup,
Format, ExpireAlertDays and the three device-identity headers are now
editable — the same five filters a group already offered over its members,
applied one step earlier. 376 nodes can become the four Dutch ones without
SSH. ExpireAlertDays keeps its three states (blank = the 3-day default,
"off" = -1) instead of being flattened.

Edit-after-create. Blocklists, allowlists and resolvers could be
configured only at creation; a typo in a URL meant delete and rebuild, and
deleting a resolver clears whichever global slot it filled. Every one now
has a row editor. `file` and `geosite` sources are offered when a list
already IS one, so opening a list the panel cannot create never becomes a
way to destroy it.

Stale local type copies. DNS.tsx and Settings.tsx carried local
Blocklist/Allowlist/Globals extensions whose comments claimed api.ts did
not type those fields; api.ts had typed them for a long time. Deleted —
the note was an invitation to declare the next field twice. (Egress.Target
was already gone.)

Along the way, three defects the work surfaced:

  * parseDomains cut comments per TOKEN, so pasting "# ads and trackers"
    contributed ads, and, trackers as three real blocked domains. Cut per
    line now.
  * FetchVia=proxy with no FetchDetour resolves to the tag `direct`
    (engine.ViaToTag), so the feed is pulled over the plain WAN and the
    provider logs the router real address — the one thing `proxy` is
    chosen to hide. The row said "via proxy" for it. It now says
    "proxy - no route" and the editor carries an amber explanation. The
    picker also gained chains, which apply.resolveVia supports for real
    and the picker excluded with a comment that misdescribed the contract.
  * Adding a subscription only saved it. apply does not fetch, and
    shater-cron is inert unless globals.enabled=1 AND the service is live,
    so on a router not yet switched on nothing would ever fill it — and
    the only Update button sat at the bottom of a collapsed panel. Adding
    now fetches, reported separately from the save, and every row carries
    Fetch now. A row with no nodes says what to press.

Nodes also gained the forward link nothing had: a node is not something a
routing rule can point at, and no page said so.

The merges live in subEdit.ts / dnsListEdit.ts / geoProvider.ts because
`node --test` cannot mount JSX. The rebuilt shapes return Complete<T>, so
a field added to api.ts fails the build in the function that has to decide
about it; the subscription merge extends instead, because five of its
fields are provider-reported state no control can show.

Verified: npm run build green; 173 tests pass; 18 mutations each killed a
named test and a probe field added to Allowlist broke the build inside
nextAllowlist; zero Cyrillic in panel/src; Chromium at 390 and 1280 with
no horizontal overflow (the detector caught a real 559px select spill at
390 before the fix).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:07:00 +03:00
omarandClaude Opus 5 447cd8cf7d fix(panel): a switched-off service is not a fault, and say so before the apply that is
The shipped config is `enabled '0'` + `kill_switch 'closed'` with nothing
applied. Every lamp in the panel was derived from what is INSTALLED and none
from whether anything was MEANT to be, so a package that installed exactly as
designed showed a crit master lamp ("Engine down"), a crit kill-switch module
("NOT IN EFFECT") and three crit pips on Apply — at a person who had not done
anything yet. Red that fires on a correct installation is red nobody reads by
the time something is actually wrong.

serviceIntent() is the missing question, and every readout that used to answer
from the installed state now asks it first: off ⇒ unlit socket and a word that
says why; on ⇒ every alarm exactly as before. A positive `off` only — an
unreadable configuration stays `unknown` and keeps its crit, because that is
the state where the LAN really is cut off.

applyRisk() is the other half. Applying an empty config with the service on
and the kill-switch closed sets route.final = block, and the tproxy divert for
the shipped `lan` inbound is installed — so every TCP connection and UDP flow
from the LAN is handed to the engine and dropped. The panel read that state
perfectly once it existed and said nothing before, with confirm_timeout at 0,
so the most dangerous apply this router does ran with no auto-rollback. The
band names the outcome, the missing rollback and the fix, and does not block
the apply.

Insights had a short-circuit for this exact job that never fired: it was gated
on logging being off, and the shipped backend is memory. Ten sections drew ten
well-mannered "nothing yet" states and not one named the switch.

Every test is mutation-checked, and each one is paired with the control that
proves the instrument can still produce the alarm.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 10:04:59 +03:00
omarandClaude Opus 5 65db309e3c fix(sub): fetch_via=proxy went out on the plain WAN from cron and at boot
`fetch_via=proxy` on a subscription means "pull this feed through the tunnel",
and it is set for exactly one reason: the provider is blocked, or the owner does
not want the provider (and every hop to it) learning the router's real address.

The panel honoured it — Applier.UpdateSubscription resolves fetch_detour against
the running engine — so testing it once from the browser showed it working. The
CLI verb did not: it logged one daemon.warn line and fetched DIRECT.
/etc/init.d/shater-cron calls exactly that verb, so every scheduled refresh and
the fetch-at-boot went out unproxied, and the only trace was a syslog line in a
log globals.log_syslog='0' switches off.

The CLI cannot do this fetch itself — only one process may own the engine — so
it now DELEGATES: a new control-socket verb `sub update <name>` runs the very
same Applier.UpdateSubscription the panel's Refresh button calls. One
implementation, so the two paths cannot drift again.

With no daemon to ask, the subscription FAILS (exit 1) instead of falling back.
The refusal is recoverable — shater-cron does not stamp the item, so it retries
after RETRY_SECS and the already-cached nodes keep working — where a silent
direct fetch is not: the disclosure has already happened. Direct subscriptions
are untouched and still need no daemon at all.

Order is load-bearing: the direct pass and its UCI write run first, then the
delegated ones, because the daemon re-reads UCI and writes back userinfo
counters a later write from this process would silently drop.

Also in this file, reported by the LuCI agent: the offline stub of `shaterd
status` published config_readable=false after a SUCCESSFUL read, telling every
consumer to disbelieve three values it had just read correctly (LuCI worked
around it by reading the field only when a daemon answered), and swallowed a
FAILED read with no trace — the inverted lie apply.Status() was fixed for, in
the one situation that matters most: a full /overlay where "not enabled" tells
the owner they switched it off themselves while the fail-closed plane holds the
LAN shut. Both halves now mirror the live path exactly.

Tests are mutation-verified in both directions, with an instrument that gives a
positive reading for BOTH "went through the tunnel" and "went direct" — a live
origin server and a live control socket in every case, so neither zero is an
artifact of the other endpoint being absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 09:53:19 +03:00
omarandClaude Opus 5 d63f1d896d fix(bridge): reassemble return-path IP fragments — the bridge dropped them
sing-tun's classifyReturn answers `returnPass` for any IP fragment, so the l3
return path never judges one. On the WireGuard endpoint a passed packet still
reaches the endpoint's own tun stack; the bridge has no second consumer — both
deliverReturn and the batch read loops offer a packet to each attached return
path and then drop whatever nobody claimed. A fragmented answer coming back
through a bridge outbound was therefore lost outright, 100% of the time.

Fragments do arrive: the return direction is fragmented by the LOCAL kernel
(conntrack defragments at PREROUTING for the NAT lookup, the output path
re-fragments to the bridge TUN's 1500-byte MTU honouring IPCB frag_max_size).
packet.go's fixReturnChecksum already recognises a fragment and declines to
touch it — the path was known to carry them.

frag_reassembly.go is a deliberate sibling of transport/wireguard/
frag_reassembly.go: same algorithm, same ceilings (64 datagrams, 1 MiB, 5 s,
non-refreshed deadline, partial overlap poisons the key), so collapsing the two
into one shared package later is mechanical. They are not shared today only
because the seam that would host the shared type — transport/wireguard/port.go
and its test suite — is owned by other work in flight.

Windows is deliberately untouched: there a fragment never reaches deliver() at
all, because classifyInbound needs a transport header to decide ours/not-ours
and WinDivert reinjects the rest into the host stack. Different function,
different defect, platform we do not ship.

protocol/tailscale gets a comment, not a fix: the one ReturnPackets call that
package makes carries BuildUnreachable replies, which are synthesised whole and
can never be fragments, and the real tunnel return path is upstream
tstun.Wrapper.Write, ahead of every seam this tree owns.

Verified: 23 tests, all 14 seeded mutations killed (including "seam removed" on
both the portable and the Linux batch loop), -race clean on linux/amd64 in
docker and on windows/amd64. The darwin seam is compile- and vet-checked only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 09:53:03 +03:00
omarandClaude Opus 5 55dea4e729 feat(stats): the DNS log says where it went out, and both logs can be searched
C1. handleEvent had QueryEvent.Outbound in its hand, used it only to compute
action(), and dropped it. The query log could name the resolver that answered
and not the channel that resolver's own packets took — the one fact an
anti-leak `detour` on a resolver exists to control.

LogEntry now carries Outbound + OutboundKind, on the ConnLogEntry.RuleKind
discipline: "" is reserved for NOT RECORDED, so the three states that all have
an empty tag stay distinct — "detour" (tag recorded), "default" (the resolver
names none, so its packets take the plain WAN), "local" (cache/optimistic/
filter block: nothing egressed at all). An unrecognised Source falls to
unrecorded, the recoverable side. Rows from older builds decode to unrecorded
and are therefore still distinguishable from a recorded no-detour row.

No rule name is attached, and that is deliberate: the DNS path has strictly
less to work with than the connection path did. A DNS *rule* picks a SERVER,
not an outbound, and the event carries no rule identity at all — only the
transport's tag. Inventing one would be a forgery.

Cost, measured: LogEntry 120 -> 152 B (+32 B/row, two string headers on
aarch64). +6.4 KB at the default ring of 200, +160 KB at 5000. Tag bodies go
through the existing intern table (maxRuleKeys=512, shared with the rule text).

C2. /api/stats/log and /api/stats/conns take q=<substring>, applied INSIDE the
store on the same walk as the seq cursor. It has to be there: Limit is applied
by the store, so post-filtering a returned page would hand back 3 rows of a
50-row page and call it a page. Substring, not regex — nothing a client can
type costs more than a linear scan.

Pagination stays honest. A filtered walk must examine rows it will not return,
so it is bounded (MaxFilterScan=20000) — and a page that stopped on that bound
is short for a reason that has nothing to do with how much data exists. That is
reported: LogPage.Truncated + ScanCursor, surfaced as X-Stats-Log-Truncated and
X-Stats-Log-Cursor. Unfiltered requests are untouched: no budget, never
truncated, same walk as before.

The logRing seam now returns LogPage/ConnPage instead of (rows, pending) so the
truncation state cannot be dropped on the floor between the ring and the API.

Tests: all mutation-verified (7 reverts, each reproduced with its message),
including the copying-variant control for the intern table — strings.Clone
passes an equality check and fails the identity check the test actually makes.
Filter coverage is both-directions (finds / does not find) on both backends,
with mem-vs-bolt parity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 09:51:17 +03:00
omarandClaude Opus 5 2dfd7caf2b fix(apply): the untunnelable note may not describe an egress the plane never bound
untunnelablePolicyWarnings opened its egress branch on
`Globals.UntunnelableEgress != ""` alone. netplane refuses far more than a
typo: UntunnelableEgressBinding fails CLOSED for any name that does not
resolve to an interface/tunnel egress WITH a device — nothing is marked in
prerouting, no forward-chain accept is rendered, addEgressRouting installs
no rule and no table, and the `untunnelable` policy decides everything by
itself. The note nevertheless opened with "...now leave through egress
"x": the kernel routes them out that interface", about a carrier that does
not exist; it even printed `(device )` once the device was interpolated.
The tail hedged the case thirty lines later, and the first sentence is what
gets read.

The branch is now gated on netplane's OWN verdict, called rather than
re-derived (apply imports netplane, so unlike model.ValidateUntunnelableEgress
there is no copy to keep in lockstep). That needs the whole model, so
collectWarnings/gatherWarnings/untunnelablePolicyWarnings take *model.Model
instead of model.Globals.

When the option is set and unbound, the note now LEADS with that fact and
then gives the ordinary policy text, because that is exactly what the router
is doing. The bound branch drops "a name that matches no interface/tunnel
egress" from its failure list — that case can no longer arrive there — and
names the device it resolved to.

Two further claims found while checking the rest of the file against the code:

- the `icmp` rung promised "IPTV and VPN passthrough work only toward
  addresses your rules route directly". Multicast crosses this router under
  NO setting (the stream is WAN-side inbound; a client's outbound multicast
  is UDP, which untunnelableFilter structurally cannot match), and the other
  three rungs all say so. One true clause was carrying one false one — the
  same sentence the `block` note records having removed for being false.
- "\"icmp\" drops it, excepting only ping/echo" understated a leak. `icmp` is
  the one rung that walks the destination plan and it ACCEPTS raw ESP/AH/GRE
  toward provably-direct destinations, so the operator was told it was
  contained while it left with the router's real address.

Ratcheted by untunnelable_egress_honesty_test.go, each assertion with a
control: the bound and unbound halves are walked in one pass, and the IPTV
and `icmp` checks fail if the matrix ever stops producing the notes they
read. The traceroute matrix grew a third egress value (set-and-bound,
set-and-unbound) so the bound branch keeps being walked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 09:45:24 +03:00
omarandClaude Opus 5 6f84d0ca7b fix(luci): read daemon_answered, and stop reading a config nobody could read
The dashboard told a live daemon from a dead one by `plane: ""` — a side
effect of the offline stub being a zero value, not a promise anyone made.
`shaterd status` now states it: daemon_answered, true on the live branch and
false on the stub. The detector reads the field first and keeps the plane test
only as the fallback for the non-atomic update window (new luci-app-shater,
old shaterd). When the two disagree the field wins; a stub carrying a plane
word must still read as "no daemon answered".

Both lists are positive and closed. A daemon_answered that is not exactly
true/false is not a verdict and falls through; a plane word this build does
not know lands in unknown. Nothing lights green or amber on a guess, and the
launcher button is still never disabled.

config_readable was already on the wire and nothing here read it. With it
false, enabled/kill_switch/panel_port are zero values: "inert (disabled)" and
a green "closed (fail-closed)" were being rendered out of placeholders, in the
one situation — a full /overlay, an interrupted commit — where the fail-closed
plane has the LAN cut off and the owner is told they did it to themselves.
Those rows now say "not known", a Configuration row carries the daemon's own
reason and its don't-switch-anything-off warning, and an absent nft table is
no longer softened to amber by an `enabled` nobody could read.

The field is consulted ONLY when a daemon answered: the offline stub reads UCI
directly and never sets ConfigReadable, so its false is a zero value while its
enabled/panel_port ARE real reads. Taking it at face value would put "could
not be read" on screen for a readable file. Same reasoning drops plane,
traffic and hash on the stub branch — the contract calls them placeholders.

panel_port is CONFIGURED, not bound: shaterd logs a panel bind failure and
carries on, and SHATER_PANEL_ADDR can switch the server off while the port is
still reported. Nothing measures a listener, so the hint, the tooltip and the
new Panel port row say the port is configured rather than checked, and its
lamp stays unlit even on a healthy router.

tests/status-readout.test.js grows the new cases and now runs under gate step
[7/7]. Mutation-checked four ways against copies: dropping the
daemon_answered branches fails 7 assertions by name, dropping the plane
fallback 5, reading config_readable without the daemon gate 5, and rendering
the placeholders as readings 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 09:40:24 +03:00
omarandClaude Opus 5 e38108a7c4 test(gate): install iproute2 in the docker lane — without ip every slot is free
test / go + panel tests (push) Successful in 15m18s
release / test gate (push) Successful in 10m57s
release / apk aarch64_cortex-a53 (push) Successful in 5m52s
release / apk x86_64 (push) Failing after 28s
release / release apk (push) Successful in 6s
This change was already in the working tree when this session started; it is
committed here because it is load-bearing and an uncommitted load-bearing file
is a trap.

netplane.L3SlotFor asks the kernel through `ip link show` and reclaims through
`ip link del`. golang:1.26 ships no iproute2, so in the docker re-exec lane
every slot read as FREE, TestIntegrationL3StaleSlotIsReclaimed stood itself
down rather than pass while proving the opposite of what it claims, and [5/7]
then failed the gate — correctly, since this environment HAS root and
/dev/net/tun and the capability guard is therefore not what skipped it.

Installing it is also what made the concurrent-namespace defect visible at all
(see 06c04c157): with no `ip` on PATH, no `ip link del` was ever issued and the
two test binaries that were destroying shater/generate's TUN looked innocent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 03:25:58 +03:00
omarandClaude Opus 5 06c04c157d fix(gate): a unit test in one package was deleting another package's TUN
`go test` runs package binaries CONCURRENTLY and every one of them shares the
host's network namespace. netplane.L3SlotFor is destructive by design — it
DELETES a candidate slot it finds occupied rather than waiting for it — and
netplane.removeL3Devices deletes both slots unconditionally. Two test binaries
reached those for real:

  shater/engine  l3slot_test.go calls l3RetargetForNext for its return value
  shater/apply   Applier.Teardown -> netplane.TeardownRouting -> removeL3Devices

Measured with an `ip` shim on PATH inside the gate container: apply.test issued
9 `ip link del shater-l3a` + 9 `ip link del shater-l3b` per run, engine.test one
per l3slot test — into the namespace where shater/generate's privileged tests
were holding a live TUN. From the other side that is

  post-start inbound/tun[l3-in]: starting TUN interface: find tun interface: Link not found
  no [shater-l3a shater-l3b] device exists after a successful Start

i.e. an intermittently red [2/7]/[4/7] in a package that did nothing wrong,
while [5/7] — which runs only `^TestIntegration`, so neither binary reaches the
slot code — passed the very same test seconds later. It only became visible when
iproute2 was installed into the gate container: without `ip` every slot read as
free and no deletion was ever issued.

Not a product defect. shaterd is one process with one engine; the running
generation's slot is excluded before anything is deleted, and nothing else on
the router calls L3SlotFor.

The kernel is faked rather than the CHOICE: making the engine's tests stub the
slot answer would delete the only place the ENGINE checks that the running
generation's slot is excluded, which is the invariant the production outage
violated. netplane.L3StubKernelForTest points the two kernel operations at an
in-memory set; engine and apply install it from TestMain (forget-proof, unlike a
per-test helper whose omission fails in a different package on some runs only).
netplane's TestL3StubKernelTakesTheSlotChoiceOffTheKernel is the control, in
both directions: stubbed, nothing reaches the exec seam; restored, the same call
does.

Mutation: with the engine TestMain reverted, the generate binary's
TestIntegrationL3* failed 8 of 8 runs beside a loop of the engine binary; with
it, 0 of 8. With L3StubKernelForTest degraded to a no-op, the control fails
naming the three escaped `ip` calls.

Also: the DoH3 ownership test's control now retries.
requireInstrumentFindsPackedQuery packed a query into a pooled buffer, released
it and demanded the scan find it — but under -race sync.Pool.Put drops one
object in four on purpose, so the control failed 18 of 60 measured runs and took
the whole -race pass down with it. Its sibling control in the same file already
retried for exactly this reason. The claim is existential ("this instrument CAN
find a released buffer"), so one success out of 32 proves it and nothing is
diluted; 0 of 60 after. What it does not buy is stated in the code: the VERDICT
is still a 3-in-4 detector under -race, which is the safe direction, and the
non-race pass runs the same test as a certainty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 03:25:46 +03:00
omarandClaude Opus 5 3654acf7fb fix(egress): tunnel was a device to the router and an unknown type to the engine
An egress type was read by two halves that never call each other. netplane's
EgressDevice accepted `tunnel`, so addEgressRouting gave it a mark, an `ip rule`,
a routing table with an unreachable floor and a prerouting mark bypass, and
`untunnelable_egress` (D26) carried ESP/AH/GRE/IGMP/SCTP out of it by kernel
routing with the engine nowhere in the path. generate's outbound switch had never
heard of `tunnel`: default arm, no outbound, so every node, group and rule bound
to the same egress was fail-closed. One name, two answers.

Refusing `tunnel` would have broken the half that works to match the half that
does not — D26's kernel egress is shipped and verified, and the generator's
refusal is already loud and fail-closed. `tunnel` is not a distinct kind either:
the data plane treats it identically to `interface` in every line that mentions
it, and the panel's own `interface` label already reads "out a specific WAN or
tunnel". So it is an ALIAS, and it is folded to `interface` ONCE, at the config
boundary (Model.NormalizeEgressTypes, called by ParseUCIExport/ReadUCI). Teaching
the generator a second string would have left two strings for the next consumer
to forget; after the fold there is one.

- model: CanonicalEgressType / EgressTypeKnown / KnownEgressTypes — a closed,
  positive registry, plus NormalizeEgressTypes on the load path. An unrecognised
  type is left as written, never defaulted: substituting `direct` for a typo
  would send traffic somewhere nobody asked for.
- model: ValidateEgresses now NAMES an unknown type at validate time. Until now
  the only notice was a generator warning raised while building an engine config,
  which said nothing about the data plane — and the two disagreed anyway.
- netplane: EgressDevice and the prerouting mgmt-bypass consult the registry
  instead of carrying their own copies of the rule. The bypass now keys off
  EgressDevice, so a device-kind egress with no interface no longer gets an
  accept for a mark addEgressRouting never installs.
- panel: the egress editor cleared Interface/Port/DPI for every type it had no
  branch for — including types it renders no field for — so opening an egress it
  labels "(unknown)", changing only the NAME and saving deleted its `interface`.
  On a `tunnel` egress that silently unbound untunnelable_egress and dropped the
  ESP/GRE carrier back to policy. A save may now only clear a field the editor
  was in a position to show.
- panel: the unknown-type hint said "This engine builds no outbound for that
  type", which was false for the one unknown type anybody had — the data plane
  was building it a routing table at that moment. It now names both halves and
  states what saving does.

Tests: TestEgressTypeMeansTheSameInBothHalves runs one table of written types
through the real boundary and then asks netplane AND generate, requiring one
verdict (external test package: generate imports netplane, so nothing inside
netplane can import generate). Mutation-checked both ways — dropping the fold
fails on `tunnel`; restoring the old EgressDevice string test reproduces the
historical split with "generate emitted outbound egress-probe = false ... want
true". Panel: egressEdit.test.ts, mutation-checked by restoring the
unconditional clear (Interface undefined, want 'wg0').

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:52:56 +03:00
omarandClaude Opus 5 3a9b3f523d fix(l3): a test that is not about the TUN must not open one
The l3_tunnel default flip (164b703a7) turned 32 ORDINARY tests in
shater/generate red — the whole CI — because every fixture with a tproxy inbound
now generates the `l3-in` TUN and engine.Apply then wants /dev/net/tun, which the
act_runner LXC guest does not have. Three PRIVILEGED tests failed too, on a host
that DOES have the device.

The proposed fix was to move the TUN inbound out of generate and have the engine
add it at apply time. Refuted, on three grounds:

- it does not fix the 32. Thirty of them fail inside engine.Apply, not box.New;
  the engine adding the inbound leaves them exactly as red, unless the l3_tunnel
  signal travels OUTSIDE option.Options — and then
- the hash gate stops seeing it. Apply's fast path is a hash of the options; a
  decision that is not in them makes toggling l3_tunnel a no-op reconcile, i.e.
  the device stays up with the option off, or never comes up with it on;
- and the `icmp "tunnel"` warning cannot move. It needs the model, and the panel
  reads it out of GenerateWithWarnings. Leaving it in a package that no longer
  makes the decision it explains is a lie generator by construction.

What the failures actually were was contention. Measured under `docker run
--cap-add NET_ADMIN --device /dev/net/tun`: run alone, all three privileged tests
PASS; run as a package, all three FAIL — and one fails by finding a `shater-l3`
device that a DNS-filter test created. There are two L3 slots and they are global
to the process. So the fix is that the engine instrument in this suite does not
open a kernel device it does not own: withoutL3Ingress, one helper, applied at
applyAndClose and at the six other call sites.

Nothing is skipped, and the ingress does not lose coverage — it gains some:

- TestL3TunnelChangesNothingButTheTunInbound (ordinary, portable) proves the
  default config MINUS the l3-in inbound is byte-identical, through the engine's
  own marshaller, to the l3_tunnel=0 config. That is what lets the 32 Starts keep
  speaking for the default config instead of merely for a config near it;
- TestL3TunInboundIsAcceptedByBoxNew (ordinary) puts the registry half of the
  privileged test on a gate that can actually run it: a slim registry that loses
  tun.RegisterInbound now fails on EVERY CI run with `type not found: tun`
  instead of only where /dev/net/tun exists. That regression changes no generated
  byte and costs a LAN-wide outage on the router;
- TestIntegrationL3StaleSlotIsReclaimed (privileged) covers what a RESTART finds:
  an engine with l3Device == "" next to a device it did not open. It must take
  the other slot, leave that one alone, and RECLAIM it on the next apply. The
  occupied slot is held by a second live engine, not planted with `ip tuntap
  add` — a planted device is PERSISTENT and therefore attachable, and the
  planted version of this test passed with netplane.L3SlotFor's reclaim loop
  deleted, i.e. proved nothing.

generate's placeholder device name is now longer than IFNAMSIZ allows. box.New
accepts it (measured), so the emitted config is still one the engine can
validate; Start refuses it and creates NO device. A caller that builds a box from
generate's output without going through engine.Apply therefore fails at once and
visibly, instead of quietly creating `shater-l3` — the one name every generation
wants, and the intermittent TUNSETIFF EBUSY that netplane/l3.go exists to refuse.

The "leaked TUN" in the sentinel's message was not a leak. Instrumented: Close
returns in ~300 µs with ZERO open /dev/net/tun fds (control: 1 fd immediately
before Close), and the device survives 3.8-4.6 s longer purely as the kernel's
deferred unregister_netdevice. On the stand (ImmortalWrt 25.12.1 r37978, kernel
6.12.94 — the router's revision) the same test takes 0.10 s, so the lag is a
nested-netns container artefact. l3GoneTimeout goes 5s -> 20s: a leak is
unbounded, so the longer budget costs one slow failure and gives up no
sensitivity.

Verification. CONTROL, the criterion that matters: without /dev/net/tun
`ok shater/generate` (was 32 failures). With `--device /dev/net/tun --cap-add
NET_ADMIN`: green, privileged tests really ran. On local_openwrt, cross-built
with the shipped tags: the WHOLE package green with every privileged test
executed, no contamination. `go build ./...`, `go vet ./shater/...` clean.

Mutation-verified, each reverted after: shortening the placeholder fails
TestL3PlaceholderCannotBecomeAKernelDevice by name; making withoutL3Ingress a
no-op brings back exactly 32 failures; gating a second config change on
l3_tunnel, and stripping nothing in the comparison, each fail
TestL3TunnelChangesNothingButTheTunInbound; removing tun.RegisterInbound fails
TestL3TunInboundIsAcceptedByBoxNew with the right hint; deleting L3SlotFor's
reclaim loop fails TestIntegrationL3StaleSlotIsReclaimed with the production
error verbatim (`TUNSETIFF: device or resource busy`); l3GoneTimeout at 1ms still
fires the leak sentinel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:49:40 +03:00
omarandClaude Opus 5 201aa7c168 fix(panel): traceroute never printed a hop — stop saying it works
Five untunnelable notes told the operator that a plain `traceroute` works,
"still follows your rules", or that the hops it prints are the tunnel's path.
Measured on the production router: it prints `* * *` and nothing else, under
every rung of the ladder — `direct` included — with the L3 ingress on or off.

There is no mechanism that could print a hop. The UDP probe is diverted by
tproxy and delivered LOCALLY to the engine's socket; local delivery is not
forwarding, so the TTL is never decremented and no router on the path is
provoked into a time-exceeded. The engine opens its own connection with a
fresh TTL, and an ICMP error raised against that has no way back to the
client's datagram. `traceroute -I` and Windows `tracert` are ICMP echo and do
work — that half of the text was true and is kept.

One shared udpTracerouteFacts now carries the symptom, the cause and the way
out, so the panel cannot fork the claim; netplane/untunnelable.go states the
same fact in the same terms.

Second correction in the same notes: the outbounds that carry an echo are not
just WireGuard/AmneziaWG. generate/route.go's l3Target is exhaustive by
adapter registration — a wireguard/AWG node AND the direct outbound behind
`direct` or an interface egress. In the commonest configuration here that is
most of the address space, and those pings answer out of the ordinary uplink
with its real address. The old text let an operator conclude either
"tunnelled" or "dropped"; it was neither.

traceroute_honesty_test.go is the ratchet: an exhaustive matrix over policy x
kill switch x L3 x egress, asserting the retired sentences never return and
that any note mentioning a trace carries the shared facts verbatim — with a
control that fails if the matrix stopped mentioning tracing at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:20:41 +03:00
omarandClaude Opus 5 78bb6a1be8 fix(netplane): a read that fails, a floor nobody checked, a flow that predates the plane
Four defects, all of the same family: something the plane relies on stops being
true and nothing says so.

1. One failed `uci -q export firewall` opened a hole AND switched off the alarm
   for it. nftZoneDevices answered nil on a read failure — the same answer as an
   empty zone — so a rule with `src: zone:lan` produced no divert line, no
   fail-closed drop and no accept_local; and uncoveredNetworkWarnings, whose job
   is to report exactly that, ran the same command, got the same nil and stayed
   silent. The read now carries its error: renderNft refuses under a closed
   kill-switch (same contract as an unusable device name) and warns under an
   open one, and the coverage check names the blindness itself.

2. RoutingPresent did not check the fail-closed floor its Apply twin installs.
   addEgressRouting/addL3Routing install three things per binding; the presence
   checks knew two. A floor that failed to install once was never retried, and
   the table fell through to `main` the first time its device went down. The
   checklist test grows clause (e) so the next mark cannot repeat it.

3. A flow established before the divert plane existed bypassed it for life:
   confirmed by conntrack while nothing diverted it, offloaded to fw4's
   flowtable, steered by netdev-ingress ahead of our prerouting hook and
   refreshed by its own packets. On the divert going from ABSENT to PRESENT —
   not on every apply — the TCP/UDP entries of flows forwarded from the divert
   devices' subnets are dropped, so they re-derive their path. Not a flush: the
   router's own addresses and LAN-to-LAN are excluded, so SSH, LuCI and the panel
   survive. Measured on the stand: 3 client flows cut, the live SSH session and
   the router's own connections untouched; `conntrack` CLI confirmed absent
   there, which is why this is ctnetlink.

4. The untunnelable text claimed Linux/macOS traceroute "still prints hops". It
   prints none, under any policy: the UDP probe is delivered locally by tproxy,
   local delivery does not decrement TTL, and no router raises time-exceeded.
   `traceroute -I` is what works.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:01:00 +03:00
omarandClaude Opus 5 ea3a4c518e test(generate): the L3 ingress is the default now — say so in the fixtures, not in 51 rewrites
The l3_tunnel default flip (164b703a7) turned 51 tests in shater/generate red.
Two premises had changed, and each is repaired where it broke rather than at the
assertion:

- ~43 fixtures build an engine-topology model with no inbounds at all and assert
  "this config produces no diagnostics". On the seeded-ON default such a model
  earns an honest `icmp "tunnel"` warning: the L3 ingress is fed only by the
  tproxy divert plane, and a model with no tproxy inbound raises none. The
  warning is TRUE of those fixtures — they are not routers. So they now say they
  run neither router-wide plane (nonDNSGlobals became plainGlobals, and gained
  the same treatment for l3_tunnel that D24 gave dns_intercept), and every
  "no warnings" assertion keeps its original strength instead of being loosened
  to "no warnings except this one".

- 8 assertions counted len(opts.Inbounds). The subject of every one of them is
  how many TPROXY LISTENERS survive a guard, and a total that also counts a
  synthetic inbound answers a different question — one whose right number
  changes whenever an unrelated global flips. They count tproxy listeners now,
  and while there they gained the assertion the count was standing in for: that
  the SURVIVOR of the clash guard is the first-declared listener, and that two
  distinct ports keep the ports their nft diverts aim at.

TestL3TunnelOffEmitsNoTunInbound had lost its meaning rather than its fixture.
It read the default and asserted "off", so after the flip it was pinning
DefaultGlobals, not l3_tunnel. It now sets the opt-out explicitly and says why
the opt-out has to keep working, and TestL3TunnelOnByDefaultEmitsTunInbound
pins the other direction — that a model which never mentions l3_tunnel gets the
ingress — which nothing in this package did.

TestSniffIsNotAnInboundField asserted "exactly 1 inbound" purely so it could
index ins[0]. It checks every emitted listener now and counts what it checked,
so the guarantee that assertion was really providing (the loop ran) survives
without a count that any future synthetic inbound breaks for no reason.

The warning text is rewritten. "l3_tunnel is on but no tproxy inbound is
enabled" accused the reader of a choice they no longer made: since the flip it
is the default, and a message that reads as "you turned this on" sends them
hunting for a switch they never touched. It now says what is not happening, that
the ingress is on by default, and names BOTH exits — a tproxy inbound restores
it, `option l3_tunnel '0'` says the router does not want it — because which one
is right is a fact about their router the generator cannot know.

model/dnsintercept_test.go had the blindness its l3 twin documented: a plain
strings.Contains is satisfied by `#option dns_intercept '1'`, and the parse half
cannot tell either, because a commented option falls back to the seed, which
since D24 is also true. A config shipping the option commented out would have
passed both halves while giving a fresh install no visible option to flip. The
check is line-wise and comment-aware now, and its "config unreadable" branch is
a Fatal instead of a Skip — a guard that skips itself is how one ends up
reporting ok while guarding nothing.

Mutation-verified, each reverted after: seeding L3Tunnel=false fails the
default test by name; removing the l3_tunnel guard fails the opt-out test;
stripping either exit from the warning fails TestL3TunnelWithoutTproxySkipped;
setting a legacy SniffEnabled on the tproxy listener fails the sniff test;
disabling the listen-clash guard fails TestDuplicateTproxyPortSkipped; freezing
the tproxy port at the default fails TestMultiLanDistinctTproxyPortsBothKept;
commenting out the shipped dns_intercept fails the shipped-config test (and the
parse half stayed silent, which is the blindness).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:54:48 +03:00
omarandClaude Opus 5 0f69880150 test(gate): a skipped test is a test that did not run — name it, or fail
Three holes, one shape: work that reads as coverage and is not.

1. shater/apply's TestApplyInstallsHoldWhenEngineFailsToStart — the only
   end-to-end test between "the engine died" and "the LAN forwards to the
   WAN in the clear" — asserted nothing. It broke the engine by pointing a
   rule-set at /nonexistent/nope.srs and stood itself down with t.Skip when
   that failed to break anything; it stopped breaking anything once
   LocalRuleSet.reloadFile began treating an unreadable file as empty.
   Measured in golang:1.26: the skip fired unconditionally and the package
   still printed `ok shater/apply`.

   It now injects the failure at the engineApply seam — the branch under
   test is applyLocked's, and a particular cause that stops causing retires
   the test silently — and COUNTS the seam calls, so applyLocked ceasing to
   go through it fails by name instead of quietly asserting something else.
   Everything else stays real: the model, generate, the kill-switch
   decision, netplane.RenderHoldNft, the latch, Status. New companion
   TestEngineApplyReallyFailsWithoutStarting is the control that the real
   engine.Apply can fail with the engine left stopped, so the simulated
   state is one this fork can be in.

   Mutation-checked both ways: drop the holdLocked call from applyLocked and
   the test fails with "0 holding planes were installed, want 1"; bypass the
   seam and it fails with "the engine-swap seam ran 0 times, want exactly 1".

2. warnings_test.go had two of the same genre. The len(genWarnings)==0
   t.Skip is now a t.Fatal — an unloadable blocklist must always warn, and a
   generate that stops saying so is the W7 regression, not a reason to stand
   down. TestStatusWarningsAlwaysNonNil pins readConfig itself: its
   "zero warnings" assertion was true on a build host only because the
   config read failed SILENTLY, so once that failure started publishing a
   critical warning the same line meant two different things in two
   environments.

3. The gate could not see any of it. It now runs the suites with -v and
   matches every `--- SKIP` against SKIP_DECLARED; an undeclared skip fails
   BY NAME, a declared one prints its reason on every run. check_skips
   proves its own instrument first (no `=== RUN` line => the check was
   reading a blank page), and it also reports on a suite that failed
   elsewhere, so a red tree cannot become a hiding place. -v costs no test
   time (38/25/24 s plain vs 38/24/24 s, warm) — only output, which is
   filtered on a green run.

Also closes the same hole one language over: [6/7] requires every non-Go
test file in the tree to be claimed by a named runner, and [7/7] runs the
ones this gate owns with a verdict by name. openwrt/luci-app-shater/tests/
status-readout.test.js — 24 assertions over the one screen an operator
reaches while the LAN is cut off — was executed by nothing at all, and
[1/7] could not report it because `go list` is its instrument. The non-Go
suites run on the HOST before the docker re-exec, so the local loop really
executes them rather than printing "did not run" every time; where there is
no node at all they are named and the notice replaces the closing banner.

Controls, all run and reverted: a planted t.Skip is caught and named; a
planted failing .test.js is caught and named; an unclaimed test file is
caught and named.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:30:39 +03:00
omarandClaude Opus 5 164b703a7d feat(l3): ping travels the tunnel by default, and every LAN zone can reach it
l3_tunnel was opt-in, and "off" had no honest win left in it. Off, a LAN ping
is decided by `untunnelable` alone and every rung is a drop (block) or a
disclosure (icmp/direct send the echo out of the WAN with the client's real
address). "Ping works" was never the state where ping was tunnelled — it was
the state where ping was leaking. On, an L3-capable outbound carries the echo
and one that is not drops it honestly: adapter.JudgeFlow returns ActionDrop for
an ICMP flow whose outbound is not a tun.Port, so no reply is forged. The price
is a standing TUN + gVisor netstack, ~2 MB RSS, and it is stated where the
option is.

The switch stays. It is a real answer on a 32/64 MB device and when bisecting
whether the L3 ingress is what broke a box — but it is now a WARNED answer:
ValidateGlobals says what the off state does to ping and names the policy that
takes over. Two combinations also changed meaning and are now reported:
untunnelable=icmp is no longer "block plus working ping" (the prerouting L3
mark claims every ICMP packet before the forward chain the echo accept lives
in, and a LAN host's ICMP errors are marked in with them and dropped in the
TUN), and the existing =direct report gains a sibling rather than standing
alone.

The fw4 seeding was the second half of the same problem. The divert set spans
every LAN inbound and every iface:/zone: rule source, but 30_shater-core seeded
a forwarding into shater_l3 for `lan` only — so on a multi-zone router ICMP
from the other zones is marked, routed, accepted by `inet shater`, and dropped
by fw4's zone policy with nothing in any log. Every zone gets a forwarding now,
guarded by a scan of the actual src/dest pairs so a re-run adds nothing. Every
zone including an uplink, because guessing which zones hold clients is wrong
somewhere and a superfluous entry authorises nothing: accept_to_shater_l3 is
`oifname "shater-l3*" accept`, and the only thing that routes a packet into
that device is our own fwmark rule.

scripts/testbed-lao.sh builds the second LAN zone this needs to be visible at
all. It is not installed by the package — that is the whole opt-in mechanism.

Verified on local_openwrt (ImmortalWrt 25.12.1 r37978): three runs of the
seeder leave exactly one forwarding per zone (lan/wan/lao) and no existing
section altered; deleting the lao forwarding removes `jump accept_to_shater_l3`
from chain forward_lao and re-seeding restores it; with the idempotency guard
disabled two runs produce nine forwardings instead of three.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:26:00 +03:00
omarandClaude Opus 5 de6fa8ebf4 fix(doh3): Close is not an ownership handoff — stop pooling the query buffer
Review found the hole and it is real. My previous fix gave the pooled buffer to
the transport and released it when the transport closed the body, on the grounds
that "http3.Transport closes the request body on every path, hence the Once".
That sentence is true about how many times the body is closed and says nothing
about when — the failure mode this project keeps writing down.

Verified against the pinned quic-go: on every error path RoundTripOpt
(http3/transport.go:167-173) closes the body the moment doRequest returns, and
doRequest (http3/client.go:338-341) waits only on the request-CANCELLATION
watchdog — close(reqDone); <-done — never on the goroutine writing the body.
Nothing in quic-go joins that goroutine. So Close is not a handoff point, and
the sync.Once stopped a double Release while doing nothing about a read after
one.

One correction to the review's severity, since it changes what we tell people:
on the failure path the bytes do not reach the resolver. Every ReadResponse
error branch (http3/stream.go:325, :336, :343, :363) calls str.CancelWrite
BEFORE RoundTripOpt closes the body, so what the writer reads out of the
recycled buffer is thrown at a cancelled stream. The disclosure primitive is the
success path only; the failure path is a read of somebody else's memory, which
is undefined behaviour and a -race finding, and not shippable either.

Fixed by not sharing at all: Pack() into memory the body owns. The alternative —
a lock around Read and Close — would also be correct and was rejected because it
keeps a released-but-referenced object alive, and that is now twice in one day
that an assumption about quic-go's internal lifetimes has been wrong.

The cost is negative, measured rather than assumed: Pack is 87 ns/op at 64 B and
1 alloc against 108 ns/op at 64 B and 1 alloc for the pooled version, because
buf.NewSize allocates the Buffer struct itself — the same 64 bytes — and then
adds Get/Put on top. The pool was never saving an allocation here.

The failure path cannot be caught on the wire, so the new test pins the cause:
a query tagged with a random needle, an exchange that fails (server never
answers; context already cancelled), then the pool drained on the goroutine
RoundTripOpt ran on, demanding the needle is not there. Mutations run without
-race: restoring pooledRequestBody fails both subtests 5/5, and blunting the
scan trips its control. -race is a separate pass, green at -count=3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:12:15 +03:00
omarandClaude Opus 5 b91fba1295 fix(l3): a covered last fragment must complete; say what the timeout really does
Two review findings on the fragment reassembler.

1. A whole datagram could vanish. entry.total was assigned before addRange
   was asked, so a last fragment (MF=0) whose range was already covered by
   MF=1 fragments answered fragInsertDuplicate and returned nil — while the
   entry was already complete(). Nothing re-examined it, because every later
   fragment is a duplicate too, so it died at its deadline with all its bytes
   present. A duplicate now falls through to the completion check: it
   contributes no bytes (held bytes still win) but it does contribute the
   total length. This is what the documented first-wins policy always
   implied; the code just did not do it.

   The sender needed is non-conforming, so the old behaviour was safe rather
   than exploitable — but it contradicted the comment three screens up, and
   that comment is the next reader's only defence.

   Also closed positively: a last fragment declaring an end BELOW the bytes
   already held now poisons the datagram instead of quietly never completing.

2. The 5 s timeout was not a memory ceiling and the comment said it was.
   sweep ran only when a NEW key was created, so once fragmented traffic
   stopped, up to fragMaxEntries entries stayed resident indefinitely.

   Both halves are fixed, and the honest one is the comment. sweep now runs
   on EVERY fragment — an O(64) scan on a path that is already the rare one —
   which releases residue as soon as any fragment arrives instead of waiting
   for an unrelated new datagram. That still does not cover total silence, so
   fragTimeout now documents the guarantee the code actually keeps: bounded
   by fragMaxEntries/fragMaxTotalBytes at all times, released on the next
   fragment, NOT "freed within 5 s".

   No timer, deliberately: it would need a goroutine with a lifecycle tied to
   something returnDeviceWrapper has no teardown hook for, and a goroutine
   that must be stopped and might not be is a failure this project has
   already paid for — to reclaim at most ~1.1 MiB that only exists after
   fragmented traffic has already happened. What bounds growth is the byte
   and entry ceiling; this timeout's job is correctness, and for that a
   check driven by the arriving fragment is exact.

   The now-unreachable per-key deadline check is removed rather than left as
   dead defence in depth.

16 mutations, all red. M15 (duplicate returns early again) reds only the
buggy case while the control and the poison case stay green, so the test is
shown able to see both an assembled datagram and a lost one. M17 (sweep back
inside the new-key branch) reds the new test while both old timeout subtests
stay green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:08:25 +03:00
omarandClaude Opus 5 df078c3205 fix(model): the write rollback may not swallow its own failure
The rollback added for the "failed import commits the deletion" defect went
through migrate.go's staged(), which drops the revert's error on the floor
(`_ = u.Revert("shater")`). That is defensible where staged() lives — a
migration that cannot revert leaves a half-migrated config, wrong but visible —
and it is not defensible here, because the delta this path stages STARTS WITH A
DELETE OF THE WHOLE PACKAGE. A revert that silently does not take leaves that
delete in /tmp/.uci, the caller is told only "import failed" and believes
nothing happened, and the next `uci commit shater` from any process publishes
an EMPTY /etc/config/shater. The guard reintroduced the exact loss it was
added to prevent.

writeUCIWith now uses its own revertStagedWrite, which reports both failures.
migrate.go's staged() is untouched: changing its signature to suit this caller
would rewrite a contract three migration paths depend on, for a hazard those
paths do not have.

The wrapped error names the CONSEQUENCE and the one command that clears it
("a staged DELETE ... will publish it ... run `uci revert shater` NOW"), not
just the fact — "revert failed" tells an operator nothing about what it costs.
ErrStagedWriteStuck makes it machine-detectable, so a caller can tell "your
change did not happen" from "your change did not happen and this router is one
unrelated `uci commit` away from an empty config".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:06:52 +03:00
omarandClaude Opus 5 d0471b2418 build(shater-core): ship the keep.d entry, or the node inventory dies at the next flash
files/ is not installed wholesale — every path in Package/shater-core/install is
explicit — so the keep.d file added alongside it would never have reached a
router. sysupgrade's "keep settings" walks /lib/upgrade/keep.d/*, and without
this entry /etc/shater/subs does not survive a flash: the restored box has its
rules and its groups and no nodes for them to point at, and the only repair is
`sub update`, which needs the internet the tunnel was going to provide.

/etc/config/shater needs no entry — it is a package conffile and sysupgrade
already keeps it that way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:02:14 +03:00
omarandClaude Opus 5 3314927bef fix(panel): stop the readouts claiming things the daemon never said
Nine places where the panel asserted more than it could know. Each was
checked against the daemon before being changed, and the two that a test
can reach are pinned by tests proven with a mutation.

MULTICAST IPTV WAS AN INSTRUCTION, AND IT WAS WRONG. The `direct` rung
said "Ping, multicast IPTV, and connecting to a VPN ... all work", so
someone who wanted IPTV read it and moved to the most open setting on the
ladder — the one that also lets a client's ESP/GRE past the proxy — and
still had no IPTV. The stream is UDP; every rule the policy emits carries
`meta l4proto != { tcp, udp }`, and the fail-closed forward chain accepts
only the RFC1918/link-local daddr sets, with no 224.0.0.0/4 among them.
The daemon says so itself in the note drawn a few pixels below. IPTV is
now stated once, and it says it does not work.

THE `block` COST LINE WAS UNCONDITIONAL, and three settings contradict
it: an open kill-switch (no drops are emitted at all), Globals.L3Tunnel
(ICMP is marked into the engine's TUN before the forward chain) and
Globals.UntunnelableEgress (ESP/AH/GRE/SCTP are routed out a named
device). The last two were not in the panel's `Globals` type, so the page
could not have been honest about them even in principle; they were added
rather than papered over with a vaguer sentence, and the copy is now
derived from all three.

THE KILL-SWITCH WAS READ WITH `=== 'closed'`. The daemon decides with
!EqualFold(TrimSpace(v), "open") and `Status.kill_switch` is the raw UCI
string, so `'Closed'`, `' closed '` and `''` — all of which BLOCK on the
router — drew OPEN, amber, "Nothing is meant to be blocked", and through
protectionState downgraded a plane-less router from crit to amber. One
normaliser now, `planeState.killSwitchClosed`, used by all five callers
that had their own spelling of it.

AN UNREADABLE CONFIG IS NOT "TURNED OFF". `enabled`, `kill_switch` and
`panel_port` are sourced from the config and are placeholders when it
could not be read (new `config_readable`). That happens on a full
/overlay or an interrupted `uci commit` — exactly when the fail-closed
plane has the LAN cut off on purpose — and the daemon publishes
plane:"hold" with enabled:false. Checking `!enabled` first rendered
"Turned off", amber, no alarm, and pointed at a Settings page backed by
the same unreadable file. The check now comes first, carries the daemon's
"do not turn anything off to fix it", and the kill-switch readout refuses
to name a policy it could not read instead of printing ARMED from "".

Also: the holding plane promises "no client TRAFFIC reaches the WAN", not
"nothing" — DNS to the router still goes to the ISP in the clear, by
design, so the daemon can recover; the stats backend is bbolt, not SQLite,
and reclaims space by rebuilding the file, not by a VACUUM that does not
exist (and skips it when the disk cannot fit the copy); the lock screen
sent people to System → shater when the menu entry is admin/services/shater,
which is the one instruction the product gives to someone who has just
lost access; and the panel port is configured, not confirmed — a failed
listen is only a log line.

RULESET.FORMAT WAS DESTROYED BY RENAMING A LIST. The edit form rebuilt
the object from its own controls and has no control for `Format`, so the
value could only be restored over SSH. It decides how a `file` list is
parsed and stops a `url` .srs being read as text; without it the list
matches nothing, the rule stops firing, and the traffic falls silently
through to the next rule. Carried now for the two sources the generator
consults it for. The same class of loss is made loud elsewhere: the two
other rebuild sites return `Complete<T>`, so adding a field to `Inbound`
or `DNSRule` fails the build in the function that has to decide.

Egress.Target is deleted: it is not in the Go model, so the "which egress
points at this node" branches could never fire, and had anything ever put
a string on it PUT would have rejected the whole write under
DisallowUnknownFields.

One layout fix on the way past: at 390px the policy plate's grid column
was sized by the select's longest option, so the sentence beside it was
clipped mid-word — which is how a line about what leaks loses its second
half.

Verified: npm run build + tsc clean; 57 tests pass; mutation-checked by
restoring the old comparison, the old check order and the old rebuild in
turn, each time watching the matching tests fail with the exact inverted
reading; browser-checked at 390 and 1280 against the mock, which now
reproduces `?ks=Closed` and `?cfg=unreadable` verbatim instead of
normalising them out of existence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:01:08 +03:00
omarandClaude Opus 5 6801146240 feat(core): back up the product state, and let the watchdog see a crash loop
Two things the box could not survive, both silent.

BACKUPS CARRIED NOTHING. No shater package put a single entry in
/lib/upgrade/keep.d, so "keep settings" and LuCI Backup took /etc/config/shater
(a conffile) and nothing else. Everything the product knows besides UCI lives in
/etc/shater: the entire node inventory (subs/*.json, hundreds of nodes on the
live router), the boot-armor arm token, the compiled blocklists. Restored onto a
new router the config looked complete and had no nodes to route to — and the
repair, `sub update`, needs the internet the tunnel was supposed to provide.

keep.d/shater-core keeps subs/, boot.nft, lists/ and alert-state.json, and names
what it refuses and why: stats.db is history bounded only by stats_disk_limit_mb
(0 = unlimited) and the archive is built in RAM; cache.db is sing-box's cache and
a stale one is worse than none; shaterd.log is a log carrying the query history
of the box it came from.

THE WATCHDOG COULD NOT SEE A CRASH LOOP. /etc/init.d/shater respawns every 5s,
forever; shater-cron escalated only after five consecutive ticks where `pidof`
found nothing. A daemon dying seconds into startup is back before the next
60s sample, so the counter reset every time — while the fail-closed plane held
the LAN shut and the panel, served by that daemon, never came up.

The tick's sleep is now spent sampling the daemon's identity (via its pidfile,
not `pidof`, which also matches the CLI verbs this loop runs) every 5s. A tick in
which 3 different daemons lived is churn; two such ticks in a row is the verdict.
A legitimate bounce replaces the daemon once and is announced twice over
(RESTART_FLAG up, ACTIVE_FLAG down), either of which discards the tick.

The action is the one the operator already chose: kill_switch=open stops the
stack, exactly as the dead-daemon path does; kill_switch=closed — and an absent
or unrecognised value, which is the documented default — reports at daemon.crit
and leaves the decision to the person, naming the command that opens the LAN.

Also drops the ruleset loop from shater_run_due. `shaterd ruleset update` has
never existed; it exited 0, so the loop stamped every url rule-set as freshly
updated and fired a reconcile for work that never happened. Now that it exits
non-zero the same loop would emit ~288 syslog lines a day per rule-set instead.
The comment says who does own the refresh, and where the gap that is left is.

Verified: sh -n and busybox `ash -n`; the pure detector driven with synthetic
sample streams under busybox ash (13 cases); shater_sample_pid against a real
/proc with a live process named shaterd as the positive control; and the whole
chain end to end against a real 2s-lifetime crash loop. Each threshold and each
veto is pinned by a mutation that makes the gate fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:58:14 +03:00
omarandClaude Opus 5 2f8c692c39 fix(l3): the TUN is a reclaimable slot — one fixed name made every apply fatal
On the production router every configuration change with l3_tunnel=1 killed the
engine and held the LAN down, three times in a row:

  19:10:33  reconcile failed: start inbound/tun[l3-in]: open tun: TUNSETIFF: device or resource busy
  19:14:02  start instance failed and could not restore previous config; engine stopped
  19:14:38  reconcile failed: TUNSETIFF: device or resource busy

A new generation had to open the device the outgoing one still held. That alone
is a failed apply; what made it an outage is that the recovery path rebuilds the
PREVIOUS config, which named the same device — so the rescue failed for exactly
the reason it was needed. A recovery path must not depend on the resource whose
contention it is recovering from.

The device is now one of two slots, chosen by the ENGINE at box-build time, on a
copy of the options taken AFTER the hash — so the stored config stays canonical
and a no-op reconcile is still a no-op. It cannot be chosen in generate: generate
runs every minute and its output is what Apply hashes, so an alternating name
there would rebuild the engine once a minute forever.

Rotation alone was NOT enough, and that was measured, not reasoned: the two-slot
build survived five applies of five kinds and then failed on 4 of 10 back-to-back
changes with the original outage in full, because a retired generation keeps its
TUN until its budgeted Close finishes. So an occupied non-current slot is now
DELETED rather than waited for — the running generation's slot is excluded first
and never touched, every other slot belongs to a box that is carrying nothing.
No bounded wait: waiting on an asynchronous kernel teardown is the race this
design removes.

The firewall never learns which slot is live — our accepts and the fw4 zone match
`shater-l3*`, verified to validate AND load on ImmortalWrt 25.12.1 / nftables
1.1.6, so the ruleset is byte-identical across a swap. Routers seeded by a
pre-slot build are migrated in place, or fw4 would silently resume dropping the
forward.

A2: turning the feature off left the device, the ip rule and table 8200 behind —
addL3Routing returned early instead of tearing down, and nothing else owns that
device. The disabled branch and TeardownRouting now remove all three.

Two smaller lies found while proving this, both measured: `ip -6 route flush`
does not take a non-unicast route, so the fail-closed floor survived and the next
add answered `File exists` — reported as a CRITICAL "this table has no floor,
traffic can leave over the plain WAN" on every apply, about a floor that was
right there; and teardown left it behind. Fixed both.

Verified on local_openwrt (ImmortalWrt 25.12.1, kernel 6.12.94 — the router's
revision) before and after, with binaries built from the same tree: the pre-fix
binary reproduces the outage and the leftovers; the fixed one survives all five
apply kinds and 12 back-to-back changes and leaves nothing behind. Ten reverted
mutations, each shown failing. See D28.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:55:44 +03:00
omar c788425cad feat(stats): the connection log now says which rule sent it there
The tracker has carried the matched route rule and the outbound chain since
upstream (common/trafficcontrol/tracker.go Rule/Chain); nothing in shater/ ever
read them, so "why did this connection go out that exit" was unanswerable from
the log and cost hours per report.

ConnLogEntry gains RuleKind/Rule/Chain. Rule is the engine rule text, not the
model rule name: nothing survives generation that ties an emitted option.Rule
back to the /etc/config/shater rule it came from, and a guessed name would be
worse than none. RuleKind keeps the two empty cases apart — "default" is a
recorded fact (nothing matched, took route.Final), "" means not recorded at all,
which is what an old persisted row decodes to.

Both fields are interned, so the ring pays 56 B/row of headers instead of a
private copy of text that is identical across every connection one rule matched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
@
2026-07-26 23:47:49 +03:00
omarandClaude Opus 5 0144282f5e fix(apply): a finding that is still true may not erase itself
Three ways this package published calm over a router that was not doing what
its config said. All three are the inverted failure: not an error raised when
things are fine, but silence when they are not.

1. Critical policy-routing findings were erased by the next no-op reconcile.
   applyDataPlaneLocked set routeWarnings only on the full path; applyLocked
   published the set unconditionally, so a minute later the fast path replaced
   it with one that no longer contained the finding. Neither surviving finding
   ("this egress CANNOT REACH ANYTHING outside its own subnet", "table could
   not be given a fail-closed floor") makes RoutingPresent false, so nothing
   brought it back: zero findings, plane full, green, over an egress carrying
   nothing. The comment on the gate claimed the previous set stood; it did not.

   planeOutcome now distinguishes "nothing was found" from "nothing was
   checked" (routeMeasured, written only by measuredRouting), and applyLocked
   carries the last MEASUREMENT forward across the fast path. A re-measurement
   still retires a finding, so this is not a latch.

2. An unreadable configuration was published as enabled=false. The panel tests
   !enabled before plane and renders "Turned off", amber, no alarm, "turn it on
   in Settings" — over a LAN the boot armor had cut off, pointing at a settings
   page backed by the same unreadable file. Status now carries config_readable
   and config_error, plus a critical finding in section "config".

3. The reason the engine failed to start existed nowhere. holdLocked logged it
   and called no publisher, and Warnings carries the last SUCCESSFUL apply — so
   plane="hold" with an empty findings list was a normal state of the product.
   The cause is recorded and published at read time while the engine is down,
   so it self-clears when the engine comes up; the boot-time arm is a warning,
   a real failure is critical.

Each fix is mutation-checked, and the route-warning test carries its control:
it sees a live finding, sees it survive the fast path, and sees a re-measured
clean state retire it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:45:28 +03:00
omarandClaude Opus 5 b642e5d8fe fix(egress): an interface egress with no interface was bound to br-lan
generate/outbound.go resolved the bind device with netplane.IfaceDevice,
whose empty-name fallback is "br-lan" — correct for an INBOUND with no
network, a black hole for an egress. netplane.EgressDevice returns "" for
the same egress on purpose (it calls br-lan "catastrophic here"), so
addEgressRouting installed no `ip rule` and no routing table for that
egress's mark, and the prerouting marking and the forward-chain accept
skipped it too.

The outbound was therefore emitted with SO_BINDTODEVICE=br-lan and a
routing mark nothing routed: every node, group and rule bound to that
egress dialled public addresses out of the LAN bridge. Not a leak — the
bind pins the socket to the LAN — but a total, silent black hole, with the
panel showing a configured, applied egress and no findings at all. The
`if dev == "" { dev = eg.Interface }` line that stood there read as a
guard against exactly this and could never execute: IfaceDevice never
returns "".

- generate now calls netplane.EgressDevice — the data plane's own
  resolution — so a bind can no longer name a device the routing was never
  installed for, and ` eth1 ` binds what the netplane routes. A device-less
  egress emits NO outbound and is reported; every reference to it then
  resolves through egressDetourOrBlock to tagBlock, so the traffic is
  blocked rather than sent out over the plain WAN.
- model.ValidateEgresses reports the same egress on the config channel
  (netplane's own skip is silent), built on model.EgressHasDevice — the
  model-side twin of EgressDevice, which ValidateUntunnelableEgress now
  shares so the two model resolutions cannot drift either.
- TestEgressDeviceResolutionParity runs one table through
  netplane.EgressDevice and model.EgressHasDevice and requires one verdict,
  the same treatment TestUntunnelableEgressResolutionLockstep gave the
  earlier validator/data-plane divergence.

Also: the UntunnelableEgress comment claimed "the panel says which, at
apply time, from whether the device is point-to-point". It does not. The
operator-facing text states both possibilities and declines to claim
either, there is no UI for the option, and isPointToPoint is consulted
only to warn that a gateway-less device can reach nothing. Said so, so the
next implementer does not read a described feature as a built one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:44:35 +03:00
omarandClaude Opus 5 35e4900769 fix(shaterd): three success reports for work that was not done
`shaterd status` fabricated a status when the daemon was unreachable and
exited 0. The stub is the same struct, printed by the same marshaller, so the
only thing that distinguished it was `plane` being "" — a value a live
Applier.Status() cannot emit. luci-app-shater was forced to key its "daemon
down" verdict off exactly that side effect, and filling `plane` in the stub for
any reason would have silently turned "dead" into "fine" on that page.

Both branches now carry an explicit "daemon_answered" boolean, and the offline
branch exits 1. The field is ADDITIVE and spliced in, not re-marshalled: every
existing key keeps its name, value and position (including plane:"" — still
emitted deliberately so dashboard.js keeps working until it moves onto the new
field), and a newer daemon's unknown fields are relayed untouched.

model.writeUCIWith committed the staged package DELETION when the import that
was supposed to refill it failed: /etc/config/shater came out empty, the caller
saw only "WriteUCI: import: ...", the next ReadUCI reported Enabled=false and
the next reconcile tore the plane down. Both error paths now revert through
migrate.go's staged() instead — the same idiom, for the same reason.

`shaterd ruleset update` printed a note and exited 0. shater-cron runs it with
output discarded and, on a zero exit, stamps the ruleset as freshly updated and
sets changed=1, so every source=url ruleset was permanently "just updated" by a
verb that fetched nothing. notImpl now exits 1 (not 2 — a caller must be able to
tell an unimplemented verb from an unknown one).

pidfilePath becomes a var so the daemon-answered / daemon-absent split is
testable without writing to the real /var/run, mirroring ctlPath.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:39:17 +03:00
omarandClaude Opus 5 d3e33294c1 fix(l3): reassemble return-path IP fragments — classifyReturn drops them
sing-tun's forwardReturn.classifyReturn refuses to judge a fragment
(flow_parse.go sets `fragment` for IPv4 MF/offset and for an IPv6
fragment extension header; flow_dispatch.go:703 answers returnPass), so
a fragmented answer coming back through a WireGuard/AmneziaWG endpoint
falls through to the endpoint's own tun stack instead of the l3 return
path, and the LAN client never sees it.

Measured on the live router: `ping -c3 -s 1400` through an AWG tunnel
with MTU 1280 is 100% loss while the WAN capture shows 3 x (1312 + 208)
in both directions — the far host answers, the peer fragments the answer
to fit the tunnel, the fragments die in classifyReturn. `-s 56` is 3/3
and PMTUD with DF works end to end, so only the fragmented return is
broken.

sing-tun is pinned upstream with no `replace`, but the fix does not need
to live there: every decrypted packet passes returnDeviceWrapper.Write
before it is offered to ReturnPackets. Reassemble there and
classifyReturn gets a whole datagram.

Hard ceilings, because this runs on a 128-256 MB router: 64 concurrent
datagrams, 1 MiB of held bytes, 64 disjoint ranges per datagram, 65535
bytes per datagram, 5 s to complete (timer starts at the first fragment
and is never refreshed). Over any ceiling evicts oldest-first.

Overlap policy: a range contained in one already held is a duplicate and
is ignored (first-wins, deterministic) because benign networks do
retransmit; any PARTIAL overlap poisons the datagram until its deadline.
No conforming fragmenter emits one, and every historical hole in this
area comes from a reassembler that tried to resolve the conflict.

The MTU of shater-l3 is untouched (65535 on purpose) and sing-tun is
untouched.

14 mutations run against the tests; each turns at least one test red,
including the two that first survived (a stale-head reuse the sweep was
covering for, and a fast-path copy).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:38:07 +03:00
omarandClaude Opus 5 32aac89139 docs: stop the docs promising a safety net that ships disarmed
Every install recipe walked the reader through `shaterd apply` + `shaterd
confirm` as if commit-confirm were armed. It is not: DefaultGlobals() never
seeds ConfirmTimeout, the shipped config carries confirm_timeout '0', and
ArmRollback returns at once on a non-positive timeout. A reader following the
README believed an apply that cut their SSH would undo itself. It would not.
README/README.en/INSTALL now arm it in the recipe and say what 0 means; the
apply-flow diagram gained the edge it always took on a stock box.

The boot armor was documented nowhere at all (`grep -rli armor --include=*.md`
returned zero) while shipping enabled and blocking LAN->WAN on every boot.
INSTALL 4 now says what it is, why SSH/LuCI stay up on purpose, every condition
under which it refuses to arm, and how to switch it off.

Also removed or corrected, each checked against the code, not inherited:

* MASQUE/CONNECT-IP is advertised in both READMEs and absent from parse,
  generate and model -- registry names it among the types deliberately left
  unregistered. Dropped, with the fork-vs-product distinction spelled out.
  The inverse too: Hysteria2/TUIC/XHTTP were tagged [T1] while shipped under
  with_quic/with_xhttp; ShadowTLS is generate+registry only, no parser.
* `direct (flow-offload on)` -- no offload/flowtable/flow_offloading anywhere
  in openwrt/, shater/ or panel/src. The product does not do this.
* shater-core deps were two releases stale in two places, one of which vouched
  for a config.buildinfo check that never covered kmod-tun. Ruling narrowed to
  what was actually checked.
* PORTING's "Full schema" -- the shipped config points at it -- was missing
  l3_tunnel and untunnelable_egress (UCI is their only path; the panel does not
  show them) and the blocklist/allowlist/device/alert sections, while listing a
  `config preset` that ReadUCI has no branch for.
* ARCHITECTURE had no L3 ingress and no kernel egress at all, though both are
  [MVP] and one creates an fw4 zone in the user's firewall config. New 3a.
* nftset-for-routing in the DNS diagram: that is the v0.1 mechanism, gone in v0.2.
* CONTEXT described a pre-Phase-1 repo and a 24.10.3 testbed. The testbed is
  ImmortalWrt 25.12.1 r37978-cd0a06bfd3fd (read off the box), which is not a
  detail: .apk does not install on 24.10 at all.
* The gate existed and no .md mentioned it. README/README.en/CONTEXT now do.
* release.yml's header still described publishing as either/or after the rolling
  pointer became unconditional. Comment only.
* Shipped /etc/config/shater: schema_version '1' against CurrentSchemaVersion=2;
  a pointer to a dns_filter line that was not in the globals block (added, '0');
  and `option sniff '1'` on the inbound -- an option the model deliberately does
  not have, which the first panel save would have silently washed out.
* lx-changelog pointed at a D25 heading that does not exist.
* ROADMAP 2b and 5 were done and unmarked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:35:06 +03:00
omarandClaude Opus 5 9e6dda22b3 fix(http3,doh3): stop releasing what is still being read
Two suspicions, both put to a test rather than to a reading. Both were real, and
neither was the leak the suspicion named — both are objects released while still
in use.

roundTripHTTP3Race ran both racers on one cancellable context and cancelled it
before returning the WINNER. quic-go and net/http reset a request's stream when
its context dies, so the caller got a response whose body stopped mid-read:
H3_REQUEST_CANCELLED (local) (read 2687 of 65536 bytes). That path is taken
whenever there is no cached HTTP/3 connection and the request is replayable —
the first request to every host, and every one after an idle close. Each racer
now has a context of its own; losers are cancelled where everything used to be,
and the winner's cancel travels with its body.

DoH3's Exchange packed the query into a POOLED buffer and released it the moment
RoundTrip returned. But http3 writes the request body on a goroutine of its own
and returns as soon as the response HEADERS arrive — the body is still being
read. With the window held open the query on the wire diverges from the query we
packed at exactly offset 8192, quic-go's copy-buffer size: everything past that
was the next pool user's memory, sent to the resolver. Not a slowdown — a data
race and a small memory-disclosure primitive. The buffer now goes back when the
transport closes the body, which http3 does on every path, and can do twice.

Both files diverge from upstream again, hours after 0a6689b29 made them
byte-identical on purpose. Upstream carries the second defect in
dns/transport/https.go too; that file is outside this audit and is named in D27
so the next person finds it instead of rediscovering it.

sing-quic moves v0.6.2-0.20260525051024 -> v0.6.4-0.20260709034545. quic.go is
byte-identical across the two, so this neither duplicates nor retires the
packet-conn ownership fix — quic-go still does not own the socket. What it does
carry is the other half of the family we took only half of: clientConn.Close in
tuic/, hysteria/ and hysteria2/ now sets a past write deadline, word for word
the fix v2rayquic already had. We ship tuic and hysteria2. Cost, measured:
+256 KiB exactly on the stripped aarch64 binary and six indirect modules for a
realm port-mapping path nothing we generate can reach.

Tests are mutation-checked: reverting each fix makes them fail, with the text
quoted above. The DoH3 test carries its own control — it first proves the pool
does hand a released buffer back and that poisoning it lands, because a clean
result from an instrument that cannot produce a dirty one proves nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:26:03 +03:00
omarandClaude Opus 5 fde4bed571 fix(luci): stop calling the daemon dead when only the engine is
`running` changed meaning on 2026-07-26 (a8970b8ac): it was a hardcoded
true and is now the ENGINE's liveness (apply.go `Running: engineUp`).
dashboard.js was last touched on 15 July and stayed in the old epoch, so
a dead engine made the page report "Daemon (shaterd): not running" in
red, advise "start the Shater service first" — the service was running —
and DISABLE the button to the panel, which is the one place the config
can be fixed. The holding plane keeps management reachable on purpose
(netplane/nft.go: "The operator can always get in to fix the config");
LuCI was the only thing taking that guarantee away.

Daemon liveness is now derived from the wire, not from `running`. "The
ubus call returned" is not enough either: `shaterd status` EXITS 0 WITH
A FABRICATED STATUS when the daemon is unreachable (cmdStatus offline
stub), and that stub is the apply.Status zero value plus a UCI read — so
it carries enabled/table/kill_switch but leaves `plane` at "", a value
no live daemon emits. A known plane word is the positive proof a daemon
answered; an explicit empty one is proof none did. Everything else —
{} from a failed call, {"error":...} from the plugin (also what a live
but WEDGED daemon produces), a pre-`plane` daemon — is unknown, and
unknown is an unlit lamp, never green. The launcher button is never
disabled again: a mint that fails already reports itself.

"Interception: active" is gone. apply.go says of `active`, verbatim:
"Never render it as 'we are proxying'" — it is the run latch that gates
hotplug and cron, it stays raised while the engine is down and the LAN
is blocked, and this page painted it green next to two more green lamps
in exactly that state. It is now "Service latch", and its lamp reports
only whether the latch agrees with globals.enabled. The row that was
missing is `plane`: full / hold (LAN->WAN BLOCKED) / none. `traffic` is
shown too, because plane=full is not "tunnelled" — a `default -> direct`
router has a full plane and no tunnel at all.

The rpcd plugin's status docstring listed five fields of fourteen and
had done since before half of them existed; it now describes the real
shape and the two fields that are easy to misread.

tests/status-readout.test.js runs the derivation against six recorded
status shapes with no browser and no router. Mutation-checked: reverting
to `st.running` fails 14 assertions including the operator-visible
"not responding - start the Shater service" over a live daemon;
restoring the "Interception: active" row fails 9; putting
openBtn.disabled back fails 1 by name; opening the closed plane list
fails 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:22:56 +03:00
omarandClaude Opus 5 cb26936ebf fix(wgdedup): merge identical WireGuard copies instead of blocking one
A rule pointing at node:awgout, which was already the first hop of the
default-route chain, took the house off the internet for two minutes.
The pass saw one private key materialised twice, kept the copy that
sorted first alphabetically, and fail-closed everything that routed
through the other one — which happened to be the default route for all
traffic.

The mechanism was right and the framing was wrong. The physical limit is
one DEVICE per key, not one mention per key. Two copies that build the
same device — same key, same peers, same address/MTU/AWG parameters and
the same dialer — are one device written down twice, and there is nothing
for them to fight over. Those are now MERGED: one survives and every
reference to the others is rewritten to it, silently. That makes the
shape the owner wanted expressible: one chain using awgout as an
intermediate hop and another using it as a terminal, both entering over
the same egress, coexisting on one device.

Identity is the marshalled options blob rather than a hand-picked field
list, so a field added to WireGuardEndpointOptions or DialerOptions later
reads as "different" instead of being silently merged.

Only a real incompatibility — different detour, different peers,
different device parameters — is still two devices, and then:

  - the survivor is chosen by WEIGHT, not by tag order: reachability from
    route.Final (the default route) dominates, breadth of use breaks
    ties, tag order only settles a true tie;
  - the warning names the consequence. "Everything that routed through X
    is fail-closed" is equally true of a stray test rule and of the whole
    house's default route, and that is what the operator read it as. It
    now says which of the three it is, measured on the finished config:
    the default route is dead, or it survives via another path, or it
    never touched the lost copy.

A merge must not rename away the subscription fetch detour: that
reference lives in the model and is resolved against the running box, so
this pass cannot rewrite it. Such tags win the survivor slot outright,
which costs nothing since every copy in a class is the same device.

Tests: identical copies coexist on one device; a real incompatibility
keeps the default-route copy even when it sorts last and says so; the
warning does not announce an outage when the default route survives
through a group, and does announce one when it dead-ends behind a
surviving exit; no duplication at all is a no-op. All seven mutations
(merge off, weight off, member-dedup off, pin off, detour-following off,
consequence collapsed, plus a positive control) fail the suite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:16:12 +03:00
omarandClaude Opus 5 8db29b6267 fix(apply): say when shaterd apply armed no safety net
`shaterd apply` exists for one reason: snapshot the last-good, apply, and arm
an automatic rollback so a change that costs you access to the router undoes
itself. It answered `{"changed":false}` and not one word about that.

On the live router (2026-07-26) that was a trap. The operator edited UCI, ran
`uci commit`, the `config.change` reload trigger had already restarted the
daemon, and the fresh daemon applied the new config on startup. By the time
`apply` ran there was nothing left to apply — and the last-good it snapshotted
as the ROLLBACK TARGET was the newly applied config itself. The watcher was
armed onto the very configuration it was meant to protect against: firing it
would have restored exactly what was already loaded. No safety net, no word
said, house offline.

The verb now answers the question it exists to answer, in a closed vocabulary:

  rollback_armed  true ONLY when a window was armed AND its target differs
                  from what is running. An armed watcher pointing at the
                  running config is not a net and is not reported as one.
  reason          applied | already-applied | nothing-to-apply | disabled |
                  commit-confirm-off | config-unreadable | apply-failed
  message         the same thing in the operator's words, never empty.

The two "nothing moved" cases are told apart where they CAN be: an
/etc/config/shater mtime later than this daemon's start, with the running
config already matching it, can only mean a reconcile beat this command to it
(reason=already-applied). Where they cannot — the `uci commit` reload trigger
is stop+start, so it moves the daemon's start past the edit — the text says
so instead of reading as success: no net, harmless if you changed nothing,
unprotected if you did, and shaterd cannot tell which.

Two silent holes surface as a side effect, both previously reported as plain
success: `confirm_timeout=0` (the SHIPPED DEFAULT in
openwrt/shater-core/files/etc/config/shater) makes ArmRollback a no-op, and a
failed post-apply ReadUCI skips the arming entirely.

Arming behaviour is byte-for-byte unchanged — this only makes its absence
visible. A real safeguard for the already-applied case is separate work.

Tests are mutation-verified three ways: reverting classifyApply to the old
{changed,error} fails 11 tests; blinding the mtime discriminator fails exactly
the discriminating one (and falls back to the honest ambiguous text); making
sameConfig always report "different" fails every invariant that forbids
claiming a net over an identical target.

NOT verified on hardware: local_openwrt was held by another agent, so the
control-socket round trip and the real mtime/daemon-start comparison have not
been exercised on a router.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:12:38 +03:00
omarandClaude Opus 5 564033cd10 fix(apply): chain: as a subscription fetch detour resolved to a name nothing answers to
`fetch_detour=chain:<X>` never worked. engine.ViaToTag maps "chain:X" to the
bare tag "X", but the generator materialises a chain as one wrapper per hop —
chain-<X>-h1..chain-<X>-hN — and routes into the LAST one. The lookup missed and
the update failed with "unknown outbound tag".

It failed CLOSED, so the feed was never pulled over the plain WAN by this path.
But the miss had a sharp edge: when a node or group happened to share the
chain's name, the lookup HIT it, and the subscription was fetched through a
completely different outbound with nothing said.

Applier.HTTPClient now resolves chain: before the engine sees it, against the
tags the RUNNING box actually holds (outbounds unioned with endpoints — a WG hop
is an endpoint and Outbounds() does not list those), mirroring the generator:
the highest-indexed chain-<X>-h<i> wrapper is the entry, and a chain that
flattens to one hop IS that hop. Every other via form is passed through
untouched.

The case the generator cannot serve is named rather than papered over: chains
are built lazily, only for a chain some enabled rule/egress/DNS detour targets,
and a fetch detour is not one of those references — so a chain nothing else
points at has no outbounds at all. That, and every other miss, is an explicit
refusal wrapping engine.ErrOutboundUnknown (the panel already maps it to 400).
Never a fall back to direct: that would put the feed and the owner's real
address on the plain WAN, which is the thing fetch_via=proxy is set to avoid.

Tests are mutation-checked. Pre-fix behaviour resolves "work"/"solo" and kills
every chain case; first-hop-instead-of-last, member-copies-count-as-hops,
dropped pass-through, and a silent direct fallback each kill their own test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:09:35 +03:00
omarandClaude Opus 5 add90b5b2f fix(panel): the DNS footnote was a grid item nobody placed
`.dns-filter-note` under the endpoint-resolver readout is a DIRECT child of
`.dns-filter-card`, so it is a grid item. With no explicit span it auto-placed
into column 1 — the toggle's `auto` track — and sized that track to its own
max-content: 237px at 390px, 322px at 1280px. That left the `1fr` copy column
with 0px, so "Network-wide ad & tracker blocking" laid out one word per line
and spilled 2px past the viewport, scrolling the whole page sideways on a
phone. On desktop the same cause parked the 52px toggle in a 322px column,
270px away from the copy it labels.

Measured at 390px: documentElement.scrollWidth 377 vs clientWidth 375. With
`grid-column: 1 / -1` on the footnote: 375/375, and the track list goes from
`237px 0px` to `52px 185px`. Cancelling just that one declaration in the live
DOM puts 377/375 and `237px 0px` straight back, so nothing else contributes.

Verified with playwright over 320/360/375/390/414/430/480/560/640/720/768/
1024/1280/1440: zero horizontal overflow at every width, with every rule
editor open, all three master toggles flipped, every source tab, and every
resolver type. No `overflow-x: hidden` anywhere — the page does not scroll
sideways because nothing overflows, not because the symptom is hidden.
Focus rings and prefers-reduced-motion re-checked and unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:06:30 +03:00
omarandClaude Opus 5 a571bd0e1a docs(claude): model is the executor's call, skills are mandatory, standards that earned their place
The old file pinned every subagent to fable — which broke the moment that
quota ran out mid-session — and spent half its length on panel scaffolding
that has been done for weeks. It said nothing about the test gate, the
testbed, or the hardware router, so none of that reached a subagent unless
it was retyped by hand into the brief.

What is new is not advice, it is the list of things whose absence cost a
day each: a test must be mutation-checked or it is decoration; an
instrument with no control proves nothing; a subagent must be told it may
refute the orchestrator, because the best results this project has had
arrived exactly that way; a formally-true sentence that reads as "it works"
is still a lie.

Skills are now a table mapping this project's areas to the skills that
cover them, with the rule that they are invoked BEFORE the work rather
than after something failed to run, and that every brief must name them —
a subagent cannot see this conversation and will not guess they exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 22:48:05 +03:00
omarandClaude Opus 5 1267d20fb8 docs: drop the L3 handoff note — it is merged, and it said to
test / go + panel tests (push) Successful in 8m33s
release / test gate (push) Successful in 8m8s
release / apk aarch64_cortex-a53 (push) Successful in 6m33s
release / apk x86_64 (push) Successful in 3m45s
release / release apk (push) Successful in 8s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 20:00:46 +03:00
omarandClaude Opus 5 35f697ed08 docs(openwrt): say why mtu_fix is inert instead of claiming an MTU we no longer set
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:52:39 +03:00
omarandClaude Opus 5 d0fb6befb1 fix(l3): the l3-in MTU is not a tunnel budget — 1420 was a forgery generator
shater-l3 was created at 1420, the WireGuard payload budget, copied one
layer too far out. It bought nothing: what actually goes into the tunnel
is sized by sing-tun's forwardToPort against Port.PortMTU(), which
already fragments to the outbound MTU without DF and answers a
well-formed `fragmentation needed` quoting it with DF. All 1420 did was
make the KERNEL split every packet above 1392 bytes of payload on its
way into the device -- and a fragment is the one thing sing-tun will not
judge. Dispatch returns on parsed.fragment before calling JudgeFlow, the
fragments reach the gVisor stack, it reassembles them, and the ICMP
forwarder's installFlow demands an unspecified port address that a
WireGuard endpoint never has. So it declined and answered the echo
itself. `ping -s 1392` honest, `ping -s 1393` a lie, and only for the
outbounds the feature exists for.

65535 rather than merely "large": no IP datagram can exceed it, so the
kernel cannot fragment at this device for any packet ever. Anything
smaller leaves a band open and re-opens the class. It is also sing-box's
own default TUN MTU on Linux.

Memory was measured, not argued. Three paired runs of the integration
test under -test.memprofilerate=1 allocate 5.41/5.48/5.47 MB at 65535
against 5.76/5.46/5.70 MB at 1420, and a -diff_base profile puts every
difference in netlink interface enumeration. Nothing in the read path
scales with the MTU: gVisor reads through fdbased.BufConfig, which
sing-tun pins to one 65535-byte view regardless. I predicted a ~1.8 MB
saving from GSO switching off above 49152 and was wrong -- protocol/tun
turns GSO back on at StartStateStart whenever a FlowOutbound exists, so
the GRO scaffolding is there at both values. The corrected reasoning is
in the constant's comment so the next reader does not redo the mistake.

The integration test now reads the MTU back off the real kernel device,
which is the assertion the value exists for: a kernel that clamped it
would restore the forgery without changing a generated byte.

D25's KNOWN HOLE block is replaced with what is genuinely left. Chiefly:
a big non-DF ping does not start WORKING, it starts failing HONESTLY --
classifyReturn declines fragments on the way back too, so the packet
really leaves, the far host really answers, and the reply is not NAT'd
home. And a client that fragments on the wire itself is still uncovered;
that is the nft carve-out's job, with a warning that conntrack defrag
may reassemble in prerouting and leave such a rule unable to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:50:08 +03:00
omarandClaude Opus 5 81c96019b5 fix(panel): let a routing rule say ICMP, instead of calling one broken
The Proto picker was a closed list of the two transports and the ten
sniffed L7 labels, and anything else drew "<value> — never matches".
The engine now routes ICMP by rule (Rule.Proto accepts icmp, icmpv4,
icmpv6), so a working ping rule was rendered as a dead one and could not
be created here at all — the operator had to hand-edit /etc/config/shater
and then watch the panel call the result broken.

Adds a third group, "Layer 3". All three spellings are offered: they are
not synonyms — icmpv4/icmpv6 pin the rule's ip_version — so hiding the
narrowing would both strand a capability outside the UI and silently
widen such a rule the first time someone edited it here.

The doc comment no longer claims the list IS generate/route.go's
sniffedProtocols; only the middle group is. ICMP goes to the emitted
rule's `network`, never to `protocol`, which is the whole reason it never
matched as a sniffed label.

An unknown value is still kept and offered as written, but the
never-matches flag is now judged on the lower-cased value, the way the
engine judges it — a hand-written `ICMP` is a live rule, not an inert one.

Verified: npm run build clean (tsc --noEmit + vite build); an icmp rule
added through the panel renders as a plain "PROTO icmp" chip; no
horizontal overflow at 360px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:26:07 +03:00
omarandClaude Opus 5 baed8ff8f2 fix(model): fwmark_base 0x7f routes the engine's own traffic into its own TUN
The panel offers fwmark_base and table_base as free hex fields under
"Advanced" and nothing has ever checked them. What makes that more than a
footgun is that the derived values are invisible from the number typed: the
L3 mark is base+0x80, so 0x7f lands it exactly on 0xff — the loop-guard mark
the engine stamps on its OWN traffic — and `ip rule fwmark 0xff lookup 8200`
then captures everything the engine sends and routes it into the engine's
TUN. The router loses the internet the moment l3_tunnel is switched on, for
a reason nothing on screen connects to a collapsed section. fwmark_base 0xff
had produced the same failure since long before the L3 offset existed.

table_base is worse and got the same treatment: its derived values can land
on the kernel's own table ids, and teardown does `ip route flush table <n>`.
It is count-sensitive (egress #i uses base+0x10+i), so the check takes the
egresses rather than living in ValidateGlobals.

Written as "derive every value this layout produces, then look for
duplicates and reserved ids" rather than as a blacklist, so a future offset
is covered by construction. The layout constants are duplicated from
netplane (the import only runs one way) and pinned by netplane's
TestMarkLayoutConstantsLockstep.

Warn-only, like every check in this file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 f80fb4dd1b fix(netplane): give every mark-driven table a floor, and check the L3 pair
Two halves of the same omission.

1. A fwmark lookup that finds an empty table does not fail — it falls
   through to main. Every mark-driven table now gets an `unreachable
   default` at the maximum metric: it loses to any real default route while
   one exists, it has no device so the kernel never garbage-collects it, and
   it turns "lookup failed, try main" into "lookup succeeded: unreachable".
   The fallthrough stops depending on somebody reading a warning at the
   moment an interface goes down. Deliberately not gated on the kill-switch:
   that switch decides whether traffic may escape the tunnel, while an egress
   binding is a statement about WHICH UPLINK, and silently substituting a
   different one is not what "fail open" was meant to permit.

   RoutingPresent's "does this table have a default route" test is tightened
   in the same breath, or the floor would answer it and turn the safety net
   into a blindfold.

2. RoutingPresent had never heard of addL3Routing. This is the same defect
   its own comment describes as already caught twice ("a presence check must
   cover everything its Apply counterpart installs"), committed a third time
   — and its trigger needs no interface to go down: editing a node URI
   restarts the engine, the kernel destroys shater-l3 and takes `default dev
   shater-l3 table 8200` with it, the rendered nft text is unchanged, so the
   fast-path skipped ApplyRouting forever and LAN ping stayed dead until
   someone restarted the daemon.

TestRoutingPresentSeesL3Table, TestEgressTableGetsFailClosedFloor and
TestEveryStampedMarkIsRoutedAndVerified all fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 b71b793681 fix(netplane): a mark says where a packet was sent, not where it went
The forward chain let untunnelable-egress traffic past the kill-switch on
the strength of its fwmark alone. `ip rule fwmark X lookup N` does not
deliver the packet to table N, it delivers the LOOKUP there — and a lookup
that finds nothing falls through to main. So when the egress interface goes
down and the kernel garbage-collects its default route, every non-TCP/UDP
packet from the LAN is still stamped, still accepted here (above the
fail-closed drop), and leaves out the plain WAN with the router's real
address. Nothing we render changes, so no apply runs and nothing notices.

Ordinary egress traffic never had this hole: the engine binds those sockets
to the device, and a dead device fails the socket. The untunnelable-egress
path is made of nothing but a mark, so the accept now carries the second
opinion instead — `meta mark X oifname "dev"`, strictly narrower than either
half, true only when the routing did what the mark asked. The comment being
replaced argued correctly that oifname ALONE would be too loose, then drew
from that the conclusion that oifname should be dropped rather than added.

Same conjunction in the holding plane, where it is theory (that plane stamps
nothing) but where a bare mark accept has no business sitting.

Also folds the egress device resolution into one EgressDevice(), because the
binding and model.ValidateUntunnelableEgress had already drifted: the
validator trimmed the interface name and the binding did not, so `option
interface '   '` gave a panel saying "the option is ignored" over a data
plane that was marking packets for a table nobody built.

TestUntunnelableEgressAcceptIsBoundToItsDevice and
TestUntunnelableEgressResolutionLockstep fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:32 +03:00
omarandClaude Opus 5 61c87ad1d9 fix(l3): guard the ICMP honest-drop at PreMatch, not inside the walk
The drop that keeps a ping from reading as tunnelled lived in
preMatchFlow, overriding the pre-declared continueResult. That covered
every exit of THAT function and none of the walk above it: the
prepareMatchMetadata error return (which arrived later, with the shared
metadata refactor), the sniff bail-outs, and the default: arm of the
rule-action switch all returned PreMatchContinue on their own.
adapter.JudgeFlow maps Continue to tun.ActionAccept, and sing-tun answers
Accept by rewriting Echo into EchoReply itself -- the exact forgery this
delta exists to remove. Narrow paths, but paths.

PreMatch is now a funnel over the renamed preMatch walk, so the guard
sits on the single return value and cannot be outgrown by a new exit.
PreMatchBypass joins the drop: sing-tun implements ActionBypass on the
nfqueue plane only, so on the TUN path it lands in the same default: arm
as Accept and forges too.

Every ICMP case has an explicit TCP/UDP twin; the JudgeFlow mapping
table is pinned outright, including the one fix that must NOT be made
there -- refusing ActionFlow for a port whose address is not unspecified
would drop every ping through WireGuard/AWG, because the forward
dispatcher and the ICMP forwarder share that function with identical
arguments and only the latter needs an unspecified address.

That leaves a real hole open, now named in D25 rather than papered over:
a FRAGMENTED echo to a WireGuard/AWG outbound is still answered by the
router. The dispatcher returns before asking for a verdict at all when
the packet is a fragment, and the reassembled packet reaches the ICMP
forwarder, whose installFlow demands the unspecified address a WireGuard
endpoint never has. The two fixes that would close it both live outside
pre-match and are written down; the Consequence paragraph is scoped
until one lands.

The stack comment in generate/inbound.go repeated the "only gvisor
really forwards ICMP" argument that D25 itself retracts -- both stacks
run the same ForwardDispatcher first. Brought in line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:21:29 +03:00
omar 4dee508e12 fix(route): let a rule say "icmp", and say when saying it is a lie
`icmp` fell through ruleMatchers' proto switch into RawDefaultRule.Protocol —
the SNIFFED-L7 field, compared against what the sniffers labelled a connection.
Nothing ever labels a flow "icmp" (PreMatch skips the sniff action for an ICMP
flow outright), so the rule was structurally valid and permanently dead. That
made the whole L3 ingress unusable on a real config: with no way to write "ICMP
goes here", every ping fell to the catch-all, which resolves to the chain's last
hop — a group of VLESS nodes that cannot carry layer 3 at all.

icmp is a NETWORK. NetworkItem.Match is a map lookup over metadata.Network, and
adapter.JudgeFlow sets that to N.NetworkICMP for BOTH ICMPv4 and ICMPv6 (one
case covers both protocol numbers), so there is exactly one network value and it
covers both families. `icmpv4`/`icmpv6` narrow that same network with an
ip_version item instead of inventing a second one: metadata.IPVersion comes from
the destination address, and an ICMPv6 packet always has an IPv6 destination —
no false positives, no false negatives.

An ICMP rule that cannot fire is not a dead setting: ICMP has no fall-through,
so route.preMatchFlow DROPS it. Four ways to get that silently are now reported:
l3_tunnel off (nothing enters the engine at all), icmpv6 with ipv6 off (neither
the nft mark nor the TUN address exists), a port matcher next to it (JudgeFlow
zeroes both ports), and a target that cannot carry layer 3 — decidable from the
model, because the capability is fixed by the outbound TYPE: only wireguard/AWG
endpoints and the direct outbound behind direct/interface egresses declare
N.NetworkICMP. A mixed group gets its own text (the answer follows group.Now()),
`block` gets none (dropping the ping IS the policy), and an unresolved target
gets none either (ruleKillFallback already said the louder thing).

Wording stays clear of shater/apply's criticalMarkers on purpose: a failed ping
is fail-CLOSED, and a cosmetic alarm is how the real one stops being read.
2026-07-26 19:18:06 +03:00
omarandClaude Opus 5 76da5134ef test(gate): the two tests that need a kernel may not skip in silence
The L3 branch adds TestIntegrationL3TunInboundStarts and
TestIntegrationL3EgressICMPIsAFlow — the only tests that prove the engine
really opens shater-l3 and that the egress outbound really is a FlowOutbound.
Both need root plus /dev/net/tun, both guard themselves with t.Skip, and the
gate could not see either: `go test` prints `ok <pkg>` whether a test ran or
skipped, so [2/5]'s per-package `ok` check is satisfied and the gate closes by
claiming it "passes every test we own". That is this script's own founding
failure (115 of 116 test files never running while CI stayed green) one level
down, and it would have shipped invisibly.

Two halves.

Where the capability CAN be granted, grant it. From a non-linux host the gate
re-execs into a container; that container now gets --cap-add NET_ADMIN and
--device /dev/net/tun, probed rather than assumed, so a plain
`scripts/run-tests.sh` on a dev box actually exercises the kernel path instead
of quietly stepping over it.

Where it cannot, say so where it cannot be missed. The act_runner is an LXC
guest whose kernel has no tun module at all (checked on 10.10.10.211:
`modprobe tun` -> "Module tun not found", /dev/net does not exist, act_runner
runs job containers with privileged:false and no container.options), so the
device cannot be handed down without reconfiguring the Proxmox host. New step
[5/5] therefore DISCOVERS every ^TestIntegration under the fork's trees — no
hand-kept list, so a privileged test written next month joins on the day it is
named — runs them with -v, and demands a verdict for each BY NAME: RAN, or
FAILED/MISSING (fatal), or SKIPPED while the environment could have run it
(fatal, because the capability guard cannot be what skipped it), or skipped for
a reason this box genuinely has — which replaces the closing banner, so the
last line of the gate can never claim coverage it does not have.
SHATER_REQUIRE_PRIVILEGED=1 makes that last case fatal for runs that can.

The discovery call carries -ldflags for the same reason every other call does:
`go test -list` links each test binary, and without -checklinkname=0 every
package pulling common/badtls fails to link. The first cut of this step omitted
it, swallowed the error, and printed "none declared" — a check against silent
skipping that was itself silently skipping. Its exit status is now inspected
and an empty list is only ever reported after a successful enumeration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 18:50:46 +03:00
omar 4c630c9a13 docs: handoff note for the L3 branch
Transient, to be deleted when omp/work merges. Everything meant to outlive the
merge is already in D25/D26 and the lx changelog; this file is the part that is
only useful while the branch is still a branch — the verification commands, the
testbed recipe, what was proven on hardware and what was not, and the six files
that will conflict on rebase.
2026-07-26 18:31:59 +03:00
omar d8dbefcd07 docs: record the AWG site-to-site path as declined, not impossible
D26's "no port-like selector" line disposes of NAT-based forwarding and nothing
else, and read alone it says "impossible" — which is false and would be
re-derived at the cost of another research pass. The endpoint is protocol-blind
in both directions, so ESP could ride it untouched with the client's own source
address and no NAT whatsoever. That was declined for two reasons worth naming:
lx-owned code in the forward hot path, and a server-side AllowedIPs prerequisite
that turns a router option into a deployment contract.
2026-07-26 18:31:59 +03:00
omar 974208fc05 docs: record the kernel egress, and retract the reason D25 gave for the ceiling
D26 writes down where the engine's boundary actually is, because the intuitive
answer is wrong and someone will look for it again: the WG/AWG forward path
never consults gVisor in either direction, so the limit is sing-tun's
ForwardDispatcher — its parser and its port-shaped NAT — and the kernel egress
was chosen because it clears that limit without a line of new hot-path code, not
because userspace "cannot". Tailscale documents the same boundary for their
userspace mode and is quoted as corroboration, with the caveat that ours sits at
the dispatcher rather than the stack.

D25 said two things that do not survive checking, and both are corrected in
place rather than left for the next reader to trip over. It blamed the netstack
for the ICMP-echo ceiling; that was the dispatcher. And it called `stack: gvisor`
mandatory because the system stack fakes ping — the system stack runs the very
same dispatcher first and only forges an echo for packets the dispatcher
declined, so gvisor is a deliberate choice (already linked via with_wireguard,
and the combination the integration test exercises), not a necessity.

The operator note says what the option buys and refuses to call an egress a
tunnel on its own say-so: with a WireGuard device it is one, with a second WAN
the destination sees that uplink's address. It also says what the option does
not fix — multicast IPTV stays broken — and that IPsec through NAT-T is ordinary
UDP that never needed any of this.
2026-07-26 18:31:59 +03:00
omar 2eb71e8244 feat(netplane,model): hand the protocols the engine will not dispatch to the kernel
ESP, AH, GRE, IGMP and SCTP cannot enter the engine, and the reason is not the
one that looks obvious. A WireGuard or AmneziaWG endpoint forwards straight past
its gVisor stack — WritePackets reads the IP version and the destination address
and hands the raw bytes to the device, and the return path offers every
decrypted packet back before the stack sees it. WireGuard would carry ESP today
if anything handed it one. What refuses is sing-tun's ForwardDispatcher: its
parser recognises TCP, UDP and ICMP echo, and its NAT wants a port-shaped
selector that ESP, AH and GRE do not have. The retracted rationale is corrected
where it was written down, not quietly dropped.

So these protocols go to the kernel instead. untunnelable_egress names an
interface or tunnel egress; prerouting stamps that egress's OWN mark on
everything that is not TCP or UDP, and addEgressRouting has already bound that
mark to a table whose default route leaves via the device. Every protocol works
because nothing in the path has to understand any of them. No new mark, no new
table, no new code in the hot path.

Whether that is a tunnel depends on the device, and nothing here claims
otherwise: a WireGuard interface is one, a second WAN is a different uplink
whose real address the far end sees.

The wide `!= { tcp, udp }` filter is safe here and stays banned for the L3
ingress, for the same reason stated in both places: there the receiver is a
dispatcher that knows four protocols, here it is the kernel. ICMP is claimed by
the L3 ingress first when both are on. The local plane keeps its exclusions —
router-addressed traffic, private destinations, ICMPv6 ND/RA — and with IPv6 off
the marking is scoped to v4, because addEgressRouting installs no v6 rule then
and a marked v6 packet would fall into the main table.

An interface egress with an empty `interface` no longer resolves: IfaceDevice
defaults to br-lan, so it passed the binding while addEgressRouting skipped it —
mark set, no rule, straight past a closed kill switch and out the default WAN.
2026-07-26 18:31:59 +03:00
omar 668cccbf24 test(generate): the L3 device name is a singleton, so wait for the kernel to take it back
Both gated tests stand an engine up on shater-l3. Run together, the second met
`TUNSETIFF: device or resource busy` and failed for a reason that had nothing to
do with what it asserts — the first had closed its box and yielded while
unregister_netdevice was still catching up. Each passed alone, which is the
shape of a fixture bug that gets rediscovered rather than fixed.

The poll that already guarded the first test is now a shared helper both call.
It stays a poll rather than a sleep for the reason it always was: the removal is
usually immediate and a fixed wait would be either flaky or slow.
2026-07-26 18:31:59 +03:00
omar 4ea4585402 test(generate): pin that ping through an interface egress is real, and byedpi's is not
An interface egress is a direct outbound carrying BindInterface and a routing
mark, and direct builds its ICMP port from the very same dialer control — so
ping routed at that egress leaves through that device, marked, like every other
packet bound to it. Nothing said so. Both halves of that sentence are one
`common.Cast[*dialer.DefaultDialer]` away from being false: if the dialer ever
stops being a DefaultDialer, icmpPort is nil, PreMatchFlow declines, and ping
through the egress degrades to a drop without a single generated byte changing.
The gated test asserts the live outbound, not the config, because that is where
the cast happens.

The failure the codegen half guards is worse than a broken ping: losing
BindInterface or the mark does not stop the echo, it sends it out the main table
over the plain WAN with the real address, which is the one thing an egress
exists to prevent.

byedpi is a SOCKS outbound and cannot be a tun.Port, so ICMP aimed at it is
dropped. That is the honest end of l3-honest-drop and it is pinned too, because
the alternative the TUN stack offers is a forged reply.
2026-07-26 18:31:59 +03:00
omar dc6d102473 docs: put a number on the second netstack, and say what it does not bound
Measured on a throwaway harness in a container: peak RSS of a process that
brought the engine up went from ~26 MB to ~28 MB with l3_tunnel on, three
paired runs. It is x86_64, idle, with an empty ICMP NAT table, so it stays
listed as unverified for the router — an indicative figure is more useful than
silence only if it says loudly what it is not.
2026-07-26 18:31:58 +03:00
omar 683afc0a47 docs: record how ping got through the tunnel, and where it stops
D25 writes down the reasoning that is expensive to reconstruct: why a TUN rather
than TPROXY, why the interface is its own with auto_route off, why gvisor is
mandatory rather than preferred, and why the ceiling is ICMP echo — a boundary
in sing-tun's flow parser and gVisor's protocol set, not an unfinished edge of
ours. It also records what carries layer 3 and what does not, that masque could
and does not, and the two things still unproven: the live-router path end to
end, and what a second gVisor NIC costs in memory on the hardware.

D17 gains one line: its claim that TPROXY cannot carry ICMP is still true, and
is no longer the end of the story.
2026-07-26 18:31:58 +03:00
omar 2c3e20512e feat(openwrt): let fw4 know the L3 tunnel device before it exists
Both nft tables run and a drop in either one wins, so our forward accept for
shater-l3 decides nothing on its own: fw4 sees a device in no zone and drops the
forward, and the feature fails with exactly the symptom it was built to fix —
ping does not work, and nothing says why.

The zone names the device directly rather than a network. fw4 resolves a zone's
networks through netifd, and a proto-none interface for a device the daemon
creates is never up and contributes nothing, so list network would compile to an
empty device set. list device compiles to a plain iifname/oifname match that is
valid before the TUN exists and starts matching the moment shaterd creates it,
with no firewall reload at enable time.

It is seeded unconditionally, not gated on l3_tunnel: uci-defaults run once, and
a zone naming an absent device is inert. Gating it would mean the option could
be switched on and never take effect. The sections are named so a re-run is a
no-op instead of a second zone, and kmod-tun joins DEPENDS because /dev/net/tun
is not on a stock image.
2026-07-26 18:31:58 +03:00
omar 51b2f04672 feat(netplane,generate): carry LAN ping through the tunnel, on a TUN of its own
Kernel TPROXY needs a socket to hand a packet to, so it moves TCP and UDP and
nothing else. Everything else reached the forward chain and met the untunnelable
policy, whose best answer was "let it out with your real address" and whose
default was "drop it" — so on a stock install ping simply did not work, and the
setting that fixed it did so by leaking.

The engine has been able to do better for a while: sing-tun's ForwardDispatcher
does real ICMP forwarding with NAT on the echo id, and a WireGuard or AmneziaWG
endpoint is a tun.Port that carries the packet for real. What was missing was a
way in, because nothing on the router could hand it an IP packet.

l3_tunnel (opt-in, off by default) adds one: the generator emits an "l3-in" TUN
inbound and prerouting fwmarks LAN ICMP into it. The interface is its own and
auto_route is off, so the main routing table is never touched and the fwmark
plus addL3Routing's ip rule are the only entrance — the TPROXY plane is byte for
byte what it was. gvisor is not a preference: the system stack forges echo
replies locally, which is the very thing this is meant to end.

Only icmp and ipv6-icmp are ever marked, and only after the local plane is out
of the way — the router itself, private destinations, and ICMPv6 ND/RA, which
mean nothing off-link and take v6 down if one neighbour probe is tunnelled.
ESP, AH, GRE, IGMP and SCTP are deliberately left alone: sing-tun's parser and
gVisor's stack know no such protocol, so marking them would black-hole the
traffic while looking like a feature. They stay with the untunnelable policy,
which also keeps its say over what happens if the ip rule fails to install.

Ping and Windows tracert now cross the tunnel; IPv6 traceroute shows only the
destination, because the return path recognises TimeExceeded for v4 alone.
2026-07-26 18:31:58 +03:00
omar f190c8251e feat(lx): stop answering ping on behalf of a tunnel that never saw it
PreMatchContinue is not "fall back to the ordinary route" the way it is for TCP
and UDP. An ICMP flow has no ordinary route: the TUN stack takes the packet back
and answers the echo itself, swapping the addresses and writing a reply
(sing-tun stack_gvisor_icmp.go). So a ping routed to any outbound that cannot
carry layer 3 — every proxy protocol; only adapter.FlowOutbound can — came back
successful, and the operator read a working tunnel off a packet that was never
sent.

That is worse than the packet loss it replaced. Loss is a fault the operator can
see and chase; a forged reply is a fault that reports itself as health, and it
reports it on the one tool anyone reaches for first.

preMatchFlow now overrides continueResult once, at the top, for
N.NetworkICMP. One hunk covers every exit that used to fall through — no such
outbound, a group whose selection is gone, an outbound whose Network() omits
icmp, an outbound that is not a FlowOutbound — and keeps the diff to three lines
against a function upstream will keep editing. JudgeFlow carries the same
verdict in its !isPort branch, because FlowOutbound and tun.Port are separate
interfaces and drift between them must not reopen the forgery.

TCP and UDP are untouched, and the test pins that as hard as it pins the drop.
2026-07-26 18:31:58 +03:00
omarandClaude Opus 5 1945404eaa fix(armor): a reboot is not someone switching the product off
test / go + panel tests (push) Successful in 5m24s
release / test gate (push) Successful in 5m24s
release / apk aarch64_cortex-a53 (push) Successful in 3m9s
release / apk x86_64 (push) Successful in 3m9s
release / release apk (push) Successful in 8s
The boot armor never armed on the router it shipped to. procd runs the
K-links on the way down with the action `shutdown`, and stop_service
classified actions with an OPEN default:

    case $action in restart|reload) keep;; *) DISARM;; esac

`shutdown` matched nobody, fell into `*`, and deleted the arm token. The
mechanism erased itself at exactly the transition it exists for, so every
boot found nothing to load. Measured on the live router, one minute apart
across a reboot:

    13:28  /etc/shater/boot.nft present
    ----   reboot
    18s    at_S22: NO_TABLE  armor_file=NO_FILE

It did not fail every time, which is worse than failing always: on the way
down `rm` from this script raced a `SaveBootArmor` driven by the ifdown
hotplug storm, and whichever landed second won. Two reboots on the same box
an hour apart gave opposite outcomes.

Both lists are now positive and CLOSED. Only `stop` disarms; only
`restart`/`reload` hand off. An action nobody thought of changes nothing,
so the default now fails toward a boot that arms when it need not have --
recoverable in the second before the daemon applies, and still gated by
shater-armor's four state refusals. The old default failed toward the
plaintext window the feature was built to close.

Also closed, found while proving the above:

  * Every restart left the LAN in the clear for 80-90ms. The exit path was
    `Teardown(); armOnExit()`, and TeardownNft DELETES the table -- two nft
    transactions with no `inet shater` between them, leaving fw4's
    `lan -> wan ACCEPT` as the only policy. Every restart, every LuCI Save
    & Apply. TeardownExiting arms first under the apply lock and skips the
    delete iff a plane actually went in; RenderHoldNft is one `nft -f` that
    REPLACES the table, so the kernel never observes its absence.
    35k-sample instrument: 7 and 6 no-table hits before, 0 across three
    runs after.

  * SaveBootArmor fsynced the payload but not the directory, so a power cut
    could lose the rename that publishes it -- a boot with no armor and no
    error anywhere.

`stop` now also reads rc.d state, so a package transaction that stops the
service is not mistaken for a person switching it off. This one does not
reproduce on apk (it runs no pre-upgrade script and never calls prerm on an
upgrade; verified with apk adbdump and 245k samples across a real reinstall)
-- it is one returning opkg lane away from being live, and the removal case
is now stated rather than implicit.

Both new tests are mutation-checked: reverting the predicate fails naming
`shutdown`; reverting the teardown fails with `did [arm delete], want [arm]`.
initscript_test.go sources the SHIPPED shell and calls the real predicates
with every action procd uses -- a comment claiming `shutdown` was handled is
what shipped last time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 17:33:53 +03:00
omarandClaude Opus 5 6476722372 fix(panel): stop shipping a fabricated router in the binary
test / go + panel tests (push) Successful in 5m26s
release / test gate (push) Successful in 5m28s
release / apk aarch64_cortex-a53 (push) Successful in 6m7s
release / apk x86_64 (push) Successful in 3m5s
release / release apk (push) Successful in 7s
mock.ts was a static import and the mock switch was read from the query string at
runtime, so the bundle that ships inside the daemon carried a complete fictional
router and a link ending in ?dev rendered it: protected, 119 of 122 nodes alive,
without a single request to the daemon. The only tell was a line in the footer.
That is worse than any wrong number — there is no data at all and nothing says
so. It is out of the production bundle now, which is 21 kB smaller for it.

Unknown state stopped reading as good news in two more places. The kill-switch
tile treated an absent plane as armed, because the check was "not none" and
undefined satisfies it — the contract in the API types says the opposite. And the
apply page announced "daemon auto-rolled back" from its own timer, while the
daemon, seeing the state generation move, disarms and says it is NOT rolling back
in the log only.

Alerts moved to Settings. They are about the kill switch, apply failures, new
devices and subscription expiry, and they lived at the bottom of the DNS page,
while Settings mentioned them in prose with nothing to click.

Findings truncation is visible now: the notice that says how many were suppressed
arrives as info, and the attention list keeps only critical and warning, so past
fifty findings the operator saw forty-nine and no hint of the rest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:39 +03:00
omarandClaude Opus 5 4078334d85 fix(stats,alert,panel): put a ceiling on everything that only grew
Four maps had no bound on a box with 512 MB that runs for months. The health
board only ever inserted — the delete exists but no path in this fork calls it —
and it lives on the engine context, so it outlives every generation. Its keys are
node tags, and providers rename nodes on each subscription refresh: about 440k
keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The
stats aggregator's server and outbound counters were the only ones with no cap,
no prune and no top-N, and one of them was handed to the panel whole on every
poll.

They are bounded now, evicting least-recently-seen, with numbers argued from this
box rather than round: the board holds 4096 against a live generation of about
1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped
silently — the same rule the log sink already follows — and a new Dropped section
in the snapshot reports all six bounded aggregates, including the three that had
been evicting without saying so.

Snapshot did O(devices × domains) under the aggregator lock, sorting five
thousand entries to show fifteen, and could read the DHCP lease file from inside
it. Meanwhile the event subscribers have 64-slot buffers that drop without a
counter, so an open Overview page cost the query log real rows. Selection is
top-K now — proven byte-identical to the old sort over 200 random trials — and
both the lease read and the row ordering happen outside the lock.

The panel server had one timeout, on headers. An unauthenticated client could
hold a goroutine, a socket and a descriptor forever by sending its body one byte
at a time; a stopped reader on the log stream held the handler, the pipe and a
child process that outlived the request. Every phase is bounded now, with the
unauthenticated route on a tighter budget than the rest, and the log stream
renewing its deadline per chunk so a slow-but-reading client is never truncated.

And the last of the detour transports: each call built a fresh one, and the alert
delivery path dropped it, pinning keep-alive sessions through the engine's own
outbounds for 90 seconds — eighteen times the budget a retiring generation gets.

The race skip is gone from the gate. The test it existed for raced in its own
clock, not in the product; that is fixed, so nothing is excluded under -race any
more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:21 +03:00
omarandClaude Opus 5 0a6689b29e fix(quic,v2ray): close the sockets quic-go was never going to close
DialEarly with a packet conn the caller made sets a flag that means quic-go does
not own it: closing the transport only stops reading from the socket. Neither DNS
transport closed it. On the QUIC one it was closed on a failed handshake and
never on success, so every redial — idle timeout, retry error, engine reload —
left a UDP socket for the life of the process. On the HTTP/3 one the library
drives its own reconnects, so the leak compounds without anything in our code
looking wrong.

That is the same shape as v2rayquic's, where offerNew overwrote the raw conn on
every reconnect without closing the previous one. Both are now owned by a watcher
tied to the connection's own context, so the socket lives exactly as long as the
connection does.

This matters more than it did last week: the shipped resolvers are DoH, and DNS
is intercepted by default now, so the whole network's query stream rides this
path on a router with 512 MB.

The same upstream commit fixes both halves. We had taken the v2ray half and not
the DNS one — the third time this session a paired fix arrived half-applied, and
the first of those cost a day of debugging. These two files are now byte-identical
to upstream so a rebase cannot reopen it.

Also from that family: websocket and httpupgrade leaked their conn on failed
handshakes, and a QUIC stream's Close did not release a blocked write.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:59 +03:00
omarandClaude Opus 5 ef22167b1a fix(apply): report the hold when the plane was armed by someone else
Booting with the armor loaded, or restarting through the handoff, left the status
saying the LAN was not being held while it was being dropped. Transient after a
successful apply, but permanent on the unreadable-config path — and there the
apply-failure alert words itself "traffic is NOT being blocked" at the exact
moment it is. That sends the operator to fix something that is not broken, past
the protection that is holding.

The table cannot be identified from here — netplane exposes no read-back and nft
does not keep comments — but identifying it is the wrong question. Holding does
not claim the holding plane is the object in the kernel; it claims the engine is
down and forwarded traffic is being dropped. A leftover full ruleset does that
too: with no engine socket the tproxy statement breaks its own rule before the
accept, so the packet reaches the forward chain unmarked and meets the primary
drop. What decides it is whether the last applied config was enabled and
fail-closed, which is exactly what the boot armor's presence already means.

So it is derived at read time rather than latched. A latch set from an inference
would have to be remembered in order to be cleared, which is the trap the active
flag already taught us. ArmHold also stops deferring to a table it cannot
inspect and installs its own render instead — the honest answer to "do not claim
a foreign table blindly" is to make it ours, and a fresh render beats a snapshot
that predates an interface rename.

Also closes the last of the detour transports: the subscription fetch took a
client and dropped it, and the exits that leak are the error ones, retried by
cron forever against a broken feed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:39 +03:00
omarandClaude Opus 5 cbda0fee0a fix(netplane): arm the fail-closed plane before the daemon can
The plane only ever existed while the daemon did. It starts at 99, after fw4 has
already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its
predecessor, migrating the schema, building the engine and reading UCI — with a
UPX-compressed binary decompressing off flash first. Every boot therefore had a
window with no protection at all, landing exactly when Wi-Fi comes up and every
client reconnects. A restart, a reload or a package upgrade opened the same
window on purpose: Teardown does not consult the kill switch, and the init script
guarantees the interval is non-empty.

The holding plane is now persisted to /etc/shater/boot.nft on every apply and
loaded by a small service at 21, right after fw4 and netifd. Its presence is the
arm token: it exists only while the last applied config was enabled AND
fail-closed, and goes away the moment either stops being true. Writes are
content-gated — the cron reconcile runs a minute — and atomic, because the one
boot that reads this file is the boot after a power cut.

The service refuses to arm four ways so it can never brick a box, and its
enabled-check reads /etc/rc.d directly rather than asking rc.common, which would
take a blocking flock in the middle of boot. On exit the daemon re-arms only for
restart and reload, read from a snapshot of rc.common's action; anything else,
including an unknown one, degrades to a real stop that also disarms.

An unreadable config used to leave the router bare forever: the arm call sat in
the branch that requires a successful read, and nothing downstream could recover
it. It now arms from the same path.

A network nobody named was neither diverted nor blocked — the divert set is built
from inbounds and rule sources, and the same set scopes the fail-closed drops. It
is now enumerated from the interfaces whose firewall zone the operator forwards
to a WAN zone — their own statement that those clients reach the internet through
this box — and reported critically, by name, with both resolutions. Deliberately
not closed automatically: this router cannot know a guest SSID was meant to be
off the tunnel, and guessing is an outage. A device name that resolved to nothing
is reported the same way, for the same reason: there is no fail-closed action
available for a device we cannot name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:16 +03:00
309 changed files with 64939 additions and 3986 deletions
+14 -12
View File
@@ -11,13 +11,12 @@
# the arch-matched artifact. (arch-specific .apk)
# - shater-core data glue, PKGARCH=all
# - luci-app-shater LuCI thin launcher, PKGARCH=all (uses feeds/luci/luci.mk)
# - byedpi ciadpi, C cross-compiled from source by the SDK (arch-specific)
#
# TARGET HARDWARE / ARCH MATRIX
# x86_64 -> the QEMU testbed VM (generic x86-64).
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 mini + BPI-R4,
# mediatek/filogic), both on 25.12 with apk-tools 3.
# Only shaterd + byedpi are arch-specific; shater-core + luci-app-shater are
# Only shaterd is arch-specific; shater-core + luci-app-shater are
# PKGARCH=all, so one build of each covers every device — but the RELEASES
# are still per-arch (see the release-apk job for why).
#
@@ -35,11 +34,15 @@
# invalidates every deployed router's trust.
#
# AUTO-RELEASE
# push a tag `vX.Y.Z` -> versioned per-arch releases `apk-vX.Y.Z-<arch>`.
# workflow_dispatch -> rolling per-arch `apk-latest-<arch>` (always-fresh
# feed). Publish uses the Gitea API via curl (ci/gitea-release.sh) — no
# external action needed. NOTE: the apk release tags deliberately do NOT start
# with `v` so publishing them cannot re-trigger this workflow's `v*` filter.
# The rolling per-arch `apk-latest-<arch>` is published on EVERY run — tag runs
# included — and then read back over the API to assert it really serves the
# version just built. A tag push `vX.Y.Z` publishes the pinnable per-arch
# `apk-vX.Y.Z-<arch>` IN ADDITION. It is not an either/or: it used to be, and
# the rolling pointer then froze at 0.2.0 while v0.2.9/v0.2.10 shipped (see the
# long comment above the `release-apk` job). Publish uses the Gitea API via curl
# (ci/gitea-release.sh) — no external action needed. NOTE: the apk release tags
# deliberately do NOT start with `v` so publishing them cannot re-trigger this
# workflow's `v*` filter.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
@@ -55,8 +58,8 @@
# into the binary's constant.Version. ci/sdk-build-apk.sh then ASSERTS that the
# built .apk really carry that version, so the failure can never be silent
# again. This is also why the build job checks out with fetch-depth: 0
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
# — `git describe` needs tags and ancestry. Every package this repo ships is
# versioned from the tag; there is no longer an exception to remember.
# CACHING (T3 — fast CI)
# All caches use actions/cache pinned to v3.3.2: the LAST release speaking the
@@ -449,7 +452,7 @@ jobs:
echo "[release-apk] arch=$arch built version=$want"
BODY="Automated apk (OpenWrt/ImmortalWrt 25.12+) package repo for \`$arch\`.
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
Packages: shaterd (per-arch), shater-core + luci-app-shater (arch=all).
This build: \`$want\`.
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
@@ -458,7 +461,6 @@ jobs:
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
apk update
apk add luci-app-shater # pulls shater-core + shaterd too
apk add byedpi # optional: ByeDPI desync egress
\`apk-latest-<arch>\` is a MOVING pointer: every release run replaces its
assets, so the same repo line keeps serving the newest build. To pin a
version instead, point the repo line at
@@ -466,7 +468,7 @@ jobs:
file must be edited by hand for each upgrade.
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
apk upgrade shaterd shater-core luci-app-shater
A bare \`apk upgrade\` reconciles EVERY installed package against every
configured repo and can downgrade unrelated system packages; naming them
upgrades only those (apk-tools 3: \"If list of packages is provided, only
+141 -47
View File
@@ -6,61 +6,155 @@
## Правила делегирования
1. ЛЮБАЯ реализация (код, тесты, конфиги, рефакторинг, отладка) выполняется
субагентами через инструмент Agent с `model: "fable"`. Сам ты правишь файлы
только в одном случае: тривиальная правка в 1–2 строки, где постановка
задачи дороже самой правки.
субагентами через инструмент Agent. Сам ты правишь файлы только в одном
случае: тривиальная правка в 1–2 строки, где постановка задачи дороже самой
правки.
2. Перед делегированием ты сам исследуешь код настолько, чтобы написать
точное ТЗ. В каждом задании субагенту обязательно указывай:
- контекст: что это за проект и над чем идёт работа;
- конкретные файлы и функции, которые нужно менять (пути, а не «найди сам»);
2. **Модель выбирает исполнитель задачи, а не привычка.** `fable` — быстрый и
дешёвый, годится для механической работы с ясным контрактом. `opus` — для
всего, где нужно рассуждение: поиск причины, аудит, дизайн, работа в чужом
коде. Если у `fable` кончилась квота — молча переходи на `opus`, это не повод
останавливать работу. Не спрашивай владельца, какую модель брать.
3. Перед делегированием ты сам исследуешь код настолько, чтобы написать точное
ТЗ. В каждом задании субагенту обязательно указывай:
- контекст: что за проект и над чем идёт работа;
- конкретные файлы и функции (пути, а не «найди сам»);
- контракт: сигнатуры, форматы данных, инварианты, что менять НЕЛЬЗЯ;
- definition of done: как проверить, что задача выполнена
(какие команды/тесты прогнать и какой ожидается результат);
- что вернуть в финальном ответе: список изменённых файлов, результаты
проверок, найденные проблемы и принятые решения.
- definition of done: какие команды прогнать и какой ждать результат;
- что вернуть: изменённые файлы, результаты проверок, найденные проблемы,
принятые решения.
3. Скиллы: при постановке задачи посмотри список доступных скиллов и ЯВНО
перечисли в ТЗ, какие скиллы субагент обязан вызвать через инструмент Skill
до начала работы (например: «сначала вызови Skill "openwrt-procd-services"
и следуй ему»). Субагент не видит наш диалог и сам не догадается — пиши
названия скиллов прямо в текст задания.
4. **Скиллы использовать по максимуму — и тебе, и агентам.** Это не
формальность: в них лежит выстраданное знание по ровно тем предметным
областям, в которых мы работаем, и игнорировать их — значит переоткрывать
чужие грабли. См. раздел «Скиллы» ниже.
4. Независимые задачи запускай ПАРАЛЛЕЛЬНО — несколько вызовов Agent в одном
сообщении, каждый с `model: "fable"`. Зависимые — последовательно, передавая
в следующее ТЗ результаты предыдущего.
5. Независимые задачи запускай ПАРАЛЛЕЛЬНО — несколько вызовов Agent в одном
сообщении. Зависимые — последовательно, передавая результаты предыдущего.
**Делишь файлы между параллельными агентами явно** и пишешь каждому, кто ещё
работает в дереве и что трогать нельзя. Запрещай им `git stash`,
`git checkout <файл>`, `git reset` — в этом проекте агент уже сносил правки
соседа через `git stash push`.
5. Приёмка: результат каждого субагента ты проверяешь сам (читаешь diff
ключевых мест, гоняешь проверки из definition of done). Если результат
не принят — не переделывай сам, а верни задачу: доработку заказывай тому же
агенту через SendMessage (у него сохранён контекст), а не новым спавном.
6. Приёмка: результат каждого субагента ты проверяешь сам — читаешь diff
ключевых мест, гоняешь проверки из definition of done. Не принимай отчёт на
слово: сегодня отчёт «тесты зелёные» дважды сопровождался тестом, который
ничего не прибивал. Если результат не принят — не переделывай сам, а верни
задачу тому же агенту через SendMessage (у него сохранён контекст).
6. Финальный отчёт пользователю: что сделано, кем (сколько агентов),
что проверено, что осталось.
7. Финальный отчёт владельцу: что сделано, сколько агентов, что проверено,
**что осталось непроверенным и почему** — последнее так же важно.
## Инженерные стандарты
Это не пожелания. Каждый пункт здесь появился после того, как его отсутствие
стоило рабочего дня.
- **Тест обязан быть проверен мутацией.** Откатить фикс → показать, что тест
падает, и с каким текстом → вернуть фикс. Тест, не падающий на сломанном коде,
не тест, а украшение.
- **Прибор без контроля не доказывает ничего.** Отрицательный результат чего-то
стоит, только если показано, что этот же прибор умеет дать положительный.
«Утечки не нашли» прибором, который не мог её увидеть, — это не результат.
- **Опровержение ценнее согласия.** В каждом ТЗ прямо разрешай субагенту
сказать «твоя версия неверна» и требуй доказательства, а не вежливости.
Лучшие результаты этого проекта приходили именно так.
- **Не обещать непроверенного.** Комментарий, предупреждение и текст в панели —
это утверждения о поведении. Если поведение не проверено, так и писать.
Формально верная фраза, которая читается как «работает», — тоже ложь.
- **Умолчание падает в восстановимую сторону.** Открытый `default:` в разборе
вариантов — источник целого класса дефектов: неучтённое значение уходит туда,
где дороже всего ошибиться. Списки делать положительными и закрытыми.
- **Проверка присутствия обязана покрывать всё, что ставит её Apply-двойник.**
Иначе идемпотентный быстрый путь становится ловушкой: «всё на месте» при
отсутствующем маршруте.
- **Никакого молчаливого скипа.** Тест, который не выполнился, обязан быть
назван поимённо в выводе гейта. Однажды CI гонял два теста из 116 файлов, и
все считали, что покрыто.
## Скиллы
**Правило: если задача касается области, по которой есть скилл, — скилл
вызывается ДО начала работы, а не после того, как что-то не заработало.**
Это относится и к тебе, и к каждому субагенту.
Субагент не видит наш диалог и сам не догадается, что скиллы существуют.
Поэтому **в каждом ТЗ перечисляй поимённо**, какие скиллы он обязан вызвать
через инструмент Skill: «сначала вызови Skill "openwrt-nftables" и Skill
"openwrt-networking", следуй им». Требуй в отчёте сказать, что именно из скилла
он применил, — так видно, вызвал он его или упомянул.
Соответствие областей этого проекта и скиллов:
| Трогаешь | Обязательные скиллы |
|---|---|
| `/etc/config/*`, `uci`, uci-defaults, парсер модели | `openwrt-uci` |
| nftables, fw4, зоны, метки, tproxy, kill-switch | `openwrt-nftables` |
| интерфейсы, мосты, VLAN, policy routing, `ip rule`, sysctl, dnsmasq | `openwrt-networking` |
| init-скрипты, procd, respawn, service triggers, boot armor | `openwrt-procd-services` |
| перехват трафика целиком (tproxy + маршрутизация + DNS) | `openwrt-transparent-proxy` |
| сборка пакетов, SDK, фид, CI, подпись, `apk`/`opkg` | `openwrt-package-build-ci`, `openwrt-native-packages` |
| LuCI-приложение, ubus/rpcd, ucode | `openwrt-luci-plugin`, `openwrt-ubus-rpcd`, `openwrt-ucode` |
| панель (React/TS) | `react-expert`, `frontend-design:frontend-design` |
| Go: конкурентность, каналы, профилирование, идиоматика | `fullstack-dev-skills:golang-pro` |
| TypeScript | `fullstack-dev-skills:typescript-pro` |
| стратегия тестирования, покрытие, тестовые данные | `fullstack-dev-skills:test-master` |
| поиск причины по логам и трассам | `fullstack-dev-skills:debugging-wizard` |
| проверка в браузере, скриншоты | `fullstack-dev-skills:playwright-expert` |
| ревью | `review`, `fullstack-dev-skills:code-reviewer` |
| безопасность | `security-review`, `fullstack-dev-skills:security-reviewer` |
| графики и визуализация данных | `dataviz` |
Список неполный — **смотри доступные скиллы под задачу**, а не только в эту
таблицу. Если скилл выглядит смежным, дешевле вызвать его и не воспользоваться,
чем не вызвать и потом отлаживать то, что там уже описано.
## Проверки
- **Гейт:** `bash scripts/run-tests.sh` — Linux в Docker, боевой набор тегов,
`-race`, и шаг, требующий вердикта по имени для привилегированных тестов.
Зелёный гейт — необходимое условие, но не достаточное: он не видит стыков с
ядром, procd и nftables.
- **Стенд:** сервер `local_openwrt` в ssh-manager — ImmortalWrt 25.12.1 той же
ревизии, что боевой роутер. Сюда — всё, что касается init-скриптов, nft,
policy routing, TUN.
- **Боевой роутер:** `mini_router` (BPI-R3), через него идёт весь домашний
трафик. Перед изменением конфигурации — резервная копия. Проверять приборно,
а не по логу: лог может печатать одно и то же в честном и в ложном случае.
## Релиз и деплой
- Тег → CI (Gitea Actions) → apk-фид → установка на роутер.
- **Обновлять только поимённо**, никогда не `apk upgrade` целиком:
`apk upgrade shaterd shater-core luci-app-shater`.
- **Не трогать кеш CI-раннера** — сборка растянется на часы.
- Число тегов на порцию работы — на твоё усмотрение, если владелец не сказал
иначе.
## Фронтенд (admin panel)
Дизайн-направление ЗАФИКСИРОВАНО: **Faceplate** (панель сетевого железа).
Полная спека, токены, компоненты и ссылка на живой эталон — в
[`docs-shater/DESIGN.md`](docs-shater/DESIGN.md). Эталон:
https://claude.ai/code/artifact/9f7c07e8-d8ac-4ae1-b113-5b25d0ba5dd2
Спека, токены и компоненты — в [`docs-shater/DESIGN.md`](docs-shater/DESIGN.md).
Эталон: https://claude.ai/code/artifact/9f7c07e8-d8ac-4ae1-b113-5b25d0ba5dd2
- **Стек:** Vite + React + TypeScript, лёгкий (SPA встраивается в бинарь —
без тяжёлых зависимостей). Расположение: папка `panel/` в корне.
- **Порядок работ:**
1. Сам (оркестратор) скаффолдишь `panel/`, переносишь токены из
`docs-shater/DESIGN.md` в `panel/src/tokens.css` один-в-один и задаёшь каркас
компонентов. Это фундамент — делай аккуратно сам или отдай ОДНОМУ агенту.
2. Дизайн-систему в компоненты: `<Faceplate> <Module> <Toggle> <Led>
<SegMeter> <QueryLog>` + кнопки — строго по эталону.
3. Страницы раздаёшь ПАРАЛЛЕЛЬНО Opus-агентам (`model: "opus"`), по одной на
агента: Overview, Nodes/Subscriptions, Routing rules, DNS/Blocklists,
Devices, Apply/Rollback.
- **В КАЖДОМ ТЗ агенту обязательно:** ссылка на `docs-shater/DESIGN.md` и на эталон;
требование сначала вызвать Skill `react-expert` и Skill
`frontend-design:frontend-design` и следовать им; список готовых компонентов,
которые он ДОЛЖЕН переиспользовать (не изобретать заново); какие токены и
семантические цвета применять; DoD — страница совпадает с языком эталона,
адаптив + фокус + reduced-motion соблюдены.
- **Не отходить от Faceplate.** Любой новый экран наследует ту же визуальную
систему. Оранжевый — только акцент; семантика good/warn/crit — отдельно.
- **Стек:** Vite + React + TypeScript в `panel/`. SPA встраивается в бинарь —
тяжёлые зависимости недопустимы.
- **Панель целиком на английском.** Ни одного символа кириллицы в `panel/src`.
- **В КАЖДОМ ТЗ на панель:** ссылка на `DESIGN.md` и на эталон; требование
сначала вызвать Skill `react-expert` и Skill
`frontend-design:frontend-design`; список существующих компонентов, которые
надо ПЕРЕИСПОЛЬЗОВАТЬ (`<Faceplate> <Module> <Toggle> <Led> <SegMeter>
<QueryLog>` и кнопки), а не изобретать заново; какие токены и семантические
цвета применять; DoD — совпадение с языком эталона, адаптив, фокус,
`prefers-reduced-motion`.
- Оранжевый — только акцент; семантика good/warn/crit — отдельно.
- **Панель не должна врать про состояние.** Значение, которое движок примет,
не может рисоваться как «never matches»; настройка, которой управляет другая
подсистема, не может описываться так, будто управляет ею.
+32 -7
View File
@@ -24,7 +24,8 @@ The engine is a **fork of [sing-box](https://github.com/SagerNet/sing-box) via
[sing-box-lx](https://github.com/Leadaxe/sing-box-lx)**, compiled into a single Go
binary `shaterd` together with the control plane, DNS filter, stats aggregator and
the web panel itself. Broad protocol set: VLESS/VMess/Trojan/Shadowsocks,
Reality/XTLS, WireGuard, **AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP, MASQUE/CONNECT-IP.
Reality/XTLS, WireGuard, **AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP — exactly what
`shater/parse` can read and `shater/registry` registers in the engine.
A thin **LuCI launcher** (mini-dashboard + "Open panel" button) hands the browser a
single-use token into the standalone SPA the daemon serves on its own port
@@ -39,8 +40,8 @@ single-use token into the standalone SPA the daemon serves on its own port
selector / chain / direct / block; node groups with balancer/observatory;
multi-hop chains; per-rule egress.
- **Fail-closed kill-switch** (dead group → block, never a silent direct leak); own
`inet shater` nft table; atomic apply with `nft -c` validation and commit-confirm
auto-rollback.
`inet shater` nft table; atomic apply with `nft -c` validation. Commit-confirm
auto-rollback exists but **ships OFF** (`confirm_timeout=0`) — arm it yourself.
- **DNS filtering & blocklists** with flexible sources (inline / file / url /
geosite), compiled `.srs` matcher; Block-DoH/DoT to stop filter bypass.
- Subscriptions (Clash / sing-box / Xray-JSON) and manual nodes; node health board.
@@ -63,13 +64,28 @@ apk update && apk add luci-app-shater # -> shater-core -> shaterd
```
`apk-latest-<arch>` is a moving pointer refreshed by every release run — install
once and `apk update && apk upgrade shaterd shater-core luci-app-shater byedpi`
once and `apk update && apk upgrade shaterd shater-core luci-app-shater`
keeps the router current. Point the repo line at `apk-vX.Y.Z-<arch>` instead to
pin a build; that file then has to be edited by hand for every upgrade.
shater ships **inert** (globals off) so install never breaks connectivity. After
configuring nodes/rules: `uci set shater.globals.enabled=1 && uci commit shater`,
then `shaterd apply` and `shaterd confirm`.
configuring nodes/rules:
```sh
uci set shater.globals.enabled=1
uci set shater.globals.confirm_timeout=120 # commit-confirm ships OFF — arm it
uci commit shater
shaterd apply && shaterd confirm
```
Without that middle line `shaterd apply` arms no auto-rollback (and says so), so an
apply that costs you SSH/LuCI access has to be undone by hand.
Once an enabled, fail-closed config has been applied, `/etc/init.d/shater-armor`
loads a saved fail-closed plane at **boot**, before the daemon exists: LAN→WAN
forwarding is blocked until `shaterd` applies, while SSH/LuCI/the panel stay
reachable on purpose (the chain hooks `forward` only). What arms it, what refuses
to arm, and how to switch it off — `INSTALL.md` §4.
## Build from source
@@ -78,13 +94,22 @@ then `shaterd apply` and `shaterd confirm`.
into `openwrt/shaterd/files/`. Details in
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
`bash scripts/run-tests.sh` is the test gate: the whole suite under the **shipped**
build tags (`scripts/router-tags.sh`), on linux (it re-execs in Docker from a
non-linux host), with `-race`, plus three machine checks against a silent skip —
the tag set may only add test files, every package with tests must report `ok` by
name, and every `TestIntegration*` must produce a verdict by name.
`scripts/check-router-tags.sh` separately proves no feature declared in
`FEATURES.md` lost a build tag it needs. A green gate is necessary but not
sufficient: it does not see the kernel, procd or nftables seams.
## Repository layout
| Path | What |
|------|------|
| `shater/` | Go control plane, DNS filter, stats aggregator, engine host |
| `panel/` | Admin SPA (Vite + React + TS) and its Go server |
| `openwrt/` | Packages: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `openwrt/` | Packages: `shaterd`, `shater-core`, `luci-app-shater` |
| `docs-shater/` | Product documentation |
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, apk feed/release scripts, CI |
| `SPECS/`, `docs-lx/` | Engine-fork constitution/specs and feature-config reference |
+62 -15
View File
@@ -27,7 +27,8 @@ BananaWRT** (Banana Pi BPI-R3, BPI-R4 и совместимые). Он проз
Go-бинарь `shaterd` вместе с control-plane, DNS-фильтром, агрегатором статистики и
самой веб-панелью. За счёт sing-box поддерживается широкий и актуальный набор
протоколов: VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, WireGuard,
**AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP, MASQUE/CONNECT-IP.
**AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP — ровно то, что умеет разобрать
`shater/parse` и что регистрирует `shater/registry` в движке.
Интеграция в OpenWrt — тонкий **LuCI-лаунчер**: мини-дашборд и кнопка «Открыть
панель», которая по одноразовому токену передаёт браузер в полноценную SPA-панель,
@@ -50,8 +51,10 @@ Go-бинарь `shaterd` вместе с control-plane, DNS-фильтром,
- **Fail-closed kill-switch**: мёртвая группа → block, а не тихая утечка мимо
прокси; собственная nft-таблица `inet shater` и свои марки/таблицы, fw4 не
трогаем.
- Атомарный apply с валидацией движком и `nft -c`, **commit-confirm** с
авто-откатом к последней рабочей конфигурации.
- Атомарный apply с валидацией движком и `nft -c`. **Commit-confirm** с
авто-откатом к последней рабочей конфигурации есть, но **на стоковой установке
выключен**: `confirm_timeout` поставляется нулём, и apply не вооружает ничего,
пока вы не зададите окно (см. «Включение»).
- Идемпотентный reconcile из hotplug/boot под flock; management-bypass
(SSH/LuCI/LAN) всегда в обход.
@@ -121,8 +124,11 @@ flowchart TB
Путь трафика: LAN-клиент → `nft tproxy` (mark → tproxy-порт) → tproxy-inbound
sing-box (сниффинг SNI/Host/QUIC) → маршрут по правилу → outbound/selector/chain
(проксировано) · direct (flow-offload) · block. Подробные диаграммы (auth-handoff,
data-plane, DNS-flow, apply-flow) — в [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md).
(проксировано) · direct (обычный маршрут, без туннеля) · block. TPROXY несёт
только TCP и UDP; ICMP и остальные протоколы — через отдельные опциональные
механизмы (`l3_tunnel`, `untunnelable_egress`, ARCHITECTURE §3a). Подробные
диаграммы (auth-handoff, data-plane, DNS-flow, apply-flow) — в
[`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md).
---
@@ -133,8 +139,8 @@ BananaWRT **25.12+**: `.apk`, индекс `packages.adb`, EC-ключ в `/etc/
Старый opkg-фид (`.ipk`, 24.10) снят — оба наших роутера на 25.12 с apk-tools 3,
бинаря `opkg` там просто нет (`docs-shater/DECISIONS.md` D22).
Пакеты ставятся по зависимостям: `shaterd` → `shater-core` → `luci-app-shater`
(+ опциональный `byedpi`). `shaterd` подтягивается автоматически как зависимость.
Пакеты ставятся по зависимостям: `shaterd` → `shater-core` → `luci-app-shater`.
`shaterd` подтягивается автоматически как зависимость.
### Фид apk
@@ -152,7 +158,6 @@ echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/a
# 3) обновляемся и ставим (shaterd подтянется как зависимость).
apk update
apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # опционально: ByeDPI desync-egress
```
Обновление — **перечисляйте пакеты явно, голый `apk upgrade` не запускайте**: без
@@ -162,14 +167,14 @@ apk add byedpi # опционально: ByeDPI desync-egress
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
apk upgrade shaterd shater-core luci-app-shater
# эквивалент, дополнительно закрепляющий пакеты в world:
# apk add -u shaterd shater-core luci-app-shater byedpi
# apk add -u shaterd shater-core luci-app-shater
```
Документация apk-tools 3 про `apk upgrade`: *«If list of packages is provided,
only those packages are upgraded along with needed dependencies»*. Проверить
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
установленные версии: `apk list -I shaterd shater-core luci-app-shater`.
> **Роллинг или фиксация — это выбор URL в `shater.list`.** `apk-latest-<arch>`
> — движущийся указатель: каждый релизный прогон заменяет его ассеты, поэтому
@@ -195,15 +200,32 @@ shater ставится **инертным** (globals выключены), чт
```sh
uci set shater.globals.enabled=1
# Предохранитель: commit-confirm поставляется ВЫКЛЮЧЕННЫМ (confirm_timeout=0),
# и без этой строки apply ничем не подстрахован. 120 с — окно на проверку связи.
uci set shater.globals.confirm_timeout=120
uci commit shater
shaterd apply # apply + вооружить commit-confirm на живом демоне
shaterd confirm # подтвердить (отменяет авто-откат)
shaterd apply # применить и вооружить авто-откат на 120 с
shaterd confirm # подтвердить в пределах окна (отменяет авто-откат)
```
`shaterd apply` печатает, вооружил ли он что-нибудь, и почему нет: при
`confirm_timeout=0` он прямо говорит, что автоматического отката НЕТ. Оставить
ноль — сознательный выбор: тогда apply, отрезавший вам SSH/LuCI, придётся
откатывать руками.
`/etc/init.d/shater enable && /etc/init.d/shater start` поднимает демона под procd.
Кнопка «Открыть панель» в LuCI чеканит одноразовый токен и передаёт браузер в
панель (`:8088` по умолчанию).
После первого же применённого включённого fail-closed конфига появляется
**загрузочная защита**: `/etc/init.d/shater-armor` (START=21) грузит сохранённый
fail-closed план ещё до старта демона, закрывая те секунды между поднятием LAN и
первым apply, когда роутер форвардил трафик в WAN открытым. Форвардинг LAN→WAN
заблокирован, пока `shaterd` не применит конфиг; SSH, LuCI и панель при этом
доступны **намеренно** — цепочка вешается только на `forward`. Чем защита
вооружается, когда отказывается вооружаться и как её снять —
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) §4.
---
## Сборка из исходников
@@ -226,6 +248,28 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
(набор build-тегов, почему `shaterd` — prebuilt-пакет, порядок CI) — в
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
### Проверка
```sh
bash scripts/run-tests.sh # полный гейт
bash scripts/run-tests.sh --no-race # без -race, для локального цикла
```
Гейт гоняет весь набор **под теми же build-тегами, с которыми собирается
роутерный бинарь** (`scripts/router-tags.sh`), на Linux (с не-Linux хоста — сам
перезапускается в Docker), с `-race`, и содержит три машинные проверки против
молчаливого скипа: набор тегов может только ДОБАВЛЯТЬ тест-файлы; каждый пакет с
тестами обязан отчитаться `ok` поимённо; каждый `TestIntegration*` обязан выдать
вердикт по имени. Причина такая: до 2026-07 релизный тракт не гонял почти ничего
— 115 тест-файлов из 116 под `shater/**` в CI не исполнялись ни разу.
Отдельно `scripts/check-router-tags.sh` проверяет, что ни одна заявленная в
`FEATURES.md` фича не потеряла нужный ей build-тег.
Зелёный гейт — необходимое, но не достаточное условие: он не видит стыков с
ядром, procd и nftables. Это проверяется на стенде (см.
[`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md)).
---
## Структура репозитория
@@ -237,7 +281,7 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
|------|---------|
| `shater/` | Go: control-plane, DNS-фильтр, агрегатор статистики, хост движка |
| `panel/` | Админ-SPA (Vite + React + TS) и её Go-сервер |
| `openwrt/` | Пакеты: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `openwrt/` | Пакеты: `shaterd`, `shater-core`, `luci-app-shater` |
| `docs-shater/` | Документация продукта (см. таблицу ниже) |
| `scripts/` | `build-shaterd.sh` — сборка ship-артефакта |
| `ci/` | Скрипты сборки apk-фида и релизов (SDK, EC-подпись, Gitea API) |
@@ -274,7 +318,10 @@ CI на **Gitea Actions** (`.gitea/workflows/release.yml`) собирает вс
shater вкомпилирует **форк движка sing-box-lx** — тонкий downstream апстрима
[SagerNet/sing-box](https://github.com/SagerNet/sing-box), добавляющий набор
клиентских фич (XHTTP, AmneziaWG 2.0, MASQUE, расширения наблюдаемости) за
build-тегами и живущий **ребейзом на каждый upstream-тег, а не merge**. Форк
build-тегами и живущий **ребейзом на каждый upstream-тег, а не merge**. Это набор
самого форка, а не shater: MASQUE/CONNECT-IP мы намеренно **не регистрируем** —
`shater/generate` его не порождает, а отказ от него и остального незадействованного
зоопарка экономит ~6 МБ бинаря и столько же RAM на роутере (`shater/registry`). Форк
разрабатывается по Spec Kit; неизменяемые принципы — в
[`SPECS/CONSTITUTION.md`](SPECS/CONSTITUTION.md), справочник фич движка — в
[`docs-lx/lx-config.ru.md`](docs-lx/lx-config.ru.md).
+145
View File
@@ -0,0 +1,145 @@
// lx:begin l3-honest-drop
package adapter
import (
"net/netip"
"testing"
"github.com/sagernet/sing-tun"
"github.com/sagernet/sing-tun/gtcpip/header"
"github.com/stretchr/testify/require"
)
// judgeFlowRouter answers PreMatch with a canned verdict; JudgeFlow reads
// nothing else off the Router.
type judgeFlowRouter struct {
Router
result PreMatchResult
}
func (r *judgeFlowRouter) PreMatch(InboundContext, []byte) PreMatchResult { return r.result }
// judgeFlowPort is the tun.Port half of a FlowOutbound. inet4 is what
// PortAddresses reports for IPv4 — the one field the two ICMP consumers in
// sing-tun disagree about (see the comment on
// TestJudgeFlowICMPToBoundPortStaysAFlow).
type judgeFlowPort struct {
Outbound
inet4 netip.Addr
}
func (o *judgeFlowPort) Tag() string { return "wg-out" }
func (o *judgeFlowPort) Type() string { return "wireguard" }
func (o *judgeFlowPort) PortAddresses() (netip.Addr, netip.Addr) {
return o.inet4, netip.Addr{}
}
func (o *judgeFlowPort) PortMTU() uint32 { return 1420 }
func (o *judgeFlowPort) AttachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) DetachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) WritePackets(packets [][]byte) error { return nil }
// judgeFlowNonPort is a FlowOutbound-shaped result that is NOT a tun.Port — the
// interface drift the second line of defense in JudgeFlow exists for.
type judgeFlowNonPort struct {
Outbound
}
func (o *judgeFlowNonPort) Tag() string { return "drifted" }
func (o *judgeFlowNonPort) Type() string { return "drifted" }
func judgeFlow(t *testing.T, protocol uint8, result PreMatchResult) tun.FlowVerdict {
t.Helper()
return JudgeFlow(
&judgeFlowRouter{result: result},
"l3-in", "tun", protocol,
netip.MustParseAddrPort("192.168.1.2:1234"),
netip.MustParseAddrPort("1.1.1.1:1234"),
nil,
)
}
const (
judgeFlowICMP = uint8(header.ICMPv4ProtocolNumber)
judgeFlowTCP = uint8(header.TCPProtocolNumber)
)
// TestJudgeFlowICMPToBoundPortStaysAFlow is the guard on the ONE fix that must
// not be made here.
//
// sing-tun has two ICMP consumers with different requirements on the port:
//
// - ForwardDispatcher.createFlow (flow_dispatch.go) needs only a VALID port
// address — it NATs the echo identifier and rewrites the source to that
// address. This is the path every unfragmented LAN ping takes, and it is
// what makes ping-through-WireGuard/AWG work at all.
// - ICMPForwarder.installFlow (stack_gvisor_icmp.go) additionally requires the
// address to be UNSPECIFIED, because it writes the packet to the port
// unmodified. A WireGuard endpoint reports its concrete interface address
// (transport/wireguard/port.go), so installFlow declines and HandlePacket
// falls through to forging the echo reply.
//
// The tempting fix — "for ICMP, refuse ActionFlow when PortAddresses() is not
// unspecified, so the verdict becomes a drop and the forgery is unreachable" —
// is applied HERE, in the one function both consumers share, with byte-identical
// arguments from either. It would therefore kill the working path too: every
// ping through WireGuard/AWG, fragmented or not, would drop, and l3_tunnel would
// carry nothing but `direct`. Keep this test failing loudly if anyone tries.
func TestJudgeFlowICMPToBoundPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.MustParseAddr("10.2.0.2")}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action,
"ICMP to a WireGuard/AWG endpoint must stay a flow: the forward dispatcher NATs it by echo identifier and this is the whole point of l3_tunnel")
require.Same(t, tun.Port(port), verdict.Port)
}
// The `direct` shape: an unspecified port address. Both consumers accept it.
func TestJudgeFlowICMPToUnspecifiedPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.IPv4Unspecified()}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action)
require.Same(t, tun.Port(port), verdict.Port)
}
// PreMatchDrop is the honest verdict and must arrive as ActionDrop: it is the
// only value (besides Reject) that stops ICMPForwarder.HandlePacket before the
// Echo -> EchoReply rewrite.
func TestJudgeFlowICMPDropReachesTheStackAsDrop(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchDrop})
require.Equal(t, tun.ActionDrop, verdict.Action)
}
// The second line of defense: a PreMatchFlow whose outbound is not a tun.Port
// must not degrade ICMP to ActionAccept, because Accept is the forged reply.
func TestJudgeFlowICMPNonPortOutboundDrops(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionDrop, verdict.Action,
"FlowOutbound and tun.Port are distinct interfaces; a drift between them must not silently re-enable the echo forger")
}
func TestJudgeFlowTCPNonPortOutboundAccepts(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionAccept, verdict.Action,
"for TCP, falling back to Accept is upstream behaviour and must stay untouched")
}
// TCP keeps every mapping it had, including the Continue -> Accept default that
// is a forgery only for ICMP.
func TestJudgeFlowTCPContinueStaysAccept(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchContinue})
require.Equal(t, tun.ActionAccept, verdict.Action)
}
func TestJudgeFlowTCPBypassStaysBypass(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchBypass})
require.Equal(t, tun.ActionBypass, verdict.Action)
}
// lx:end l3-honest-drop
+11
View File
@@ -75,7 +75,18 @@ func JudgeFlow(router Router, inbound string, inboundType string, network uint8,
case PreMatchFlow:
port, isPort := result.Outbound.(tun.Port)
if !isPort {
// lx:begin l3-honest-drop
// Second line of defense behind route.(*Router).preMatchFlow: a
// PreMatchFlow result already implies the outbound is an
// adapter.FlowOutbound, but FlowOutbound and tun.Port are distinct
// interfaces, and a drift between them must not degrade ICMP to
// ActionAccept — the TUN stack would then forge the echo reply
// itself instead of admitting the tunnel cannot carry the packet.
if networkName == N.NetworkICMP {
return tun.FlowVerdict{Action: tun.ActionDrop}
}
return tun.FlowVerdict{Action: tun.ActionAccept}
// lx:end l3-honest-drop
}
verdict := tun.FlowVerdict{Action: tun.ActionFlow, Port: port, UDPTimeout: result.UDPTimeout, NewTracker: result.NewTracker}
if result.Destination.IsValid() {
+6 -1
View File
@@ -110,7 +110,12 @@ docker run --rm --volumes-from "$(hostname)" \
# --- 2) sanity: the per-arch apk repo dir must be complete -------------------
[ -s "$OUT/packages.adb" ] || { echo "[apk-feed] ERROR: $OUT/packages.adb missing/empty" >&2; exit 4; }
apks=$(find "$OUT" -maxdepth 1 -name '*.apk' | wc -l)
[ "$apks" -ge 4 ] || { echo "[apk-feed] ERROR: expected >=4 .apk in $OUT, found $apks" >&2; exit 5; }
# Three, since D29 removed byedpi: shaterd, shater-core, luci-app-shater. The
# count lives in TWO scripts — sdk-build-apk.sh asserts what it collected out of
# bin/, this one asserts what reached the feed dir. v0.2.22 shipped with only the
# first one updated and the aarch64 lane died here on `found 3`, so if the set of
# packages ever changes again, change it in both.
[ "$apks" -ge 3 ] || { echo "[apk-feed] ERROR: expected >=3 .apk in $OUT, found $apks" >&2; exit 5; }
if [ -n "${KEY_APK:-}" ] && [ ! -s "$OUT/shater-apk.pem" ]; then
echo "[apk-feed] ERROR: signed feed but shater-apk.pem missing from $OUT" >&2; exit 6
fi
+14 -13
View File
@@ -30,7 +30,7 @@ echo "[apk-sdk] sdk=$SDK_URL"
# Package version derived from the git tag by ci/version.sh (bug B4). Forwarded
# to the unprivileged build user on the `su` line at the bottom of this file;
# openwrt/{shaterd,shater-core,luci-app-shater}/Makefile pick it up from the
# environment. byedpi keeps upstream ByeDPI's own version (see its Makefile).
# environment. All three are versioned from the tag — there is no exception.
echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -135,7 +135,7 @@ if ! ./scripts/feeds update -a; then
./scripts/feeds update -a
fi
echo "[apk-sdk] feeds install (prefer shater feed)"
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
./scripts/feeds install -p shater shaterd shater-core luci-app-shater
# --- strip the SDK's generated per-package `default m` blocks ----------------
# Run 60 settled the question that runs 58 and 59 left open. Writing an explicit
@@ -232,7 +232,7 @@ if [ -s .config.sdk ]; then
.config.sdk | tee -a .config | sed 's/^/[apk-sdk] /' || true
fi
for p in shaterd shater-core byedpi luci-app-shater; do
for p in shaterd shater-core luci-app-shater; do
echo "CONFIG_PACKAGE_$p=m" >> .config
done
# Route source downloads through OpenWrt's fast CDN mirror FIRST. Some upstreams
@@ -336,15 +336,15 @@ fi
echo "[apk-sdk] cache settings after defconfig:"
grep -E '^CONFIG_(LOCALMIRROR|DOWNLOAD_FOLDER)=' .config | sed 's/^/[apk-sdk] /' || true
echo "[apk-sdk] our packages after defconfig:"
grep -E '^CONFIG_PACKAGE_(shaterd|shater-core|byedpi|luci-app-shater)=' .config \
grep -E '^CONFIG_PACKAGE_(shaterd|shater-core|luci-app-shater)=' .config \
| sed 's/^/[apk-sdk] /' || true
# Each of our 4 must have SURVIVED defconfig. If kconfig dropped one, it is
# Each of our 3 must have SURVIVED defconfig. If kconfig dropped one, it is
# because a symbol it `select`s (a DEPENDS entry) does not exist in the installed
# feeds — with the old append-everything .config that was masked by the SDK
# pre-selecting half the distro. `make package/<p>/compile` would then die with a
# cryptic "No rule to make target", far from the real cause.
for p in shaterd shater-core byedpi luci-app-shater; do
for p in shaterd shater-core luci-app-shater; do
grep -q "^CONFIG_PACKAGE_$p=m" .config || {
echo "[apk-sdk] ERROR: $p is NOT selected after defconfig."
echo " kconfig dropped it -> one of its DEPENDS is missing from the"
@@ -368,7 +368,7 @@ done
grep -m5 '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/ /' || true
exit 11; }
for p in shaterd shater-core byedpi luci-app-shater; do
for p in shaterd shater-core luci-app-shater; do
echo "[apk-sdk] === build $p ==="
make "package/$p/compile" V=s -j"$(nproc)"
done
@@ -384,24 +384,25 @@ anyapk=$(find bin -type f -name '*.apk' | wc -l)
[ "$anyapk" -gt 0 ] || {
echo "[apk-sdk] ERROR: no .apk produced under bin/ (wrong/older SDK? found $(find bin -type f -name '*.ipk' | wc -l) .ipk)";
find bin -maxdepth 4 -type d || true; exit 6; }
# Collect ONLY our 4 packages' .apk (apk filenames carry NO arch:
# Collect ONLY our 3 packages' .apk (apk filenames carry NO arch:
# `<name>-<ver>-r<rel>.apk`). NOT a blanket `*.apk` copy — the SDK bin/ can hold
# prebuilt base/kmod .apk that would bloat the index and be signed under our key.
found=0
for p in shaterd shater-core byedpi luci-app-shater; do
for p in shaterd shater-core luci-app-shater; do
for a in $(find bin -type f -name "${p}-*.apk"); do
cp -f "$a" "$OUT/"; found=$((found+1))
done
done
[ "$found" -ge 4 ] || { echo "[apk-sdk] ERROR: expected >=4 of OUR .apk, collected $found"; echo "[apk-sdk] (all .apk under bin/:)"; find bin -type f -name '*.apk' | head -20; exit 6; }
[ "$found" -ge 3 ] || { echo "[apk-sdk] ERROR: expected >=3 of OUR .apk, collected $found"; echo "[apk-sdk] (all .apk under bin/:)"; find bin -type f -name '*.apk' | head -20; exit 6; }
echo "[apk-sdk] collected $found of our .apk"
# --- assert the tag-derived version actually reached the packages -------------
# B4's failure mode is a wrong-but-plausible version shipping silently, so the
# env -> make hand-off is verified, not trusted: each of our three tag-versioned
# packages must be named `<name>-<ver>-r<rel>.apk`. byedpi is excluded on purpose
# (it carries upstream ByeDPI's own version). This runs BEFORE `apk mkndx`, so a
# stale version can never even reach the index.
# packages must be named `<name>-<ver>-r<rel>.apk`. Every package this repo ships
# is tag-versioned, so the check covers all of them with no exception to
# remember. This runs BEFORE `apk mkndx`, so a stale version can never even
# reach the index.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
+3 -1
View File
@@ -38,7 +38,9 @@
# a dispatch build of the tagged commit itself identical to the release build of
# that same commit — which is the truth: same tree, same binary.
#
# `byedpi` is deliberately NOT versioned from our tag — see openwrt/byedpi/Makefile.
# Every package this repo ships is versioned from the tag. There used to be one
# exception (an external tool carrying its upstream's own version); it is gone
# with the package, and nothing here has to remember it any more.
#
# USAGE
# ci/version.sh # or --env: eval-able / $GITHUB_ENV-able lines
+223
View File
@@ -0,0 +1,223 @@
//go:build with_quic
package httpclient
import (
"context"
stdTLS "crypto/tls"
"io"
"net"
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
sbTLS "github.com/sagernet/sing-box/common/tls"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
)
// raceProbePayload is large enough that it cannot ride along in the response
// headers: the caller has to read the body off the QUIC stream AFTER
// roundTripHTTP3Race has returned. That is the whole point of the test.
const raceProbePayload = 64 * 1024
var _ N.Dialer = (*plainDialer)(nil)
type plainDialer struct{}
func (d *plainDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
return (&net.Dialer{}).DialContext(ctx, network, destination.String())
}
func (d *plainDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
// splitDialer sends the HTTP/3 racer and the HTTP/2 racer to two different
// listeners, so a test can decide which one of them wins without having to bind
// a TCP and a UDP socket on the same port number.
type splitDialer struct {
udp M.Socksaddr
tcp M.Socksaddr
}
func (d *splitDialer) DialContext(ctx context.Context, network string, _ M.Socksaddr) (net.Conn, error) {
destination := d.tcp
if network == N.NetworkUDP {
destination = d.udp
}
return (&net.Dialer{}).DialContext(ctx, network, destination.String())
}
func (d *splitDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
func startH3Server(t *testing.T, handler http.Handler) M.Socksaddr {
t.Helper()
certificate, err := sbTLS.GenerateKeyPair(nil, nil, nil, "localhost")
if err != nil {
t.Fatal(err)
}
listener, err := quic.ListenAddrEarly("127.0.0.1:0", &stdTLS.Config{
Certificates: []stdTLS.Certificate{*certificate},
NextProtos: []string{http3.NextProtoH3},
MinVersion: stdTLS.VersionTLS13,
}, nil)
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: handler}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
return M.ParseSocksaddr(listener.Addr().String())
}
func newRaceProbeTransport(t *testing.T, serverAddr M.Socksaddr) (*http3FallbackTransport, string) {
return newRaceProbeTransportWithDialer(t, &plainDialer{}, serverAddr)
}
func newRaceProbeTransportWithDialer(t *testing.T, dialer N.Dialer, serverAddr M.Socksaddr) (*http3FallbackTransport, string) {
t.Helper()
baseTLSConfig, err := sbTLS.NewClient(context.Background(), logger.NOP(), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
})
if err != nil {
t.Fatal(err)
}
h2Fallback, err := newHTTP2FallbackTransport(dialer, baseTLSConfig, option.HTTP2Options{})
if err != nil {
t.Fatal(err)
}
inner, err := newHTTP3FallbackTransport(dialer, baseTLSConfig, h2Fallback, option.QUICOptions{}, 300*time.Millisecond)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { inner.Close() })
return inner.(*http3FallbackTransport), "https://" + serverAddr.String() + "/probe"
}
// TestHTTP3RaceWinnerBodyStaysReadable pins that the response handed back by the
// HTTP/3 race is a LIVE response: its body must still be readable after
// roundTripHTTP3Race returns. Cancelling the context the winner was issued on
// resets its QUIC stream, so a "successful" round trip would hand the caller a
// response it can never read.
func TestHTTP3RaceWinnerBodyStaysReadable(t *testing.T) {
payload := make([]byte, raceProbePayload)
for i := range payload {
payload[i] = byte(i)
}
serverAddr := startH3Server(t, http.HandlerFunc(func(writer http.ResponseWriter, request *http.Request) {
writer.Header().Set("Content-Type", "application/octet-stream")
writer.Write(payload)
}))
transport, url := newRaceProbeTransport(t, serverAddr)
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
request, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
t.Fatal(err)
}
// No cached HTTP/3 connection yet and a bodyless GET is replayable, so this
// takes the racing path.
response, err := transport.RoundTrip(request)
if err != nil {
t.Fatal("round trip: ", err)
}
defer response.Body.Close()
if response.ProtoMajor != 3 {
t.Fatalf("expected the HTTP/3 racer to win, got HTTP/%d.%d", response.ProtoMajor, response.ProtoMinor)
}
body, err := io.ReadAll(response.Body)
if err != nil {
t.Fatalf("the race winner's body died with the race: %v (read %d of %d bytes)", err, len(body), len(payload))
}
if len(body) != len(payload) {
t.Fatalf("short body: got %d bytes, want %d", len(body), len(payload))
}
}
// TestHTTP3RaceFallbackWinnerBodyStaysReadableAndH3LoserIsCancelled covers the
// other half of the race: the HTTP/2 fallback wins, so its body must survive the
// race, and the HTTP/3 racer that lost must be torn down instead of being left
// to run to completion on the caller's behalf.
func TestHTTP3RaceFallbackWinnerBodyStaysReadableAndH3LoserIsCancelled(t *testing.T) {
payload := make([]byte, raceProbePayload)
for i := range payload {
payload[i] = byte(i)
}
h3Started := make(chan struct{}, 1)
h3Cancelled := make(chan struct{}, 1)
// The HTTP/3 handler never answers, so the fallback wins on the timer.
h3Addr := startH3Server(t, http.HandlerFunc(func(_ http.ResponseWriter, request *http.Request) {
select {
case h3Started <- struct{}{}:
default:
}
<-request.Context().Done()
select {
case h3Cancelled <- struct{}{}:
default:
}
}))
h2Server := httptest.NewUnstartedServer(http.HandlerFunc(func(writer http.ResponseWriter, _ *http.Request) {
writer.Header().Set("Content-Type", "application/octet-stream")
writer.Write(payload)
}))
h2Server.EnableHTTP2 = true
h2Server.StartTLS()
t.Cleanup(h2Server.Close)
transport, _ := newRaceProbeTransportWithDialer(t, &splitDialer{
udp: h3Addr,
tcp: M.ParseSocksaddr(h2Server.Listener.Addr().String()),
}, h3Addr)
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
request, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://localhost:443/probe", nil)
if err != nil {
t.Fatal(err)
}
response, err := transport.RoundTrip(request)
if err != nil {
t.Fatal("round trip: ", err)
}
if response.ProtoMajor != 2 {
t.Fatalf("expected the HTTP/2 fallback to win, got HTTP/%d.%d", response.ProtoMajor, response.ProtoMinor)
}
body, err := io.ReadAll(response.Body)
if err != nil {
t.Fatalf("the fallback winner's body died with the race: %v (read %d of %d bytes)", err, len(body), len(payload))
}
response.Body.Close()
if len(body) != len(payload) {
t.Fatalf("short body: got %d bytes, want %d", len(body), len(payload))
}
select {
case <-h3Started:
case <-time.After(5 * time.Second):
t.Fatal("the HTTP/3 racer never reached the server, the test proves nothing about cancelling it")
}
select {
case <-h3Cancelled:
case <-time.After(5 * time.Second):
t.Fatal("the losing HTTP/3 request was left running after the fallback won")
}
}
+67 -19
View File
@@ -6,6 +6,7 @@ import (
"context"
stdTLS "crypto/tls"
"errors"
"io"
"net/http"
"sync"
"time"
@@ -168,32 +169,65 @@ func (t *http3FallbackTransport) roundTripHTTP3(request *http.Request) (*http.Re
return t.roundTripHTTP3Race(request, authority)
}
// cancelOnBodyClose releases a racer's context when the caller is done with the
// response it won. The race cannot release it on the way out: the body is read
// after RoundTrip returns, and the context the request was issued on is what
// keeps its stream alive.
type cancelOnBodyClose struct {
io.ReadCloser
cancel context.CancelFunc
cancelOnce sync.Once
}
func (b *cancelOnBodyClose) Close() error {
err := b.ReadCloser.Close()
b.cancelOnce.Do(b.cancel)
return err
}
func withCancelOnBodyClose(response *http.Response, cancel context.CancelFunc) *http.Response {
if response == nil || response.Body == nil {
cancel()
return response
}
response.Body = &cancelOnBodyClose{ReadCloser: response.Body, cancel: cancel}
return response
}
func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, authority string) (*http.Response, error) {
ctx, cancel := context.WithCancel(request.Context())
defer cancel()
type result struct {
response *http.Response
err error
h3 bool
}
results := make(chan result, 2)
startRoundTrip := func(request *http.Request, useH3 bool) {
request = request.WithContext(ctx)
var (
response *http.Response
err error
)
if useH3 {
response, err = t.h3Transport.RoundTrip(request)
} else {
response, err = t.h2FallbackRoundTrip(request)
}
results <- result{response: response, err: err, h3: useH3}
// Each racer runs on a context of its own. A context shared by both cannot be
// cancelled when one of them wins: quic-go and net/http reset the winner's
// stream on cancellation, so the caller would be handed a response whose body
// stops mid-read with H3_REQUEST_CANCELLED. Only losers are cancelled here;
// the winner's cancel travels with its body and fires on Close.
startRoundTrip := func(useH3 bool) context.CancelFunc {
ctx, cancel := context.WithCancel(request.Context())
raceRequest := cloneRequestForRetry(request).WithContext(ctx)
go func() {
var (
response *http.Response
err error
)
if useH3 {
response, err = t.h3Transport.RoundTrip(raceRequest)
} else {
response, err = t.h2FallbackRoundTrip(raceRequest)
}
results <- result{response: response, err: err, h3: useH3}
}()
return cancel
}
goroutines := 1
received := 0
var fallbackCancel context.CancelFunc
h3Cancel := startRoundTrip(true)
drainRemaining := func() {
cancel()
for range goroutines - received {
go func() {
loser := <-results
@@ -203,7 +237,6 @@ func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, autho
}()
}
}
go startRoundTrip(cloneRequestForRetry(request), true)
timer := time.NewTimer(t.fallbackDelay)
defer timer.Stop()
var (
@@ -215,20 +248,28 @@ func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, autho
case <-timer.C:
if goroutines == 1 {
goroutines++
go startRoundTrip(cloneRequestForRetry(request), false)
fallbackCancel = startRoundTrip(false)
}
case raceResult := <-results:
received++
if raceResult.err == nil {
winnerCancel := fallbackCancel
if raceResult.h3 {
t.clearH3Broken(authority)
winnerCancel = h3Cancel
if fallbackCancel != nil {
fallbackCancel()
}
} else {
h3Cancel()
}
drainRemaining()
return raceResult.response, nil
return withCancelOnBodyClose(raceResult.response, winnerCancel), nil
}
if raceResult.h3 {
t.markH3Broken(authority)
h3Err = raceResult.err
h3Cancel()
if goroutines == 1 {
goroutines++
if !timer.Stop() {
@@ -237,14 +278,21 @@ func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, autho
default:
}
}
go startRoundTrip(cloneRequestForRetry(request), false)
fallbackCancel = startRoundTrip(false)
}
} else {
fallbackErr = raceResult.err
if fallbackCancel != nil {
fallbackCancel()
}
}
if received < goroutines {
continue
}
h3Cancel()
if fallbackCancel != nil {
fallbackCancel()
}
drainRemaining()
switch {
case h3Err != nil && fallbackErr != nil:
+227 -18
View File
@@ -6,6 +6,8 @@ import (
"encoding/binary"
"math/rand"
"net"
"net/netip"
"slices"
"strings"
"time"
@@ -47,6 +49,33 @@ func (c *Conn) Write(b []byte) (n int, err error) {
}()
serverName := IndexTLSServerName(b)
if serverName != nil {
// The SNI extension carries a LIST of names; MyServerName.Length is
// the length of the FIRST entry while MyServerName.ServerName is
// everything left in the extension. Plan the cuts over the first
// entry only: a second entry would otherwise be handed to the
// public suffix list as if it were part of the name.
name := serverName.ServerName
if serverName.Length >= 0 && serverName.Length < len(name) {
name = name[:serverName.Length]
}
// Packet fragmentation pays half a second per cut, record
// fragmentation pays microseconds — see the budget constants.
budget := recordCutBudget
if c.splitPacket {
budget = packetCutBudget
}
splitIndexes := cutOffsets(name, budget, rand.Intn)
if len(splitIndexes) == 0 {
// Nothing inside this name can be cut — it is empty or a single
// byte, so there is no offset that leaves a non-empty piece on
// both sides. Write the ClientHello as it stands: the loop
// below reads b[:splitIndexes[0]] unconditionally and would
// panic on an empty plan.
return c.Conn.Write(b)
}
for i := range splitIndexes {
splitIndexes[i] += serverName.Index
}
if c.splitPacket {
if c.tcpConn != nil {
err = c.tcpConn.SetNoDelay(true)
@@ -55,24 +84,6 @@ func (c *Conn) Write(b []byte) (n int, err error) {
}
}
}
splits := strings.Split(serverName.ServerName, ".")
currentIndex := serverName.Index
if publicSuffix := publicsuffix.List.PublicSuffix(serverName.ServerName); publicSuffix != "" {
splits = splits[:len(splits)-strings.Count(serverName.ServerName, ".")]
}
if len(splits) > 1 && splits[0] == "..." {
currentIndex += len(splits[0]) + 1
splits = splits[1:]
}
var splitIndexes []int
for i, split := range splits {
splitAt := rand.Intn(len(split))
splitIndexes = append(splitIndexes, currentIndex+splitAt)
currentIndex += len(split)
if i != len(splits)-1 {
currentIndex++
}
}
var buffer bytes.Buffer
for i := 0; i <= len(splitIndexes); i++ {
var payload []byte
@@ -133,6 +144,204 @@ func (c *Conn) Write(b []byte) (n int, err error) {
return c.Conn.Write(b)
}
// labelSpan is the half-open byte range [start, end) of one DNS label inside a
// server name, relative to the first byte of that name.
type labelSpan struct {
start int
end int
}
// How many cuts Write may spend on one ClientHello. The two numbers differ
// because the two modes cost completely different things per cut — MEASURED
// 2026-07-27 on this tree, loopback peer, product default fallbackDelay
// (500 ms), one ClientHello per row:
//
// cuts tls_fragment (*net.TCPConn) tls_fragment (proxy conn) tls_record_fragment
// 1 502 ms 500 ms <1 ms
// 2 1.004 s 1.001 s <1 ms
// 4 2.008 s 2.002 s <1 ms
// 8 4.015 s 4.003 s 539 µs
// 21 10.540 s 10.509 s 525 µs
//
// So a cut in the PACKET modes costs half a second of connection setup, and it
// costs that on BOTH branches: writeAndWaitAck sleeps the whole fallbackDelay
// anyway whenever the ACK comes back inside 20 ms (its "under transparent
// proxy" case), and N.UnwrapReader only reaches the *net.TCPConn when nothing
// in the chain transforms the stream — which a proxy protocol conn always does,
// so a proxied egress takes the flat-500 ms branch regardless of RTT. The
// number of labels is chosen by whoever picked the hostname, so "a cut in every
// label" made a 253-byte SNI worth ~42 s of one connection's setup.
//
// In tls_record_fragment nothing waits at all: the whole ClientHello leaves in
// ONE write, split into more TLS records. 21 cuts cost 525 µs and 105 bytes of
// record headers, and a real server (1.1.1.1) completed the handshake with the
// ClientHello in 22 records in the same 77 ms it took with 2. That mode is
// where "cut every label" was always affordable — and it is the mode the field
// measurement that started this was taken in.
//
// recordCutBudget is therefore not a cost limit but a shape limit: real names
// have one to three labels outside the public suffix, so 4 never binds on real
// traffic, while a hostile 253-byte name cannot turn one ClientHello into 85
// records that no ordinary client would ever emit.
const (
packetCutBudget = 1
recordCutBudget = 4
)
// cutOffsets plans where the ClientHello must be cut, in byte offsets relative
// to the FIRST BYTE OF THE SERVER NAME, spending at most budget cuts. randIntn
// is math/rand's Intn in production; a test hands in its own to make the plan
// deterministic.
//
// A cut is only a cut if it lands STRICTLY INSIDE a label. Offset 0 of a label
// is that label's own boundary: it leaves the label — the very string the DPI
// box matches on — whole in the following segment. That is not theory. The old
// code drew rand.Intn(len(label)), so a one-byte label could only ever produce
// offset 0, and on the measured provider (blocks by name in the handshake)
// m.youtube.com, tv.youtube.com and www.youtube.com were all blocked with the
// cut sitting uselessly at the start of "m"/"tv"/"www", while the name itself
// travelled intact in one segment. Hence a label shorter than two bytes carries
// no cut at all.
func cutOffsets(name string, budget int, randIntn func(n int) int) []int {
var offsets []int
for _, span := range cutCandidates(name) {
if span.end-span.start < 2 {
continue // no interior offset exists
}
offsets = append(offsets, cutInside(span, randIntn))
if len(offsets) >= budget {
break
}
}
if len(offsets) == 0 && len(name) >= 2 {
// No candidate label was long enough to cut on its own (a.b.co.uk).
// Cut the name somewhere rather than hand it over in one piece: a
// matcher looking for the whole FQDN still fails across the split, even
// though no single label was severed.
offsets = append(offsets, cutInside(labelSpan{start: 0, end: len(name)}, randIntn))
}
slices.Sort(offsets) // candidates are returned by priority, the wire wants order
return offsets
}
// cutInside draws an offset strictly inside span, from its MIDDLE THIRD.
//
// Every interior offset severs the label, but not equally well: a cut one byte
// in leaves "outube" of "youtube", and a matcher keyed on a suffix or on a
// six-byte substring still reads it. The middle leaves two short, unremarkable
// halves. The draw stays random inside that third — a fixed point (say, exactly
// the middle of the longest label) would be a constant a middlebox vendor can
// special-case in one line, and the whole family of fragmentation tricks lives
// on making reassembly the only counter.
func cutInside(span labelSpan, randIntn func(n int) int) int {
lo, hi := span.start+1, span.end-1 // the interior offsets, both inclusive
if margin := (span.end - span.start - 1) / 3; margin > 0 {
lo += margin
hi -= margin
}
return lo + randIntn(hi-lo+1)
}
// cutCandidates returns the labels of name that a cut may land in, MOST WORTH
// CUTTING FIRST — which matters because the budget above is small.
//
// First is the registrable label: the one immediately left of the public
// suffix. That is the label a name-based blocklist keys on ("youtube" of
// youtube.com, www.youtube.com and studio.youtube.com alike, "ytimg" of
// i9.ytimg.com, "example" of a.b.example.co.uk), and severing it also breaks
// any match on the whole FQDN, so one cut covers both matchers. It is chosen by
// STRUCTURE, from the public suffix list — not by length, which is the trap the
// old code fell into from the other side: in cdn-static-assets.youtube.com the
// longest label is not the blocked one.
//
// The rest follow longest-first: among labels we have no structural reason to
// rank, a long one is likelier to be a distinctive token than "www", "m" or
// "tv". They are only reached when the budget allows more than one cut, or when
// the registrable label is too short to cut.
//
// The public suffix itself is dropped because it is shared by everything under
// it and carries none of the blocked word. WIDENING this set needs no proof,
// NARROWING it does, so an input the public suffix list has no opinion about (a
// trailing dot, an unmanaged TLD, a name that IS a suffix) keeps every label.
// No branch here ends up with nothing to cut except the empty name, which has
// nothing to cut by construction.
func cutCandidates(name string) []labelSpan {
spans := labelSpans(name)
suffix := publicsuffix.List.PublicSuffix(name)
switch {
case len(spans) == 0:
// name == "". Nothing to cut; Write sends the ClientHello unchanged.
case isIPLiteral(name):
// An IP literal is not a name (RFC 6066 forbids it in SNI) and its dots
// do not separate labels, so the public suffix list has nothing to say
// about it — it returns the literal itself. Treat the whole literal as
// one token: there is no name for a DPI box to read here, but the
// caller asked for a fragmented handshake and gets one.
return []labelSpan{{start: 0, end: len(name)}}
case suffix != "" && len(suffix) < len(name) && strings.HasSuffix(name, "."+suffix):
// The ordinary case, and the one the old arithmetic got wrong: it
// subtracted the number of dots in the WHOLE NAME, which — labels being
// always one more than dots — left exactly one label, the FIRST, for
// every name in existence. Subtract the number of labels in the SUFFIX
// instead: "com" is one ("www.youtube.com" keeps www + youtube),
// "co.uk" is two ("a.b.co.uk" keeps a + b).
if keep := len(spans) - strings.Count(suffix, ".") - 1; keep > 0 {
spans = spans[:keep]
}
// Everything else — suffix == "" (a trailing dot, which the list
// declines to parse), suffix == name (the name IS a public suffix:
// "com", "co.uk", "localhost"), or a suffix that is somehow not a tail
// of the name — keeps every label. Cutting inside a suffix costs a
// segment and hides nothing that was not already hidden; NOT cutting is
// the expensive mistake.
}
return byCutPriority(spans)
}
// byCutPriority puts the registrable label first and orders the rest
// longest-first. It never drops a span, so the budget — not this — decides how
// many labels are actually cut.
func byCutPriority(spans []labelSpan) []labelSpan {
if len(spans) < 2 {
return spans
}
out := make([]labelSpan, 0, len(spans))
out = append(out, spans[len(spans)-1])
rest := make([]labelSpan, len(spans)-1)
copy(rest, spans[:len(spans)-1])
slices.SortStableFunc(rest, func(a, b labelSpan) int {
return (b.end - b.start) - (a.end - a.start)
})
return append(out, rest...)
}
// labelSpans splits name on '.' and returns the byte range of each label.
// Empty labels (a leading, trailing or doubled dot) come back as zero-width
// spans and are dropped by cutOffsets, which is what keeps a name like
// ".youtube.com" away from rand.Intn(0) — that combination panicked.
func labelSpans(name string) []labelSpan {
if name == "" {
return nil
}
var spans []labelSpan
start := 0
for i := 0; i <= len(name); i++ {
if i == len(name) || name[i] == '.' {
spans = append(spans, labelSpan{start: start, end: i})
start = i + 1
}
}
return spans
}
func isIPLiteral(name string) bool {
_, err := netip.ParseAddr(name)
return err == nil
}
func (c *Conn) ReaderReplaceable() bool {
return true
}
+742
View File
@@ -0,0 +1,742 @@
package tf
// Cut planning: which label of the SNI gets a cut, where inside it, and how
// many cuts one ClientHello is allowed to cost.
//
// WHY THIS FILE EXISTS (2026-07-27)
// Conn.Write used to compute the labels to cut as
//
// splits = splits[:len(splits)-strings.Count(serverName.ServerName, ".")]
//
// which is identically splits[:1] for EVERY name, labels being always one
// more than dots. Exactly one label was ever cut — the LEFTMOST — so on a
// provider that blocks by the name in the handshake, youtube.com passed (its
// first label IS the blocked word) while m./tv./www./music./studio.youtube.com
// were all blocked, the cut sitting inside "m"/"tv"/"www" while "youtube"
// travelled whole in the next segment. Measured on the router.
//
// Two more halves of the same defect:
// - the offset came from rand.Intn(len(label)), whose 0 is the label's own
// boundary and severs nothing. For a one-byte label that is the ONLY
// value it can take;
// - an EMPTY label reached rand.Intn(0) and panicked the process. Reachable
// from the LAN: route/conn.go wraps the outbound with this and the
// ClientHello it fragments is the client's. See
// TestWriteDoesNotPanicOnAServerNameChosenFromTheLAN.
//
// Every test below fails on the old expressions — see the mutation log.
import (
"crypto/tls"
"encoding/binary"
"io"
"math/rand"
"net"
"strings"
"testing"
"time"
"github.com/stretchr/testify/require"
)
// --- deterministic draws -----------------------------------------------------
// minRand takes the lowest offset a label allows, maxRand the highest. Between
// them they pin BOTH ends of the range, which is where the interesting failures
// live: it is the ends that decide whether the label is severed or only touched.
func minRand(int) int { return 0 }
func maxRand(n int) int { return n - 1 }
func fixedRand(v int) func(int) int {
return func(n int) int {
if v >= n {
return n - 1
}
return v
}
}
// severedLabel names the label that a cut at offset o splits in two, or says
// why it splits none. This is the assertion vocabulary of the whole file: the
// question is never "which number came out" but "which word did we break".
func severedLabel(name string, o int) string {
switch {
case o <= 0 || o >= len(name):
return "!outside the name"
case name[o] == '.' || name[o-1] == '.':
return "!a label boundary, nothing severed"
}
start := strings.LastIndexByte(name[:o], '.') + 1
end := len(name)
if i := strings.IndexByte(name[o:], '.'); i >= 0 {
end = o + i
}
return name[start:end]
}
func severedLabels(name string, offsets []int) []string {
var out []string
for _, o := range offsets {
out = append(out, severedLabel(name, o))
}
return out
}
// --- which label is cut ------------------------------------------------------
func TestCutOffsetsCutTheRegistrableLabel(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
packet []string // labels severed with the packet budget (1 cut)
record []string // ... and with the record budget (4 cuts), in wire order
why string
}{
{name: "youtube.com", packet: []string{"youtube"}, record: []string{"youtube"}},
{name: "www.youtube.com", packet: []string{"youtube"}, record: []string{"www", "youtube"},
why: "THE regression: the old code cut www and shipped youtube whole"},
{name: "m.youtube.com", packet: []string{"youtube"}, record: []string{"youtube"},
why: "a one-byte label has no interior offset and carries no cut"},
{name: "tv.youtube.com", packet: []string{"youtube"}, record: []string{"tv", "youtube"}},
{name: "music.youtube.com", packet: []string{"youtube"}, record: []string{"music", "youtube"}},
{name: "studio.youtube.com", packet: []string{"youtube"}, record: []string{"studio", "youtube"}},
{name: "cdn-static-assets.youtube.com", packet: []string{"youtube"}, record: []string{"cdn-static-assets", "youtube"},
why: "the LONGEST label is not the blocked one — structure decides, not length"},
{name: "foo.bar.baz.youtube.com", packet: []string{"youtube"}, record: []string{"foo", "bar", "baz", "youtube"},
why: "four candidates, four cuts, and the budget stops there"},
{name: "a.b.c.d.e.youtube.com", packet: []string{"youtube"}, record: []string{"youtube"},
why: "five one-byte labels: the budget is never even reached"},
{name: "i9.ytimg.com", packet: []string{"ytimg"}, record: []string{"i9", "ytimg"}},
{name: "example.co.uk", packet: []string{"example"}, record: []string{"example"},
why: "co.uk is TWO labels of public suffix"},
{name: "a.b.example.co.uk", packet: []string{"example"}, record: []string{"example"}},
{name: "example.com.br", packet: []string{"example"}, record: []string{"example"}},
{name: "site.pp.ru", packet: []string{"site"}, record: []string{"site"},
why: "pp.ru is a private two-label suffix"},
{name: "localhost", packet: []string{"localhost"}, record: []string{"localhost"},
why: "unmanaged TLD: the list returns the whole name, so cut it"},
{name: "com", packet: []string{"com"}, record: []string{"com"}},
{name: "co.uk", packet: []string{"uk"}, record: []string{"co", "uk"},
why: "the name IS the suffix: keep every label rather than cut nothing"},
{name: ".youtube.com", packet: []string{"youtube"}, record: []string{"youtube"},
why: "leading dot: the empty label is skipped, NOT fed to rand.Intn(0)"},
{name: "youtube.com.", packet: []string{"youtube"}, record: []string{"youtube", "com"},
why: "trailing dot: the list declines to parse it, so every label stays a candidate"},
{name: "WWW.YouTube.COM", packet: []string{"YouTube"}, record: []string{"WWW", "YouTube"}},
{name: "ab", packet: []string{"ab"}, record: []string{"ab"}},
} {
t.Run(tc.name, func(t *testing.T) {
for _, draw := range []struct {
label string
fn func(int) int
}{{"lowest", minRand}, {"highest", maxRand}} {
got := severedLabels(tc.name, cutOffsets(tc.name, packetCutBudget, draw.fn))
require.Equal(t, tc.packet, got, "%s draw, packet budget: %s", draw.label, tc.why)
got = severedLabels(tc.name, cutOffsets(tc.name, recordCutBudget, draw.fn))
require.Equal(t, tc.record, got, "%s draw, record budget: %s", draw.label, tc.why)
}
})
}
}
// TestCutOffsetsSeverTheBlockedLabel is the field measurement turned into an
// instrument. On the measured provider these names differ only in the label in
// front of "youtube", and five of the six were blocked. What has to hold — for
// every draw and both budgets, not for most of them — is that the byte range of
// the blocked word straddles a cut.
func TestCutOffsetsSeverTheBlockedLabel(t *testing.T) {
t.Parallel()
for _, tc := range []struct{ name, blocked string }{
{"youtube.com", "youtube"},
{"m.youtube.com", "youtube"},
{"tv.youtube.com", "youtube"},
{"www.youtube.com", "youtube"},
{"music.youtube.com", "youtube"},
{"studio.youtube.com", "youtube"},
{"cdn-static-assets.youtube.com", "youtube"},
{"i9.ytimg.com", "ytimg"},
{"a.b.example.co.uk", "example"},
} {
t.Run(tc.name, func(t *testing.T) {
start := strings.Index(tc.name, tc.blocked)
require.GreaterOrEqual(t, start, 0)
end := start + len(tc.blocked)
// Every draw the label can take, not a sample: the range is small
// enough to enumerate, so there is no "it passed 1000 times" here.
for draw := 0; draw < len(tc.name); draw++ {
for _, budget := range []int{packetCutBudget, recordCutBudget} {
offsets := cutOffsets(tc.name, budget, fixedRand(draw))
severed := false
for _, o := range offsets {
if o > start && o < end {
severed = true
}
}
require.True(t, severed,
"draw %d, budget %d: %q got cuts at %v (%v), none inside %q [%d,%d)",
draw, budget, tc.name, offsets, severedLabels(tc.name, offsets), tc.blocked, start, end)
}
}
})
}
}
// TestCutOffsetsStayInTheMiddleThird: every interior offset severs the label,
// but not equally well — one byte in leaves "outube" of "youtube", which a
// matcher keyed on a substring still reads. Both halves must keep at least
// (width-1)/3 + 1 bytes.
func TestCutOffsetsStayInTheMiddleThird(t *testing.T) {
t.Parallel()
for _, name := range []string{
"youtube.com", "www.youtube.com", "cdn-static-assets.youtube.com",
"music.youtube.com", "example.co.uk", "ab.example.com", "localhost",
} {
t.Run(name, func(t *testing.T) {
for draw := 0; draw < 64; draw++ {
for _, budget := range []int{packetCutBudget, recordCutBudget} {
for _, o := range cutOffsets(name, budget, fixedRand(draw)) {
label := severedLabel(name, o)
require.NotContains(t, label, "!", "draw %d: cut at %d in %q severed nothing", draw, o, name)
start := strings.Index(name, label)
width := len(label)
margin := (width-1)/3 + 1
require.GreaterOrEqual(t, o-start, margin,
"draw %d: cut at %d leaves only %d byte(s) of %q on the left", draw, o, o-start, label)
require.GreaterOrEqual(t, start+width-o, margin,
"draw %d: cut at %d leaves only %d byte(s) of %q on the right", draw, o, start+width-o, label)
}
}
}
})
}
}
// TestCutOffsetsRespectTheBudget: the budget is what bounds a hostile name's
// cost — a measured 500 ms of connection setup per cut in the packet modes.
func TestCutOffsetsRespectTheBudget(t *testing.T) {
t.Parallel()
var long strings.Builder
for i := 0; i < 40; i++ {
long.WriteString("lb.")
}
long.WriteString("example.com") // 40 cuttable labels plus the registrable one
for _, budget := range []int{1, 2, 3, 4} {
require.Len(t, cutOffsets(long.String(), budget, rand.Intn), budget, "budget %d", budget)
}
require.Len(t, cutOffsets(long.String(), packetCutBudget, rand.Intn), 1,
"a 253-byte SNI must not be able to buy more than one 500 ms wait")
require.Len(t, cutOffsets(long.String(), recordCutBudget, rand.Intn), 4,
"nor more than five records")
}
// TestCutOffsetsFallBackWhenNoLabelCanBeCut covers the names where NO candidate
// label has an interior offset. Severing a label is impossible there, so what
// is checked is that a cut still happens and still lands inside the buffer: a
// matcher keyed on the whole FQDN fails across it.
func TestCutOffsetsFallBackWhenNoLabelCanBeCut(t *testing.T) {
t.Parallel()
for _, name := range []string{"a.b.co.uk", "x.pp.ru", "a.b.c.d", "1.2.3.4", "::1", "x.com"} {
t.Run(name, func(t *testing.T) {
for draw := 0; draw < len(name)+4; draw++ {
for _, budget := range []int{packetCutBudget, recordCutBudget} {
offsets := cutOffsets(name, budget, fixedRand(draw))
require.NotEmpty(t, offsets, "%q went out in one piece", name)
require.Greater(t, offsets[0], 0)
require.Less(t, offsets[len(offsets)-1], len(name))
}
}
})
}
}
// TestCutOffsetsAreOrderedAndDistinct: the write loop slices b between
// consecutive offsets, so anything out of order or repeated is an empty or
// negative segment on the wire. Candidates come back in PRIORITY order, which
// is not wire order — this is the test that the sort is not forgotten.
func TestCutOffsetsAreOrderedAndDistinct(t *testing.T) {
t.Parallel()
for _, name := range []string{
"www.youtube.com", "foo.bar.baz.youtube.com", "cdn-static-assets.youtube.com",
"a.bb.ccc.dddd.example.com", "youtube.com.", ".youtube.com", "co.uk",
} {
t.Run(name, func(t *testing.T) {
for i := 0; i < 200; i++ {
offsets := cutOffsets(name, recordCutBudget, rand.Intn)
prev := 0
for _, o := range offsets {
require.Greater(t, o, prev, "%q: offsets %v are not strictly increasing", name, offsets)
prev = o
}
require.Less(t, prev, len(name))
}
})
}
}
// --- the arithmetic must not panic on anything ------------------------------
// adversarialNames is the closed list of shapes that reach the arithmetic from
// outside: empty and one-byte names, every position a dot can take, names that
// are nothing but dots, names at the 253-byte limit, and bytes that are not
// ASCII at all. The unit test, the fuzz seed corpus and the end-to-end test all
// draw from it, so all three see the same inputs.
func adversarialNames() []string {
return []string{
"", "a", ".", "..", "...", "....",
".com", "com.", ".com.", ".youtube.com", "youtube.com.", ".youtube.com.",
"a..b.example.com", "..youtube..com..", "-.-.-.-", "-", "--",
"xn--p1ai", "test.xn--p1ai", "xn--", ".xn--p1ai.",
"\xff\xfe.example.com", "\x00\x00.com", "пример.рф", "\xff",
strings.Repeat("a", 253),
strings.Repeat("ab.", 84) + "a", // 253 bytes, 85 labels
strings.Repeat(".", 253),
strings.Repeat("a.", 126) + "a",
"1.2.3.4", "::1", "::ffff:1.2.3.4", "fe80::1%eth0", "0.0.0.0",
}
}
func TestCutOffsetsSurviveEveryAdversarialName(t *testing.T) {
t.Parallel()
for _, name := range adversarialNames() {
t.Run(strings.ToValidUTF8(name, "?"), func(t *testing.T) {
for i := 0; i < 100; i++ {
for _, budget := range []int{packetCutBudget, recordCutBudget} {
offsets := cutOffsets(name, budget, rand.Intn) // must not panic
require.LessOrEqual(t, len(offsets), budget)
if len(name) >= 2 {
require.NotEmpty(t, offsets, "%q is long enough to cut and was not cut", name)
} else {
require.Empty(t, offsets, "%q has no offset that leaves bytes on both sides", name)
}
prev := 0
for _, o := range offsets {
require.Greater(t, o, prev)
require.Less(t, o, len(name))
prev = o
}
}
}
})
}
}
// FuzzCutOffsets is the open half of the audit above: the closed list says what
// we thought of, this says whether anything else reaches rand.Intn with a
// non-positive argument or produces an offset the write loop cannot slice at.
// Under plain `go test` it runs the seed corpus, which is that closed list.
func FuzzCutOffsets(f *testing.F) {
for _, name := range adversarialNames() {
f.Add(name, packetCutBudget)
f.Add(name, recordCutBudget)
}
for _, name := range []string{"www.youtube.com", "a.b.example.co.uk", "localhost"} {
f.Add(name, 1)
f.Add(name, 4)
}
f.Fuzz(func(t *testing.T, name string, budget int) {
if budget < 0 {
budget = -budget
}
budget = budget%recordCutBudget + 1 // 1..4, never zero or negative
offsets := cutOffsets(name, budget, rand.Intn)
if len(offsets) > budget {
t.Fatalf("%q: %d offsets for a budget of %d", name, len(offsets), budget)
}
if len(name) >= 2 && len(offsets) == 0 {
t.Fatalf("%q (%d bytes) was handed over in one piece", name, len(name))
}
prev := 0
for _, o := range offsets {
if o <= prev || o >= len(name) {
t.Fatalf("%q: offsets %v are not strictly increasing inside [1,%d)", name, offsets, len(name))
}
prev = o
}
})
}
// --- end to end: what actually goes out on the wire -------------------------
// fakeConn records every Write. It is deliberately NOT a *net.TCPConn, which is
// also the common production case (the outbound is usually a proxy stream whose
// reader transforms the bytes, so N.UnwrapReader stops there), so Conn.Write
// takes the sleep-instead-of-ACK path — hence the 1ns fallback delay the tests
// below pass to NewConn.
type fakeConn struct {
writes [][]byte
}
func (c *fakeConn) Read([]byte) (int, error) { return 0, io.EOF }
func (c *fakeConn) Close() error { return nil }
func (c *fakeConn) LocalAddr() net.Addr { return &net.TCPAddr{} }
func (c *fakeConn) RemoteAddr() net.Addr { return &net.TCPAddr{} }
func (c *fakeConn) SetDeadline(time.Time) error { return nil }
func (c *fakeConn) SetReadDeadline(time.Time) error { return nil }
func (c *fakeConn) SetWriteDeadline(time.Time) error { return nil }
func (c *fakeConn) Write(b []byte) (int, error) {
c.writes = append(c.writes, append([]byte(nil), b...))
return len(b), nil
}
// clientHelloFor produces a real ClientHello for serverName by letting
// crypto/tls build one and capturing the first write.
func clientHelloFor(t *testing.T, serverName string) []byte {
t.Helper()
rec := &fakeConn{}
_ = tls.Client(rec, &tls.Config{ServerName: serverName, MinVersion: tls.VersionTLS12}).Handshake()
require.NotEmpty(t, rec.writes, "crypto/tls wrote no ClientHello for %q", serverName)
hello := rec.writes[0]
// Control on the instrument: the parser this package ships must find the
// name we asked for, otherwise the assertions below prove nothing.
sni := IndexTLSServerName(hello)
require.NotNil(t, sni, "IndexTLSServerName found no SNI in the generated ClientHello")
require.Equal(t, serverName, sni.ServerName)
return hello
}
// buildClientHello assembles a ClientHello by hand around a server_name_list of
// the given entries. crypto/tls will not emit a name with a leading or trailing
// dot, an IP literal, an empty name or a second list entry — hostnameInSNI
// rewrites or refuses all of them — and those are exactly the shapes a
// forwarded ClientHello from a LAN client can carry.
func buildClientHello(t *testing.T, entries ...string) []byte {
t.Helper()
var list []byte
for _, e := range entries {
list = append(list, sniNameDNSHostnameType)
list = binary.BigEndian.AppendUint16(list, uint16(len(e)))
list = append(list, e...)
}
extBody := binary.BigEndian.AppendUint16(nil, uint16(len(list)))
extBody = append(extBody, list...)
ext := binary.BigEndian.AppendUint16(nil, sniExtensionType)
ext = binary.BigEndian.AppendUint16(ext, uint16(len(extBody)))
ext = append(ext, extBody...)
extensions := binary.BigEndian.AppendUint16(nil, uint16(len(ext)))
extensions = append(extensions, ext...)
body := []byte{0x03, 0x03} // client_version TLS 1.2
body = append(body, make([]byte, 32)...) // random
body = append(body, 0x00) // session_id length
body = append(body, 0x00, 0x02, 0x13, 0x01) // cipher_suites
body = append(body, 0x01, 0x00) // compression_methods
body = append(body, extensions...)
handshake := []byte{handshakeType, byte(len(body) >> 16), byte(len(body) >> 8), byte(len(body))}
handshake = append(handshake, body...)
record := []byte{contentType, 0x03, 0x01}
record = binary.BigEndian.AppendUint16(record, uint16(len(handshake)))
record = append(record, handshake...)
// Control on the instrument: this hand-built record must parse the way a
// real one does, or the tests below are testing a straw man.
sni := IndexTLSServerName(record)
require.NotNil(t, sni, "hand-built ClientHello did not parse")
require.Equal(t, len(entries[0]), sni.Length, "Length must be the FIRST entry")
require.Equal(t, entries[0], string(record[sni.Index:sni.Index+sni.Length]))
return record
}
// patchSNI rewrites the server name inside a ClientHello in place. from and to
// must be the same length, so every length field in the record stays valid.
func patchSNI(t *testing.T, hello []byte, from, to string) []byte {
t.Helper()
require.Equal(t, len(from), len(to), "patchSNI cannot change the length")
at := IndexTLSServerName(hello)
require.NotNil(t, at)
require.Equal(t, from, at.ServerName)
out := append([]byte(nil), hello...)
copy(out[at.Index:], to)
sni := IndexTLSServerName(out)
require.NotNil(t, sni)
require.Equal(t, to, sni.ServerName)
return out
}
// segments returns, for one recorded run, the payload of every segment written
// and the absolute offsets in hello at which the cuts fell.
func segments(t *testing.T, hello []byte, writes [][]byte, recordFragment bool) ([][]byte, []int) {
t.Helper()
var payloads [][]byte
for _, w := range writes {
if !recordFragment {
payloads = append(payloads, w)
continue
}
// A record-fragmented write is one or more TLS records: 3 bytes of the
// original header, a 2-byte length, then the payload.
for len(w) > 0 {
require.GreaterOrEqual(t, len(w), recordLayerHeaderLen, "truncated record header")
require.Equal(t, hello[:3], w[:3], "record header is not the ClientHello's own")
n := int(binary.BigEndian.Uint16(w[3:5]))
require.LessOrEqual(t, recordLayerHeaderLen+n, len(w), "record length runs past the write")
payloads = append(payloads, w[recordLayerHeaderLen:recordLayerHeaderLen+n])
w = w[recordLayerHeaderLen+n:]
}
}
require.NotEmpty(t, payloads, "Write returned without putting anything on the wire")
// Cut offsets are the cumulative payload lengths, shifted past the record
// header that the first fragment drops.
offset := 0
if recordFragment {
offset = recordLayerHeaderLen
}
var cuts []int
for _, p := range payloads[:len(payloads)-1] {
offset += len(p)
cuts = append(cuts, offset)
}
return payloads, cuts
}
type writeMode struct {
name string
splitPacket bool
splitRecord bool
recordFraming bool
segmentPerCall bool // one Write call per segment
budget int
}
var writeModes = []writeMode{
{name: "tls_fragment", splitPacket: true, segmentPerCall: true, budget: packetCutBudget},
{name: "tls_record_fragment", splitRecord: true, recordFraming: true, budget: recordCutBudget},
{name: "both", splitPacket: true, splitRecord: true, recordFraming: true, segmentPerCall: true, budget: packetCutBudget},
}
// TestWriteSeversTheBlockedLabelOnTheWire is the end-to-end control: not "the
// planner returned nice numbers" but "the bytes that left the socket have the
// blocked label straddling a segment boundary", for every mode the presets
// expose, over many real random draws.
func TestWriteSeversTheBlockedLabelOnTheWire(t *testing.T) {
t.Parallel()
for _, mode := range writeModes {
for _, tc := range []struct{ name, blocked string }{
{"youtube.com", "youtube"}, // the ONE name the old code got right
{"www.youtube.com", "youtube"}, // the regression
{"m.youtube.com", "youtube"}, // one-byte label in front
{"music.youtube.com", "youtube"},
{"cdn-static-assets.youtube.com", "youtube"},
{"a.b.example.co.uk", "example"}, // two-label public suffix
} {
t.Run(mode.name+"/"+tc.name, func(t *testing.T) {
t.Parallel()
hello := clientHelloFor(t, tc.name)
sniAt := IndexTLSServerName(hello).Index
start := sniAt + strings.Index(tc.name, tc.blocked)
end := start + len(tc.blocked)
for i := 0; i < 100; i++ {
out := &fakeConn{}
n, err := NewConn(out, t.Context(), mode.splitPacket, mode.splitRecord, time.Nanosecond).Write(hello)
require.NoError(t, err)
require.Equal(t, len(hello), n, "Write must report the length of the buffer it was given")
_, cuts := segments(t, hello, out.writes, mode.recordFraming)
require.NotEmpty(t, cuts, "the ClientHello went out in one piece")
require.LessOrEqual(t, len(cuts), mode.budget, "more cuts than this mode's budget")
severed := false
for _, c := range cuts {
if c > start && c < end {
severed = true
}
}
require.True(t, severed,
"run %d: %q left with cuts at %v, none inside %q [%d,%d)",
i, tc.name, cuts, tc.blocked, start, end)
}
})
}
}
}
// TestWriteReassemblesToTheOriginalClientHello: cutting may change how the
// bytes are packaged and nothing else. Byte-for-byte, plus the length Write
// reports, plus the segment count implied by the plan.
func TestWriteReassemblesToTheOriginalClientHello(t *testing.T) {
t.Parallel()
for _, mode := range writeModes {
for _, serverName := range []string{
"www.youtube.com", "youtube.com", "a.b.example.co.uk", "localhost", "a",
"foo.bar.baz.youtube.com",
} {
t.Run(mode.name+"/"+serverName, func(t *testing.T) {
t.Parallel()
hello := clientHelloFor(t, serverName)
for i := 0; i < 50; i++ {
out := &fakeConn{}
n, err := NewConn(out, t.Context(), mode.splitPacket, mode.splitRecord, time.Nanosecond).Write(hello)
require.NoError(t, err)
require.Equal(t, len(hello), n)
payloads, cuts := segments(t, hello, out.writes, mode.recordFraming)
var joined []byte
for _, p := range payloads {
require.NotEmpty(t, p, "empty segment: a cut of zero length went out on the wire")
joined = append(joined, p...)
}
want := hello
if mode.recordFraming {
// The record header is re-emitted per fragment, so what
// must survive is the handshake body.
want = hello[recordLayerHeaderLen:]
}
require.Equal(t, want, joined, "run %d: the reassembled ClientHello differs from the original", i)
if mode.segmentPerCall {
require.Len(t, out.writes, len(cuts)+1, "one Write call per segment")
} else {
require.Len(t, out.writes, 1, "record fragmentation without packet fragmentation is a single write")
}
}
})
}
}
}
// TestWriteDoesNotPanicOnAServerNameChosenFromTheLAN is the regression for a
// PROCESS DEATH that shipped in v0.2.21.
//
// route/conn.go wraps the outbound connection with this Conn and fragments the
// ClientHello the LAN client sent, so the server name is chosen by the client,
// not by us. An empty label made cutOffsets call rand.Intn(0) — "panic: invalid
// argument to Intn" — and a Go panic in a connection goroutine takes the whole
// daemon with it. Two ordinary ways to produce one:
//
// "youtube.com." a fully qualified name with the root dot, which curl and
// every browser will happily send, and for which the public
// suffix list returns "" so the trailing empty label survived;
// ".youtube.com" a leading dot, which nothing legitimate sends but nothing
// stops a client from writing into its own ClientHello.
//
// With the kill switch armed the daemon's death is not a slow connection, it is
// a dark LAN until procd restarts it — into the same request.
func TestWriteDoesNotPanicOnAServerNameChosenFromTheLAN(t *testing.T) {
t.Parallel()
for _, tc := range []struct{ from, to, why string }{
{from: "youtube.comx", to: "youtube.com.", why: "the FQDN root dot — a legitimate name"},
{from: "xyoutube.com", to: ".youtube.com", why: "leading dot"},
{from: "xyoutube.comx", to: ".youtube.com.", why: "both"},
{from: "ax.example.com", to: "a..example.com", why: "a doubled dot mid-name"},
{from: "xxxxxxxxxxxx", to: "............", why: "nothing but dots"},
{from: "1x2x3x4", to: "1.2.3.4", why: "an IP literal, which RFC 6066 forbids in SNI"},
{from: "xxx", to: "::1", why: "an IPv6 literal"},
{from: "a", to: "a", why: "one byte: no cut exists, and the empty plan must not be indexed"},
{from: "ab", to: "ab", why: "two bytes: exactly one interior offset"},
{from: "\xff\xfe.example.com", to: "\xff\xfe.example.com", why: "bytes that are not ASCII"},
} {
t.Run(strings.ToValidUTF8(tc.to, "?"), func(t *testing.T) {
t.Parallel()
hello := patchSNI(t, clientHelloFor(t, tc.from), tc.from, tc.to)
for _, mode := range writeModes {
for i := 0; i < 50; i++ {
out := &fakeConn{}
n, err := NewConn(out, t.Context(), mode.splitPacket, mode.splitRecord, time.Nanosecond).Write(hello)
require.NoError(t, err, "%s: %s", mode.name, tc.why)
require.Equal(t, len(hello), n, "%s: %s", mode.name, tc.why)
payloads, _ := segments(t, hello, out.writes, mode.recordFraming)
var joined []byte
for _, p := range payloads {
require.NotEmpty(t, p)
joined = append(joined, p...)
}
want := hello
if mode.recordFraming {
want = hello[recordLayerHeaderLen:]
}
require.Equal(t, want, joined, "%s run %d: %s", mode.name, i, tc.why)
}
}
})
}
}
// TestWriteHandlesServerNameListsCryptoTLSWillNotEmit reaches the shapes that
// need a hand-built record: a zero-length name, a name at the 253-byte limit,
// and a list carrying a SECOND entry — which MyServerName.ServerName includes
// and MyServerName.Length does not, so cut planning must run on the first entry
// alone or it feeds the public suffix list bytes that belong to no name.
func TestWriteHandlesServerNameListsCryptoTLSWillNotEmit(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
title string
entries []string
cut bool // must the ClientHello leave in more than one piece?
}{
{title: "empty name", entries: []string{""}, cut: false},
{title: "one byte", entries: []string{"a"}, cut: false},
{title: "253 bytes, 85 labels", entries: []string{strings.Repeat("ab.", 84) + "a"}, cut: true},
{title: "253 bytes, one label", entries: []string{strings.Repeat("a", 253)}, cut: true},
{title: "two entries", entries: []string{"www.youtube.com", "evil.example.com"}, cut: true},
{title: "two entries, first empty", entries: []string{"", "www.youtube.com"}, cut: false},
{title: "trailing dot", entries: []string{"youtube.com."}, cut: true},
} {
t.Run(tc.title, func(t *testing.T) {
t.Parallel()
hello := buildClientHello(t, tc.entries...)
for _, mode := range writeModes {
for i := 0; i < 20; i++ {
out := &fakeConn{}
n, err := NewConn(out, t.Context(), mode.splitPacket, mode.splitRecord, time.Nanosecond).Write(hello)
require.NoError(t, err, mode.name)
require.Equal(t, len(hello), n, mode.name)
payloads, cuts := segments(t, hello, out.writes, mode.recordFraming)
require.Equal(t, tc.cut, len(cuts) > 0, "%s: expected cut=%v, got %d cut(s)", mode.name, tc.cut, len(cuts))
require.LessOrEqual(t, len(cuts), mode.budget, mode.name)
var joined []byte
for _, p := range payloads {
require.NotEmpty(t, p)
joined = append(joined, p...)
}
want := hello
if mode.recordFraming {
want = hello[recordLayerHeaderLen:]
}
require.Equal(t, want, joined, "%s run %d", mode.name, i)
}
}
})
}
}
// TestWriteCutsTheFirstEntryOfTheServerNameList: with two entries the cut must
// land inside "youtube" of the FIRST one. Planning over the whole remainder of
// the extension would hand the public suffix list a string that is not a name
// and put the cut somewhere else entirely.
func TestWriteCutsTheFirstEntryOfTheServerNameList(t *testing.T) {
t.Parallel()
hello := buildClientHello(t, "www.youtube.com", "cdn-static-assets.example.com")
sni := IndexTLSServerName(hello)
start := sni.Index + strings.Index("www.youtube.com", "youtube")
end := start + len("youtube")
for i := 0; i < 200; i++ {
out := &fakeConn{}
_, err := NewConn(out, t.Context(), true, false, time.Nanosecond).Write(hello)
require.NoError(t, err)
_, cuts := segments(t, hello, out.writes, false)
require.Len(t, cuts, 1)
require.Greater(t, cuts[0], start, "run %d: cut at %d is outside youtube [%d,%d)", i, cuts[0], start, end)
require.Less(t, cuts[0], end, "run %d: cut at %d is outside youtube [%d,%d)", i, cuts[0], start, end)
}
}
// TestWriteWithoutSNIIsUntouched: the fast path must stay a straight pass, and
// the second and later writes must never be re-planned.
func TestWriteWithoutSNIIsUntouched(t *testing.T) {
t.Parallel()
payload := []byte("not a tls record at all")
out := &fakeConn{}
conn := NewConn(out, t.Context(), true, true, time.Nanosecond)
n, err := conn.Write(payload)
require.NoError(t, err)
require.Equal(t, len(payload), n)
require.Len(t, out.writes, 1)
require.Equal(t, payload, out.writes[0])
hello := clientHelloFor(t, "www.youtube.com")
n, err = conn.Write(hello)
require.NoError(t, err)
require.Equal(t, len(hello), n)
require.Len(t, out.writes, 2, "a ClientHello after the first write must not be fragmented")
require.Equal(t, hello, out.writes[1])
}
+141
View File
@@ -0,0 +1,141 @@
// lx:begin health-board
package urltest
import (
"strconv"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/adapter"
)
// captureEvictions swaps the eviction notice sink for the duration of a test and
// returns a func that reads back everything reported.
func captureEvictions(t *testing.T) func() []string {
t.Helper()
var (
mu sync.Mutex
msgs []string
)
orig := boardEvictionLog
boardEvictionLog = func(m string) {
mu.Lock()
msgs = append(msgs, m)
mu.Unlock()
}
t.Cleanup(func() { boardEvictionLog = orig })
return func() []string {
mu.Lock()
defer mu.Unlock()
return append([]string(nil), msgs...)
}
}
// TestBoardHoldsAGenerationWithoutEvicting is the "what it holds" half of the
// bound. A live generation on this box is ~1200 tags (≈380 nodes plus their
// per-group egress copies and chain hops); the board must carry that — and a
// second generation's worth of overlap during a subscription rename — with no
// eviction at all, or the ceiling would be silently degrading real health data.
func TestBoardHoldsAGenerationWithoutEvicting(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
const generation = 1200
for gen := 0; gen < 2; gen++ {
for i := 0; i < generation; i++ {
s.StoreURLTestHistory("gen"+strconv.Itoa(gen)+"-node-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 20})
}
}
if got := s.Evicted(); got != 0 {
t.Fatalf("two full generations (%d tags) evicted %d entries; the board must hold them",
2*generation, got)
}
if msgs := read(); len(msgs) != 0 {
t.Fatalf("unexpected eviction notices: %v", msgs)
}
// Everything is still readable.
if s.LoadURLTestHistory("gen0-node-0") == nil {
t.Fatalf("the first tag of the first generation was lost without an eviction")
}
}
// TestBoardEvictsOldestAndSaysSo is the "what happens when it overflows" half.
// Overflow must (a) actually bound the map, (b) drop the LEAST RECENTLY MEASURED
// tags — on this box, exactly the ones no config names any more — and (c) be
// audible: a silent eviction is a health board quietly forgetting nodes it is
// still being asked about.
func TestBoardEvictsOldestAndSaysSo(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
base := time.Now().Add(-24 * time.Hour)
// Stale generation first: measured a day ago, nothing since.
const stale = 1500
for i := 0; i < stale; i++ {
s.StoreURLTestHistory("stale-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: base.Add(time.Duration(i) * time.Millisecond), Delay: 30})
}
if s.Evicted() != 0 {
t.Fatalf("evicted before the ceiling was reached")
}
// Now push past the ceiling with fresh measurements.
for i := 0; i <= maxBoardEntries; i++ {
s.StoreURLTestHistory("fresh-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 15})
}
if got := s.Evicted(); got == 0 {
t.Fatalf("board grew past %d entries without evicting anything — it is still unbounded", maxBoardEntries)
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("board holds %d entries, above the %d ceiling", size, maxBoardEntries)
}
// The day-old generation is what went, not the fresh one.
if s.LoadURLTestHistory("stale-0") != nil {
t.Fatalf("the oldest observation survived while newer ones were dropped")
}
if s.LoadURLTestHistory("fresh-"+strconv.Itoa(maxBoardEntries)) == nil {
t.Fatalf("the newest measurement was evicted")
}
msgs := read()
if len(msgs) == 0 {
t.Fatalf("entries were evicted with no notice — eviction must never be silent")
}
m := msgs[0]
for _, want := range []string{"health board full", "evicted", "re-probed"} {
if !strings.Contains(m, want) {
t.Fatalf("eviction notice %q does not say %q", m, want)
}
}
}
// TestBoardEvictionThroughMarkFailed pins the OTHER write path. MarkFailed is how
// a dead node is recorded, and a flood of dead renamed nodes is exactly the shape
// of the leak — so it has to prune too, not just the success path.
func TestBoardEvictionThroughMarkFailed(t *testing.T) {
captureEvictions(t)
s := NewHistoryStorage()
for i := 0; i <= maxBoardEntries; i++ {
s.MarkFailed("dead-" + strconv.Itoa(i))
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("MarkFailed grew the board to %d, above the %d ceiling", size, maxBoardEntries)
}
if s.Evicted() == 0 {
t.Fatalf("MarkFailed never prunes — the failure path is still unbounded")
}
}
// lx:end health-board
+118
View File
@@ -10,11 +10,128 @@
package urltest
import (
"sort"
"strconv"
"sync"
"time"
"github.com/sagernet/sing-box/adapter"
"github.com/sagernet/sing-box/log"
)
// --- board capacity ---------------------------------------------------------
//
// The board is the one structure in the daemon whose key space is chosen by
// somebody else. Its keys are outbound TAGS, and on this box a tag is a node
// NAME straight out of the subscription — plus the derived per-group egress
// copies ("group-<g>-m<i>-<node>") and per-chain hop copies the probe planner
// creates for the same nodes. Providers rename their nodes freely, so a daily
// subscription refresh introduces a whole new generation of keys, while the
// store itself is pinned to the ENGINE's context (shater/engine.New) and so
// outlives every generation and every Apply — by design, so health survives a
// config change.
//
// Nothing ever removed a key. DeleteURLTestHistory exists but no shater path
// calls it (only daemon/ and clashapi/, which this fork does not run), so the
// map was strictly append-only for the life of the process — and the process is
// expected to live for months.
//
// The arithmetic: ~380 nodes, and a config with a couple of egress-bound groups
// plus a handful of chains puts a LIVE generation at roughly 380 base tags +
// 2x380 group copies + ~100 chain copies ≈ 1200 keys. One new generation per day
// is ~440k keys a year, at ~200 B per entry (map bucket + a tag string that is
// routinely 30-50 B with flag emoji, + a 56 B URLTestHistory) ≈ 88 MB of a
// 512 MB box — spent entirely on nodes that no longer exist.
const (
// maxBoardEntries is the hard ceiling. 4096 is ~3.4 live generations, so the
// board comfortably holds the current config plus the overlap while a
// subscription refresh swaps names, and still costs under a megabyte. A tighter
// bound would start evicting tags the running config actually uses; a looser one
// would stop being a bound in any useful sense.
maxBoardEntries = 4096
// keepBoardEntries is the prune target: drop a quarter at a time so the
// O(n log n) selection is amortised over ~1024 inserts instead of running on
// every probe once the board is full.
keepBoardEntries = 3072
)
// boardEvictionLog reports an eviction. A package var so tests can capture it;
// production leaves it writing to the process log, which under procd is the same
// syslog/logsink stream every other daemon line lands in.
//
// Eviction is NEVER silent. It is not free either: an evicted tag reverts to
// "untested" and its next probe re-measures it, so a board that evicts entries
// belonging to the LIVE config is a board whose ceiling is too low — and the only
// way anyone finds that out is this line.
var boardEvictionLog = func(msg string) { boardLogger().Warn(msg) }
// pruneLocked drops the least-recently-OBSERVED entries when the board exceeds
// maxBoardEntries. "Least recently observed" is max(LastOK, LastFail): the entry
// nothing has measured for the longest is, on this box, precisely a tag that no
// longer exists in any config — a renamed node, a removed group copy, a retired
// chain hop. Caller holds access.
func (s *HistoryStorage) pruneLocked() {
if len(s.delayHistory) <= maxBoardEntries {
return
}
type kv struct {
tag string
seen time.Time
}
all := make([]kv, 0, len(s.delayHistory))
for tag, h := range s.delayHistory {
seen := h.LastOK
if h.LastFail.After(seen) {
seen = h.LastFail
}
all = append(all, kv{tag, seen})
}
sort.Slice(all, func(i, j int) bool { return all[i].seen.Before(all[j].seen) })
drop := len(all) - keepBoardEntries
var oldest time.Time
for i := 0; i < drop; i++ {
if i == 0 {
oldest = all[i].seen
}
delete(s.delayHistory, all[i].tag)
}
s.evicted += uint64(drop)
msg := "urltest: health board full (" + strconv.Itoa(maxBoardEntries) + " tags) — evicted " +
strconv.Itoa(drop) + " least-recently-measured entries (" + strconv.FormatUint(s.evicted, 10) +
" total since start); they revert to untested and will be re-probed"
if !oldest.IsZero() {
msg += "; oldest observation was " + time.Since(oldest).Truncate(time.Second).String() + " ago"
}
boardEvictionLog(msg)
}
// Evicted reports how many entries the capacity bound has dropped since the store
// was created. Nonzero means the board reached maxBoardEntries at least once.
func (s *HistoryStorage) Evicted() uint64 {
if s == nil {
return 0
}
s.access.RLock()
defer s.access.RUnlock()
return s.evicted
}
// boardLogger is the process-wide fallback logger for eviction notices. The store
// is built from a plain constructor with no logger in sight (box.New, the daemon,
// shater/engine all call NewHistoryStorage()), so rather than change that
// signature everywhere the notice goes to the standard logger — which on the
// router is the daemon's own stderr, i.e. the same sink logsink owns.
var (
boardLogOnce sync.Once
boardLog log.ContextLogger
)
func boardLogger() log.ContextLogger {
boardLogOnce.Do(func() { boardLog = log.StdLogger() })
return boardLog
}
// HealthVerdict classifies a stored history entry at read time.
type HealthVerdict int
@@ -54,6 +171,7 @@ func (s *HistoryStorage) MarkFailed(tag string) {
updated.Delay = previous.Delay
}
s.delayHistory[tag] = updated
s.pruneLocked()
s.notifyUpdated()
s.access.Unlock()
}
+9
View File
@@ -21,6 +21,10 @@ type HistoryStorage struct {
access sync.RWMutex
delayHistory map[string]*adapter.URLTestHistory
updateHooks []*observable.Subscriber[struct{}]
// evicted counts entries dropped by the capacity bound (board_lx.go). The map
// is keyed by outbound tags chosen by a subscription provider, so it needs a
// ceiling; see the comment on maxBoardEntries.
evicted uint64
}
func NewHistoryStorage() *HistoryStorage {
@@ -71,6 +75,11 @@ func (s *HistoryStorage) StoreURLTestHistory(tag string, history *adapter.URLTes
}
// lx:end health-board
s.delayHistory[tag] = history
// lx:begin health-board — the map is keyed by provider-chosen tags and the
// store outlives every engine generation, so it must bound itself here: no
// shater path ever calls DeleteURLTestHistory. See maxBoardEntries.
s.pruneLocked()
// lx:end health-board
s.notifyUpdated()
s.access.Unlock()
}
+90 -3
View File
@@ -10,6 +10,7 @@ import (
"net/url"
"strconv"
"sync"
"sync/atomic"
"time"
"github.com/sagernet/sing-box/adapter"
@@ -171,6 +172,73 @@ func (t *HTTPSTransport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
return response, nil
}
// requestBuffer owns the pooled buffer that backs one DoH query.
//
// Both transports behind HTTPSTransportWrapper write the request body on a
// goroutine of their own and return from RoundTrip as soon as the response
// HEADERS arrive: net/http's write loop is still copying out of the body a
// bufferful at a time (4 KiB of write buffer, or io.Copy's 32 KiB once it hands
// the body to the connection), and http2's writeRequestBody has read only the
// first max-frame-size bytes of it. Returning the buffer to the pool at that
// point handed live memory to the next caller while the query was still going
// out — everything past that first copy left the router as whatever that caller
// had written there. A data race, and a memory-disclosure primitive aimed at the
// resolver. Measured, not reasoned: with the write parked mid-query the bytes on
// the wire diverge from the bytes we packed at exactly one copy buffer in.
//
// Ownership is counted rather than handed over once, because a retry holds two
// bodies at a time and the two transports order that differently:
// http.Transport.rewindBody CLOSES the old body before asking GetBody for a
// new one, while http2's shouldRetryRequest asks GetBody first and closes the
// old body on a goroutine. exchange keeps a count of its own until RoundTrip
// returns — the only window in which either can call GetBody — so neither
// ordering can free the buffer under the other. If a transport ever fails to
// close a body, the count never reaches zero and the buffer is simply not
// reused: garbage, not corruption.
type requestBuffer struct {
buffer *buf.Buffer
raw []byte
refs atomic.Int32
}
func newRequestBuffer(buffer *buf.Buffer, raw []byte) *requestBuffer {
holder := &requestBuffer{buffer: buffer, raw: raw}
holder.refs.Store(1)
return holder
}
// body hands out a reader over the packed query as one more owner. It refuses
// once the buffer is back in the pool, so a late caller gets an error instead
// of a reader over memory that now belongs to somebody else.
func (b *requestBuffer) body() (*pooledRequestBody, bool) {
for {
refs := b.refs.Load()
if refs < 1 {
return nil, false
}
if b.refs.CompareAndSwap(refs, refs+1) {
return &pooledRequestBody{Reader: bytes.NewReader(b.raw), owner: b}, true
}
}
}
func (b *requestBuffer) release() {
if b.refs.Add(-1) == 0 {
b.buffer.Release()
}
}
type pooledRequestBody struct {
*bytes.Reader
owner *requestBuffer
closeOne sync.Once
}
func (b *pooledRequestBody) Close() error {
b.closeOne.Do(b.owner.release)
return nil
}
func (t *HTTPSTransport) exchange(ctx context.Context, message *mDNS.Msg) (*mDNS.Msg, error) {
exMessage := *message
exMessage.Id = 0
@@ -181,11 +249,31 @@ func (t *HTTPSTransport) exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
requestBuffer.Release()
return nil, err
}
request, err := http.NewRequestWithContext(ctx, http.MethodPost, t.destination.String(), bytes.NewReader(rawMessage))
queryBuffer := newRequestBuffer(requestBuffer, rawMessage)
// Drops the count exchange holds once RoundTrip is done with the request;
// the bodies handed to the transport keep their own until it closes them.
defer queryBuffer.release()
requestBody, _ := queryBuffer.body() // cannot fail: the count above is ours
request, err := http.NewRequestWithContext(ctx, http.MethodPost, t.destination.String(), requestBody)
if err != nil {
requestBuffer.Release()
requestBody.Close()
return nil, err
}
// http.NewRequestWithContext infers both only for the body types it knows,
// and pooledRequestBody is not one of them. Upstream got them for free from
// *bytes.Reader; GetBody is what lets a POST be replayed when a pooled
// connection turns out to have been closed under us. Being unknown to
// net/http also costs one packet on the HTTP/1.1 leg: isKnownInMemoryReader
// no longer recognises the body, so the request headers are flushed before
// the query instead of travelling with it.
request.ContentLength = int64(len(rawMessage))
request.GetBody = func() (io.ReadCloser, error) {
retryBody, ok := queryBuffer.body()
if !ok {
return nil, E.New("DoH request buffer already released")
}
return retryBody, nil
}
request.Header = t.headers.Clone()
request.Header.Set("Content-Type", MimeType)
request.Header.Set("Accept", MimeType)
@@ -193,7 +281,6 @@ func (t *HTTPSTransport) exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
currentTransport := t.transport
t.transportAccess.Unlock()
response, err := currentTransport.RoundTrip(request)
requestBuffer.Release()
if err != nil {
return nil, err
}
@@ -0,0 +1,541 @@
package transport
import (
"bytes"
"context"
"errors"
"io"
"net"
"net/http"
"net/http/httptest"
"net/url"
"os"
"strconv"
"sync"
"sync/atomic"
"testing"
"time"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
mDNS "github.com/miekg/dns"
"golang.org/x/net/http2"
)
// The request body of a DoH query is backed by a POOLED buffer. Neither
// transport behind HTTPSTransportWrapper is done with that body when RoundTrip
// returns: net/http hands the request to a write loop of its own and returns as
// soon as the response HEADERS have been read, and golang.org/x/net/http2 writes
// the body on the goroutine that runs writeRequest while roundTrip waits on
// respHeaderRecv. Returning the buffer to the pool at that point hands live
// memory to the next caller while the query is still being written to the wire,
// and what goes out is whatever that next caller put there.
//
// Both tests below force a window that is normally microseconds wide to stay
// open, and drain the pool while it is open:
//
// - HTTP/1.1: the client connection stops accepting writes past the request
// headers, so net/http's write loop is parked having copied only the first
// io.Copy buffer (32 KiB) of the query.
// - HTTP/2: the server pins a 1 KiB stream receive window and does not read
// the body, so writeRequestBody is parked in awaitFlowControl having copied
// only the first max-frame-size bytes of the query.
//
// In both, the server sends the response HEADERS first and withholds the
// response BODY until the pool has been drained, so Exchange has returned from
// RoundTrip — and released the buffer, on the broken build — while the query is
// still going out.
//
// Both queries are padded past the transport's copy buffer on purpose. Below it
// the transport lifts the whole query out of the pooled buffer in a single Read
// that RACES the release rather than provably following it, and a test built on
// that race would be a coin toss. The ownership defect is the same at every
// size; only its deterministic proof needs the padding.
const (
// Past io.Copy's 32 KiB buffer, which is the granularity net/http moves a
// request body at (persistConnWriter.ReadFrom -> io.Copy), and still inside
// buf.MaxPooledBufferSize so the buffer really comes from the pool.
httpsH1PaddedQuerySize = 40000
// Past http2's max frame size, which is how much of the body
// writeRequestBody lifts into its scratch buffer per round.
httpsH2PaddedQuerySize = 20000
// Pinned on the HTTP/2 server so the client cannot write the whole body
// before the response headers come back.
httpsPinnedStreamWindow = 1024
// Pinned too: Go's HTTP/2 server advertises a 1 MiB max frame size by
// default, and the client sizes its body-copy buffer from that — with the
// default it would slurp a 20 KB query in one Read and the divergence would
// be hidden by the copy size rather than absent. 16384 is the protocol
// minimum and what real resolvers advertise.
httpsPinnedMaxFrameSize = 16384
// How many times the HTTP/2 scenario is repeated; see the test.
httpsH2Rounds = 8
// How long to wait after the response headers before draining the pool, so
// that Exchange has certainly returned from RoundTrip.
httpsReleaseSettleDelay = 200 * time.Millisecond
)
// httpsPaddedQuery returns a query and the exact bytes HTTPSTransport.exchange
// packs for it.
func httpsPaddedQuery(t *testing.T, padding int) (*mDNS.Msg, []byte) {
t.Helper()
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
opt := new(mDNS.OPT)
opt.Hdr.Name = "."
opt.Hdr.Rrtype = mDNS.TypeOPT
opt.Option = append(opt.Option, &mDNS.EDNS0_PADDING{Padding: make([]byte, padding)})
message.Extra = append(message.Extra, opt)
onWire := *message
onWire.Id = 0
onWire.Compress = true
expected, err := onWire.Pack()
if err != nil {
t.Fatal(err)
}
return message, expected
}
func httpsTestReply(t *testing.T) []byte {
t.Helper()
query := new(mDNS.Msg)
query.SetQuestion("example.com.", mDNS.TypeA)
response := new(mDNS.Msg)
response.SetReply(query)
raw, err := response.Pack()
if err != nil {
t.Fatal(err)
}
return raw
}
// httpsPoisonPool takes buffers of one size class out of the pool and fills them
// with a pattern no DNS message contains. They are returned, not released: the
// caller holds them so nothing can hand them back while the check runs.
func httpsPoisonPool(size int, count int) []*buf.Buffer {
poison := make([]*buf.Buffer, 0, count)
for range count {
buffer := buf.NewSize(size)
poison = append(poison, buffer)
free := buffer.FreeBytes()
for i := range free {
free[i] = 0xEE
}
}
return poison
}
func httpsReleaseAll(buffers []*buf.Buffer) {
for _, buffer := range buffers {
buffer.Release()
}
}
// httpsRequirePoisonReachesReleasedBuffer is the CONTROL for the tests below. A
// clean result there means nothing unless this instrument is shown to be able to
// produce a dirty one: it must be true that a buffer released while its bytes
// are still referenced comes back out of the pool and gets overwritten. If that
// stops holding — a different allocator, a pool that zeroes, a size class that
// is not pooled at all — the tests below would go green on broken code.
//
// Retried, because under -race sync.Pool.Put drops one object in four on
// purpose. That same dice roll is why the checks below are 3-in-4 detectors
// under -race and certainties without it; it can only make a broken build look
// clean, never a clean build look broken.
func httpsRequirePoisonReachesReleasedBuffer(t *testing.T, size int, pattern []byte) {
t.Helper()
for range 32 {
control := buf.NewSize(size)
free := control.FreeBytes()
if len(free) < len(pattern) {
t.Fatalf("control failed: a %d-byte buffer came back %d bytes long", size, len(free))
}
copy(free, pattern)
alias := free[:len(pattern)]
control.Release()
held := httpsPoisonPool(size, 8)
poisoned := !bytes.Equal(alias, pattern)
httpsReleaseAll(held)
if poisoned {
return
}
}
t.Fatal("control failed: poisoning the pool never touched a released buffer, so a clean result below would prove nothing")
}
// httpsTestDialer hands HTTPSTransportWrapper a connection to a local test
// server, optionally wrapped.
type httpsTestDialer struct {
target string
wrap func(net.Conn) net.Conn
access sync.Mutex
conns []net.Conn
}
func (d *httpsTestDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := (&net.Dialer{}).DialContext(ctx, "tcp", d.target)
if err != nil {
return nil, err
}
var wrapped net.Conn = conn
if d.wrap != nil {
wrapped = d.wrap(conn)
}
d.access.Lock()
d.conns = append(d.conns, conn)
d.access.Unlock()
return wrapped, nil
}
func (d *httpsTestDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return nil, os.ErrInvalid
}
func (d *httpsTestDialer) closeAll() {
d.access.Lock()
defer d.access.Unlock()
for _, conn := range d.conns {
conn.Close()
}
}
// httpsGatedConn stops accepting writes once limit bytes have gone out, until
// the gate is opened. HTTP/1.1 has no flow-control knob to park the writer with,
// so the connection provides one.
type httpsGatedConn struct {
net.Conn
limit int64
written atomic.Int64
gate chan struct{}
}
func (c *httpsGatedConn) Write(p []byte) (int, error) {
if c.written.Load()+int64(len(p)) > c.limit {
select {
case <-c.gate:
case <-time.After(30 * time.Second):
return 0, errors.New("gated conn: nobody opened the gate")
}
}
n, err := c.Conn.Write(p)
c.written.Add(int64(n))
return n, err
}
// httpsSlowServer is the handler both tests share: response HEADERS first, then
// nothing until the pool has been drained, then the request body, then the
// response body.
type httpsSlowServer struct {
reply []byte
served atomic.Int32
warmups int32
headersSent chan struct{}
bodyGate chan struct{}
received chan []byte
readErr chan error
}
func newHTTPSSlowServer(reply []byte) *httpsSlowServer {
return &httpsSlowServer{
reply: reply,
headersSent: make(chan struct{}, 1),
bodyGate: make(chan struct{}),
received: make(chan []byte, 1),
readErr: make(chan error, 1),
}
}
func (s *httpsSlowServer) ServeHTTP(writer http.ResponseWriter, request *http.Request) {
if s.served.Add(1) <= s.warmups {
// Warm-up: answer normally, so the connection is established and the
// client has applied the server's SETTINGS before the query that
// matters goes out.
io.Copy(io.Discard, request.Body)
writer.Header().Set("Content-Type", MimeType)
writer.Header().Set("Content-Length", strconv.Itoa(len(s.reply)))
writer.Write(s.reply)
return
}
// Without this, net/http's HTTP/1.1 server drains up to 256 KB of the
// request body before it will write response headers, precisely so that a
// half-duplex client cannot deadlock. That would consume the query before
// the client is anywhere near done sending it, and there would be nothing
// left in flight to catch. Full duplex is how a resolver that answers from
// cache before reading the whole query behaves; HTTP/2 is full duplex
// already and returns an error here, which is fine.
http.NewResponseController(writer).EnableFullDuplex()
writer.Header().Set("Content-Type", MimeType)
// Content-Length matters: without it Exchange falls into io.ReadAll and
// waits for the end of the response, which this handler is about to
// withhold on purpose.
writer.Header().Set("Content-Length", strconv.Itoa(len(s.reply)))
writer.WriteHeader(http.StatusOK)
writer.(http.Flusher).Flush()
s.headersSent <- struct{}{}
// A real resolver would be reading the query by now. Withholding it is what
// keeps the client parked mid-body while the pool is drained.
<-s.bodyGate
body, err := io.ReadAll(request.Body)
s.readErr <- err
s.received <- body
writer.Write(s.reply)
}
// drainPoolOnceHeadersAreOut waits for the response headers, gives Exchange time
// to return from RoundTrip, drains the size class the query buffer came from —
// on this goroutine, so a buffer released on the way out lands in our hands and
// not somewhere harmless — and only then lets the server read the query.
func (s *httpsSlowServer) drainPoolOnceHeadersAreOut(bufferSize int) <-chan []*buf.Buffer {
poisoned := make(chan []*buf.Buffer, 1)
go func() {
<-s.headersSent
time.Sleep(httpsReleaseSettleDelay)
poisoned <- httpsPoisonPool(bufferSize, 32)
close(s.bodyGate)
}()
return poisoned
}
func (s *httpsSlowServer) requireQueryOnWire(t *testing.T, expected []byte) {
t.Helper()
var sent []byte
select {
case sent = <-s.received:
case <-time.After(30 * time.Second):
t.Fatal("the server never received the request body")
}
if err := <-s.readErr; err != nil {
t.Fatal("reading the request body: ", err)
}
if bytes.Equal(sent, expected) {
return
}
firstDiff := -1
for i := 0; i < len(sent) && i < len(expected); i++ {
if sent[i] != expected[i] {
firstDiff = i
break
}
}
t.Fatalf("the query on the wire is not the query we packed: %d of %d bytes received, first difference at offset %d — "+
"the pooled request buffer was reused while the transport was still reading it", len(sent), len(expected), firstDiff)
}
// TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP1 proves that the query an
// HTTP/1.1 resolver receives is the query we asked to send, even when the pool
// is drained the instant the response headers arrive.
func TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP1(t *testing.T) {
message, expected := httpsPaddedQuery(t, httpsH1PaddedQuerySize)
bufferSize := 1 + message.Len()
httpsRequirePoisonReachesReleasedBuffer(t, bufferSize, expected)
handler := newHTTPSSlowServer(httpsTestReply(t))
server := httptest.NewServer(handler)
t.Cleanup(server.Close)
dialer := &httpsTestDialer{
target: server.Listener.Addr().String(),
wrap: func(conn net.Conn) net.Conn {
// One 4 KiB flush of net/http's write buffer gets through, which is
// what carries the request headers to the server, and the write loop
// parks on the next one — still holding the query.
return &httpsGatedConn{Conn: conn, limit: 4096, gate: handler.bodyGate}
},
}
t.Cleanup(dialer.closeAll)
// Scheme http puts HTTPSTransportWrapper on its HTTP/1.1 leg, the one it
// also falls back to whenever a resolver does not negotiate h2.
destination := &url.URL{Scheme: "http", Host: "doh.invalid", Path: "/dns-query"}
dnsTransport := &HTTPSTransport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTPS, "test-doh-h1", nil),
logger: logger.NOP(),
dialer: dialer,
destination: destination,
headers: http.Header{},
transport: NewHTTPSTransportWrapper(dialer, M.ParseSocksaddr(server.Listener.Addr().String()), destination),
}
t.Cleanup(func() { dnsTransport.Close() })
poisoned := handler.drainPoolOnceHeadersAreOut(bufferSize)
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, message); err != nil {
t.Fatal("exchange: ", err)
}
defer httpsReleaseAll(<-poisoned)
handler.requireQueryOnWire(t, expected)
}
// TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP2 does the same over h2,
// the leg every resolver that speaks HTTP/2 lands on.
func TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP2(t *testing.T) {
message, expected := httpsPaddedQuery(t, httpsH2PaddedQuerySize)
bufferSize := 1 + message.Len()
httpsRequirePoisonReachesReleasedBuffer(t, bufferSize, expected)
// Repeated because a buffer released on the goroutine running Exchange
// usually lands in that P's private sync.Pool slot, which the goroutine
// draining the pool cannot steal: one round catches a broken build about
// half the time, eight catch it better than 99 times in 100. Every round
// must come back clean.
for round := range httpsH2Rounds {
if !t.Run(strconv.Itoa(round), func(t *testing.T) {
httpsH2Round(t, message, expected, bufferSize)
}) {
return
}
}
}
func httpsH2Round(t *testing.T, message *mDNS.Msg, expected []byte, bufferSize int) {
handler := newHTTPSSlowServer(httpsTestReply(t))
// x/net/http2 may put the first request on the wire before it has applied
// the server's SETTINGS, and would then overrun the 1 KiB window this test
// pins and be reset with FLOW_CONTROL_ERROR. One small query first settles
// that: reading its response proves the SETTINGS frame ahead of it was
// processed.
handler.warmups = 1
listener, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { listener.Close() })
h2server := &http2.Server{
MaxUploadBufferPerStream: httpsPinnedStreamWindow,
MaxReadFrameSize: httpsPinnedMaxFrameSize,
}
go func() {
for {
conn, acceptErr := listener.Accept()
if acceptErr != nil {
return
}
go h2server.ServeConn(conn, &http2.ServeConnOpts{Handler: handler})
}
}()
dialer := &httpsTestDialer{target: listener.Addr().String()}
t.Cleanup(dialer.closeAll)
// Scheme https keeps HTTPSTransportWrapper on its h2 leg. The dialer hands
// back a plain connection, which x/net/http2 speaks prior-knowledge h2 over;
// TLS adds nothing this test is about.
destination := &url.URL{Scheme: "https", Host: "doh.invalid", Path: "/dns-query"}
dnsTransport := &HTTPSTransport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTPS, "test-doh-h2", nil),
logger: logger.NOP(),
dialer: dialer,
destination: destination,
headers: http.Header{},
transport: NewHTTPSTransportWrapper(dialer, M.ParseSocksaddr(listener.Addr().String()), destination),
}
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
warmup := new(mDNS.Msg)
warmup.SetQuestion("warmup.invalid.", mDNS.TypeA)
if _, err = dnsTransport.Exchange(ctx, warmup); err != nil {
t.Fatal("warm-up exchange: ", err)
}
poisoned := handler.drainPoolOnceHeadersAreOut(bufferSize)
if _, err = dnsTransport.Exchange(ctx, message); err != nil {
t.Fatal("exchange: ", err)
}
defer httpsReleaseAll(<-poisoned)
handler.requireQueryOnWire(t, expected)
}
// TestHTTPSRequestBufferSurvivesRewind covers the second owner a retry creates.
// net/http rewinds a dead connection's request by CLOSING the body it has and
// then asking GetBody for another one (rewindBody), while x/net/http2 asks
// GetBody first and closes the old body on a goroutine (shouldRetryRequest,
// closeReqBodyLocked). Either ordering frees the buffer under the retry if the
// first Close is what returns it to the pool, and the retry then sends whatever
// the next pool user wrote — the same disclosure, one attempt later.
func TestHTTPSRequestBufferSurvivesRewind(t *testing.T) {
message, expected := httpsPaddedQuery(t, httpsH2PaddedQuerySize)
bufferSize := 1 + message.Len()
httpsRequirePoisonReachesReleasedBuffer(t, bufferSize, expected)
exMessage := *message
exMessage.Id = 0
exMessage.Compress = true
requestBuffer := buf.NewSize(bufferSize)
rawMessage, err := exMessage.PackBuffer(requestBuffer.FreeBytes())
if err != nil {
t.Fatal(err)
}
queryBuffer := newRequestBuffer(requestBuffer, rawMessage)
defer queryBuffer.release()
first, ok := queryBuffer.body()
if !ok {
t.Fatal("the first body was refused while exchange still holds the buffer")
}
// The transport got some of the query out before the connection turned out
// to be dead, then closed the body.
if _, err = io.CopyN(io.Discard, first, 128); err != nil {
t.Fatal(err)
}
first.Close()
// GetBody, as the retry would call it.
second, ok := queryBuffer.body()
if !ok {
t.Fatal("GetBody was refused after the first body was closed: the retry has no query left to send")
}
poison := httpsPoisonPool(bufferSize, 32)
defer httpsReleaseAll(poison)
retried, err := io.ReadAll(second)
if err != nil {
t.Fatal(err)
}
if !bytes.Equal(retried, expected) {
firstDiff := -1
for i := 0; i < len(retried) && i < len(expected); i++ {
if retried[i] != expected[i] {
firstDiff = i
break
}
}
t.Fatalf("the retried query is not the query we packed: %d of %d bytes, first difference at offset %d — "+
"closing the first body returned the buffer to the pool while the retry still needed it", len(retried), len(expected), firstDiff)
}
second.Close()
}
// TestHTTPSRequestBufferRefusesBodyAfterRelease pins the recoverable end of the
// contract: once the buffer really is back in the pool, GetBody must hand out an
// error rather than a reader over memory that now belongs to somebody else.
func TestHTTPSRequestBufferRefusesBodyAfterRelease(t *testing.T) {
requestBuffer := buf.NewSize(64)
rawMessage := requestBuffer.FreeBytes()[:8]
queryBuffer := newRequestBuffer(requestBuffer, rawMessage)
body, ok := queryBuffer.body()
if !ok {
t.Fatal("the first body was refused while the caller still holds the buffer")
}
body.Close()
body.Close() // http3 and net/http both manage to close a body twice
queryBuffer.release()
if _, ok = queryBuffer.body(); ok {
t.Fatal("a body was handed out over a buffer that is already back in the pool")
}
}
+29 -5
View File
@@ -126,6 +126,12 @@ func (t *HTTP3Transport) newTransport() *http3.Transport {
conn.Close()
return nil, dialErr
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-quicConn.Context().Done()
conn.Close()
}()
return quicConn, nil
},
TLSClientConfig: t.tlsConfig,
@@ -156,15 +162,34 @@ func (t *HTTP3Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
exMessage := *message
exMessage.Id = 0
exMessage.Compress = true
requestBuffer := buf.NewSize(1 + message.Len())
rawMessage, err := exMessage.PackBuffer(requestBuffer.FreeBytes())
// NOT a pooled buffer, deliberately — the request body must own memory this
// transport can never hand back.
//
// quic-go writes the request body on a goroutine of its own (http3's
// doRequest spawns it and goes on to block in ReadResponse), and NOTHING ever
// joins that goroutine. On the success path sendRequestBody closes the body
// when it is finished, but on every error path RoundTripOpt closes it as soon
// as doRequest returns — and doRequest waits only on the request-cancellation
// watchdog, not on the writer. So there is no moment at which this code can
// know the body is no longer being read, and therefore no moment at which it
// may return a pooled buffer. Releasing on Close looks like an ownership
// handoff and is not one.
//
// Owning it costs nothing here, measured rather than assumed: for a typical
// query (a 36-byte name, A record) Pack is 87 ns/op at 64 B and 1 alloc,
// against 108 ns/op at 64 B and 1 alloc for packing into a pooled buffer. The
// pool never avoided an allocation on this path — buf.NewSize allocates the
// Buffer struct itself, the same 64 bytes the message needs — it only added
// Get/Put on top. This path is hot in queries, not in bytes.
//
// The response buffer below stays pooled: it is read to completion and
// unpacked before Exchange returns, and nothing outlives it.
rawMessage, err := exMessage.Pack()
if err != nil {
requestBuffer.Release()
return nil, err
}
request, err := http.NewRequestWithContext(ctx, http.MethodPost, t.destination.String(), bytes.NewReader(rawMessage))
if err != nil {
requestBuffer.Release()
return nil, err
}
request.Header = t.headers.Clone()
@@ -174,7 +199,6 @@ func (t *HTTP3Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
currentTransport := t.transport
t.transportAccess.Unlock()
response, err := currentTransport.RoundTrip(request)
requestBuffer.Release()
if err != nil {
return nil, err
}
@@ -0,0 +1,426 @@
package quic
import (
"bytes"
"context"
"crypto/rand"
"crypto/tls"
"io"
"net"
"net/http"
"net/url"
"strconv"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing-box/dns/transport"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
mDNS "github.com/miekg/dns"
)
// The request body of a DoH3 query used to be backed by a POOLED buffer. quic-go
// sends that body on a goroutine of its own which outlives RoundTrip (http3's
// doRequest spawns it and returns as soon as the response HEADERS arrive), and
// NOTHING joins that goroutine, so there is no moment at which the transport may
// hand the buffer back.
//
// Two tests, because the two paths are observable in different ways.
//
// - On the SUCCESS path the body keeps flowing, so the damage is visible on the
// wire: TestHTTP3ExchangeRequestBufferOutlivesRoundTrip pins a 2 KB server
// stream window and answers before reading the body, so the client is still
// writing when Exchange returns, and compares what the server received.
//
// - On the FAILURE and CANCELLATION paths the damage is not visible on the wire
// at all: every ReadResponse error in quic-go calls str.CancelWrite BEFORE
// RoundTripOpt closes the body, so whatever the writer reads afterwards is
// thrown at a dead stream. What is left is a read of memory that belongs to
// somebody else. TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory therefore
// pins the CAUSE instead of the symptom: the bytes of a query must never end
// up in a buffer this transport can return to the pool.
const (
// Big enough to need more than one 8 KiB read out of the request body
// (http3's bodyCopyBufferSize), small enough to still come from the pool
// (buf.MaxPooledBufferSize).
paddedQuerySize = 20000
// Pinned on the server so the client cannot write the whole body before the
// response comes back.
pinnedStreamWindow = 2048
// Padding for the marked query of the ownership test. Only has to be
// distinctive and pooled, not large.
markedQueryPadding = 4096
markedQueryNeedle = 64
// How deep to drain a size class when looking for the needle.
poolScanDepth = 64
// How many times a CONTROL may repeat before it gives up.
//
// Both controls in this file assert the same thing — a buffer released while
// its bytes are still referenced comes back out of the pool — and under
// `-race` that is a DICE ROLL, not a certainty: sync.Pool.Put drops one
// object in four on purpose (runtime_randn(4) == 0, sync/pool.go). Measured
// in golang:1.26 with `go test -race -count=60`: the single-attempt control
// failed 18 times out of 60, i.e. the gate's -race pass had a ~30% chance of
// going red on a tree with nothing wrong with it.
//
// A retry is the honest repair rather than a papering-over, because the
// control's claim is EXISTENTIAL — "this instrument is able to find a
// released, still-referenced buffer" — and one success proves it. It is not
// an average over attempts, so nothing is diluted by taking more than one.
// 32 attempts leave a (1/4)^32 chance of a false alarm.
//
// What this does NOT do, said plainly: it does not make the VERDICT below
// certain under -race. The same 1-in-4 drop means a scan that comes back
// clean has a 1-in-4 chance of being clean because the pool threw the
// evidence away. That direction is the safe one — it can only let a broken
// build look clean, never make a clean build look broken — and the -race
// pass is not the only one that runs this test: [2/7] of scripts/run-tests.sh
// runs the same file WITHOUT -race, where both the control and the verdict
// are certainties.
controlAttempts = 32
)
func paddedQuery(t *testing.T) (*mDNS.Msg, []byte) {
t.Helper()
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
opt := new(mDNS.OPT)
opt.Hdr.Name = "."
opt.Hdr.Rrtype = mDNS.TypeOPT
opt.Option = append(opt.Option, &mDNS.EDNS0_PADDING{Padding: make([]byte, paddedQuerySize)})
message.Extra = append(message.Extra, opt)
// Exactly what HTTP3Transport.Exchange puts on the wire.
onWire := *message
onWire.Id = 0
onWire.Compress = true
expected, err := onWire.Pack()
if err != nil {
t.Fatal(err)
}
return message, expected
}
// poisonPool takes buffers of one size class out of the pool and fills them with
// a pattern no DNS message contains. The buffers are returned, not released: the
// caller holds them so nothing can hand them back while the check runs.
func poisonPool(size int, count int) []*buf.Buffer {
poison := make([]*buf.Buffer, 0, count)
for range count {
buffer := buf.NewSize(size)
poison = append(poison, buffer)
free := buffer.FreeBytes()
for i := range free {
free[i] = 0xEE
}
}
return poison
}
func releaseAll(buffers []*buf.Buffer) {
for _, buffer := range buffers {
buffer.Release()
}
}
// requirePoisonReachesReleasedBuffer is the CONTROL for the test below. A clean
// result there means nothing unless this instrument is shown to be able to
// produce a dirty one: it must be true that a buffer released while its bytes
// are still referenced comes back out of the pool and gets overwritten. If this
// stops holding — a different allocator, a pool that zeroes, a size class that
// is not pooled at all — the test below would go green on broken code.
//
// Retried, because under -race sync.Pool.Put drops one object in four on
// purpose. That same dice roll is why the check below is a 3-in-4 detector under
// -race and a certainty without it; it can only make a broken build look clean,
// never a clean build look broken. See controlAttempts.
func requirePoisonReachesReleasedBuffer(t *testing.T, size int, pattern []byte) {
t.Helper()
for range controlAttempts {
control := buf.NewSize(size)
free := control.FreeBytes()
if len(free) < len(pattern) {
t.Fatalf("control failed: a %d-byte buffer came back %d bytes long", size, len(free))
}
copy(free, pattern)
alias := free[:len(pattern)]
control.Release()
held := poisonPool(size, 8)
poisoned := !bytes.Equal(alias, pattern)
releaseAll(held)
if poisoned {
return
}
}
t.Fatal("control failed: poisoning the pool never touched a released buffer, so a clean result below would prove nothing")
}
// TestHTTP3ExchangeRequestBufferOutlivesRoundTrip proves that the query the
// server receives is the query we asked to send, even when the pool is drained
// the instant Exchange returns.
func TestHTTP3ExchangeRequestBufferOutlivesRoundTrip(t *testing.T) {
message, expected := paddedQuery(t)
bufferSize := 1 + message.Len()
requirePoisonReachesReleasedBuffer(t, bufferSize, expected)
drainGate := make(chan struct{})
received := make(chan []byte, 1)
mux := http.NewServeMux()
mux.HandleFunc("/dns-query", func(writer http.ResponseWriter, request *http.Request) {
// Answer BEFORE reading the request body. A real resolver would not, but
// any peer, middlebox or loss pattern that delays the body has the same
// effect, and this makes the window deterministic.
response := new(mDNS.Msg)
response.SetReply(testQuery())
rawResponse, err := response.Pack()
if err != nil {
writer.WriteHeader(http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", transport.MimeType)
// Content-Length matters here: without it Exchange falls into io.ReadAll
// and waits for the stream FIN, which this handler is about to withhold.
writer.Header().Set("Content-Length", strconv.Itoa(len(rawResponse)))
writer.Write(rawResponse)
writer.(http.Flusher).Flush()
<-drainGate
body, _ := io.ReadAll(request.Body)
received <- body
})
listener, err := quic.ListenAddrEarly("127.0.0.1:0", testServerTLSConfig(t, []string{http3.NextProtoH3}), &quic.Config{
InitialStreamReceiveWindow: pinnedStreamWindow,
MaxStreamReceiveWindow: pinnedStreamWindow,
InitialConnectionReceiveWindow: 1 << 16,
MaxConnectionReceiveWindow: 1 << 16,
})
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: mux}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3-buffer", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: M.ParseSocksaddr(listener.Addr().String()),
tlsConfig: &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
},
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if _, err = dnsTransport.Exchange(ctx, message); err != nil {
t.Fatal("exchange: ", err)
}
// Exchange has returned, the body is still in flight. Drain the size class it
// came from, on this very goroutine, so a buffer released on the way out lands
// in our hands and not somewhere harmless. The buffers are held until after
// the comparison below.
poison := poisonPool(bufferSize, 32)
defer releaseAll(poison)
close(drainGate)
var sent []byte
select {
case sent = <-received:
case <-time.After(20 * time.Second):
t.Fatal("the server never received the request body")
}
if !bytes.Equal(sent, expected) {
firstDiff := -1
for i := 0; i < len(sent) && i < len(expected); i++ {
if sent[i] != expected[i] {
firstDiff = i
break
}
}
t.Fatalf("the query on the wire is not the query we packed: %d of %d bytes received, first difference at offset %d — "+
"the pooled request buffer was reused while quic-go was still reading it", len(sent), len(expected), firstDiff)
}
}
// markedQuery builds a query whose EDNS0 padding carries a random tag, so the
// packed bytes contain a needle that can be searched for in pool memory and
// cannot collide with anything else.
func markedQuery(t *testing.T) (*mDNS.Msg, []byte) {
t.Helper()
padding := make([]byte, markedQueryPadding)
if _, err := rand.Read(padding); err != nil {
t.Fatal(err)
}
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
opt := new(mDNS.OPT)
opt.Hdr.Name = "."
opt.Hdr.Rrtype = mDNS.TypeOPT
opt.Option = append(opt.Option, &mDNS.EDNS0_PADDING{Padding: padding})
message.Extra = append(message.Extra, opt)
return message, padding[:markedQueryNeedle]
}
// poolHoldsNeedle drains one size class of the buffer pool and reports whether
// any buffer in it still carries the needle. It must run on the goroutine that
// released the buffer: sync.Pool keeps a per-P private slot that no other P can
// steal from, and on the paths this test covers the release happens inline in
// RoundTripOpt, on the caller's own goroutine.
func poolHoldsNeedle(size int, needle []byte, count int) bool {
held := make([]*buf.Buffer, 0, count)
defer func() { releaseAll(held) }()
var found bool
for range count {
buffer := buf.NewSize(size)
held = append(held, buffer)
if bytes.Contains(buffer.FreeBytes(), needle) {
found = true
}
}
return found
}
// requireInstrumentFindsPackedQuery is the CONTROL. It does exactly what the old
// Exchange did — pack a query into a pooled buffer and release it — and demands
// that the scan below FINDS the needle. Without it, "the pool does not hold the
// query" would also be the verdict for a scan that can never find anything.
//
// Retried for the same reason its sibling control above is, and it was NOT
// before: under -race sync.Pool.Put drops one object in four, so a single
// attempt made this control — and with it the whole -race pass of the gate —
// fail on 18 of 60 measured runs with nothing wrong in the tree. A fresh
// needle is packed on each attempt, so a later one cannot be answered by an
// earlier one's bytes. See controlAttempts for what the retry does and does not
// buy.
func requireInstrumentFindsPackedQuery(t *testing.T) {
t.Helper()
for range controlAttempts {
message, needle := markedQuery(t)
size := 1 + message.Len()
exMessage := *message
exMessage.Id = 0
exMessage.Compress = true
buffer := buf.NewSize(size)
if _, err := exMessage.PackBuffer(buffer.FreeBytes()); err != nil {
t.Fatal(err)
}
buffer.Release()
if poolHoldsNeedle(size, needle, poolScanDepth) {
return
}
}
t.Fatalf("control failed: %d times in a row, a query packed into a pooled buffer and released was NOT "+
"found by the scan, so a clean verdict below would prove nothing. Under -race sync.Pool.Put drops "+
"one object in four, which is what the retries absorb; this many consecutive misses is something "+
"else — a pool that zeroes on Put, a size class that stopped being pooled, or buf.Buffer no longer "+
"handing its array back at all", controlAttempts)
}
// TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory pins the ownership rule the
// failure paths depend on.
//
// quic-go's http3.Transport closes the request body on every error path
// (transport.go RoundTripOpt) the moment doRequest returns, and doRequest waits
// only on the request-cancellation watchdog — never on the goroutine writing the
// body. So releasing the buffer when the body is closed is not an ownership
// handoff, and the only safe arrangement is for the query never to live in pool
// memory at all.
//
// This test encodes THAT design. A future guarded-pool design (a lock around
// Read and Close, refusing reads after release) would also be correct and would
// fail this test on purpose — it would have to replace it, and say so.
func TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory(t *testing.T) {
requireInstrumentFindsPackedQuery(t)
// A UDP socket nobody answers on: the handshake runs to the context deadline
// instead of being refused, which is the shape a router sees when the tunnel
// carrying its resolver drops.
blackhole, err := net.ListenUDP("udp", &net.UDPAddr{IP: net.IPv4(127, 0, 0, 1)})
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { blackhole.Close() })
for _, testCase := range []struct {
name string
ctx func(t *testing.T) (context.Context, context.CancelFunc)
}{
{
// RoundTripOpt closes the body after the handshake gives up.
name: "server never answers",
ctx: func(t *testing.T) (context.Context, context.CancelFunc) {
return context.WithTimeout(context.Background(), 500*time.Millisecond)
},
},
{
// The cancellation watchdog fires, then RoundTripOpt closes the body.
name: "context already cancelled",
ctx: func(t *testing.T) (context.Context, context.CancelFunc) {
ctx, cancel := context.WithCancel(context.Background())
cancel()
return ctx, func() {}
},
},
} {
t.Run(testCase.name, func(t *testing.T) {
message, needle := markedQuery(t)
size := 1 + message.Len()
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3-ownership", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: M.ParseSocksaddr(blackhole.LocalAddr().String()),
tlsConfig: &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
},
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := testCase.ctx(t)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, message); err == nil {
t.Fatal("expected the exchange to fail; this test is about the failure path")
}
// Same goroutine that ran RoundTripOpt, so the per-P private slot a
// release would have landed in is the one being drained.
if poolHoldsNeedle(size, needle, poolScanDepth) {
t.Fatal("the bytes of the query came back out of the buffer pool: the request body was packed into pooled " +
"memory and released while quic-go's body writer could still be reading it")
}
})
}
}
+351
View File
@@ -0,0 +1,351 @@
package quic
import (
"context"
"crypto/tls"
"net"
"net/http"
"net/url"
"sync"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
sbTLS "github.com/sagernet/sing-box/common/tls"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing-box/dns/transport"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
mDNS "github.com/miekg/dns"
)
var _ N.Dialer = (*trackingDialer)(nil)
// These tests pin down who owns the UDP socket handed to quic-go.
//
// quic-go's Dial/DialEarly take a net.PacketConn but do NOT take ownership of
// it: quic.setupTransport() builds a Transport with createdConn=false, and
// Transport.Close() then only calls conn.SetReadDeadline(time.Now()) instead of
// conn.Close(). So every QUIC connection torn down here — idle timeout, a
// retryable error, an engine reload calling Reset() — used to strand the UDP
// socket that carried it for the rest of the process's life. On a router that
// resolves through DoQ/DoH3 for months that is an unbounded fd leak.
//
// Both tests reconnect once and assert the socket from the FIRST connection is
// actually closed. Without the `<-conn.Context().Done() -> rawConn.Close()`
// watchdogs in quic.go / http3.go they fail on that assertion.
type trackedConn struct {
net.Conn
closeOnce sync.Once
closed chan struct{}
}
func (c *trackedConn) Close() error {
c.closeOnce.Do(func() { close(c.closed) })
return c.Conn.Close()
}
// trackingDialer hands out real UDP sockets and remembers every one of them.
type trackingDialer struct {
access sync.Mutex
conns []*trackedConn
}
func (d *trackingDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := (&net.Dialer{}).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn, closed: make(chan struct{})}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *trackingDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
func (d *trackingDialer) count() int {
d.access.Lock()
defer d.access.Unlock()
return len(d.conns)
}
func (d *trackingDialer) at(index int) *trackedConn {
d.access.Lock()
defer d.access.Unlock()
return d.conns[index]
}
func (d *trackingDialer) closeAll() {
d.access.Lock()
defer d.access.Unlock()
for _, conn := range d.conns {
conn.Close()
}
}
func requireClosed(t *testing.T, conn *trackedConn, what string) {
t.Helper()
select {
case <-conn.closed:
case <-time.After(5 * time.Second):
t.Fatalf("%s: the UDP socket of the retired QUIC connection was never closed — quic-go does not own it, we must", what)
}
}
func requireDialed(t *testing.T, dialer *trackingDialer, want int) {
t.Helper()
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
if dialer.count() >= want {
return
}
time.Sleep(10 * time.Millisecond)
}
t.Fatalf("expected at least %d dial(s), got %d", want, dialer.count())
}
func testServerTLSConfig(t *testing.T, nextProtos []string) *tls.Config {
t.Helper()
certificate, err := sbTLS.GenerateKeyPair(nil, nil, nil, "localhost")
if err != nil {
t.Fatal(err)
}
return &tls.Config{
Certificates: []tls.Certificate{*certificate},
NextProtos: nextProtos,
MinVersion: tls.VersionTLS13,
}
}
func testClientTLSConfig(t *testing.T, nextProtos []string) sbTLS.Config {
t.Helper()
config, err := sbTLS.NewClient(context.Background(), logger.NOP(), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
})
if err != nil {
t.Fatal(err)
}
config.SetNextProtos(nextProtos)
return config
}
// startDoQServer serves a minimal DoQ responder and returns its address.
func startDoQServer(t *testing.T) M.Socksaddr {
t.Helper()
listener, err := quic.ListenAddr("127.0.0.1:0", testServerTLSConfig(t, []string{"doq"}), nil)
if err != nil {
t.Fatal(err)
}
ctx, cancel := context.WithCancel(context.Background())
t.Cleanup(func() {
cancel()
listener.Close()
})
go func() {
for {
conn, acceptErr := listener.Accept(ctx)
if acceptErr != nil {
return
}
go func(conn *quic.Conn) {
for {
stream, streamErr := conn.AcceptStream(ctx)
if streamErr != nil {
return
}
go func(stream *quic.Stream) {
defer stream.Close()
request, readErr := transport.ReadMessage(stream)
if readErr != nil {
return
}
response := new(mDNS.Msg)
response.SetReply(request)
transport.WriteMessage(stream, 0, response)
}(stream)
}
}(conn)
}
}()
return M.ParseSocksaddr(listener.Addr().String())
}
func testQuery() *mDNS.Msg {
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
return message
}
func TestQUICTransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
serverAddr := startDoQServer(t)
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeQUIC, "test-doq", nil),
dialer: dialer,
serverAddr: serverAddr,
tlsConfig: testClientTLSConfig(t, []string{"doq"}),
connection: transport.NewConnPool(transport.ConnPoolOptions[*quic.Conn]{
Mode: transport.ConnPoolSingle,
IsAlive: func(conn *quic.Conn) bool {
return conn != nil && !common.Done(conn.Context())
},
Close: func(conn *quic.Conn, _ error) {
conn.CloseWithError(0, "")
},
}),
}
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
// Retire the connection the way a retryable error or an engine reload does.
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
// The reconnect must still work, on a fresh socket.
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err := dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func TestHTTP3TransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
mux := http.NewServeMux()
mux.HandleFunc("/dns-query", func(writer http.ResponseWriter, request *http.Request) {
message, err := readRequestMessage(request)
if err != nil {
writer.WriteHeader(http.StatusBadRequest)
return
}
response := new(mDNS.Msg)
response.SetReply(message)
rawResponse, err := response.Pack()
if err != nil {
writer.WriteHeader(http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", transport.MimeType)
writer.Write(rawResponse)
})
listener, err := quic.ListenAddrEarly("127.0.0.1:0", testServerTLSConfig(t, []string{http3.NextProtoH3}), nil)
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: mux}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
serverAddr := M.ParseSocksaddr(listener.Addr().String())
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
stdConfig := &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
}
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: serverAddr,
tlsConfig: stdConfig,
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err = dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func readRequestMessage(request *http.Request) (*mDNS.Msg, error) {
defer request.Body.Close()
rawMessage := make([]byte, 4096)
n, err := readFull(request.Body, rawMessage)
if err != nil {
return nil, err
}
var message mDNS.Msg
err = message.Unpack(rawMessage[:n])
if err != nil {
return nil, err
}
return &message, nil
}
func readFull(reader interface{ Read([]byte) (int, error) }, buffer []byte) (int, error) {
var total int
for total < len(buffer) {
n, err := reader.Read(buffer[total:])
total += n
if err != nil {
if total > 0 {
return total, nil
}
return total, err
}
}
return total, nil
}
+12
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"os"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/sing-box/adapter"
@@ -117,6 +118,12 @@ func (t *Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS.Msg,
rawConn.Close()
return nil, E.Cause(err, "establish QUIC connection")
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-earlyConnection.Context().Done()
rawConn.Close()
}()
return earlyConnection, nil
})
if err != nil {
@@ -144,6 +151,11 @@ func (t *Transport) exchange(ctx context.Context, message *mDNS.Msg, conn *quic.
return nil, E.Cause(err, "open stream")
}
defer stream.CancelRead(0)
stopWatch := context.AfterFunc(ctx, func() {
stream.CancelRead(0)
_ = stream.SetWriteDeadline(time.Now())
})
defer stopWatch()
err = transport.WriteMessage(stream, 0, message)
if err != nil {
stream.Close()
+44
View File
@@ -12,6 +12,50 @@ as GitHub **pre-releases** and never become "Latest".
#### Unreleased (shater)
**`l3-honest-drop` — ICMP routed to an L4-only outbound is dropped, not
forged** — ships with `shaterd` (part of the shater L3 ingress,
`docs-shater/DECISIONS.md` D25), not as an lx release tag; recorded here because
it edits two upstream files. Without it the TUN stack answers an unroutable echo
ITSELF — sing-tun's `ICMPForwarder.HandlePacket` rewrites Echo→EchoReply
whenever the flow judgment comes back Accept (`stack_gvisor_icmp.go`) — so a
ping routed to vless/vmess/… would read as a working tunnel while the packet
never left the router.
* **`route/route.go` (`PreMatch`)** — the pre-match walk was renamed to
`preMatch` and the exported `PreMatch` became a thin FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for
`N.NetworkICMP`. An earlier version overrode `continueResult` inside
`preMatchFlow` instead; that covered only the exits reaching that function and
left three of the walk's own exits forging — the `prepareMatchMetadata` error
return, the sniff bail-outs, and the `default:` arm of the rule-action switch
(every action pre-match has no arm for: `hijack-dns`, `direct`, …). A guard on
the single return value cannot be outgrown by a new exit. `PreMatchBypass` is
folded in because sing-tun implements `ActionBypass` on the nfqueue plane only
— on the TUN path it lands in the same `default:` arm as Accept, i.e. forges.
* **`adapter/router.go` (`JudgeFlow`, the `!isPort` branch)** — ICMP returns
`ActionDrop` where it fell through to `ActionAccept`. Second line of defense:
`adapter.FlowOutbound` and `tun.Port` are distinct interfaces, and a drift
between them must not quietly re-enable the forged reply.
* **TCP/UDP behaviour is unchanged** — `PreMatchContinue` still means "take the
ordinary connection route" for both, `PreMatchBypass` still means bypass, and
the `!isPort` fallthrough still returns `ActionAccept` for them; pinned by
`route/prematch_icmp_lx_test.go` and `adapter/judgeflow_icmp_lx_test.go`
(both inside the marker), each ICMP case having an explicit TCP/UDP twin.
* **NOT covered: a FRAGMENTED echo to a WireGuard/AWG outbound is still
forged** — sing-tun's `ForwardDispatcher.Dispatch` returns before asking for a
verdict at all when `parsed.fragment`, and the reassembled packet reaches
`ICMPForwarder.HandlePacket`, whose `installFlow` demands an UNSPECIFIED port
address that a WireGuard endpoint never has. Fixing it inside `JudgeFlow`
is NOT possible — both consumers call it with identical arguments and the
working path needs the concrete address. Full chain, the two viable fixes and
the trap are in `docs-shater/DECISIONS.md` D25, under "What is still NOT
covered, said plainly", item 2.
* **Rebase cost: two small marked blocks** (`lx:begin/end l3-honest-drop`, a
wrapper function in `route/route.go` and one branch body in
`adapter/router.go`) plus the two self-contained test files — carried across
an upstream rebase by eye. Note that `PreMatch`'s own body now lives in
`preMatch`, so an upstream change to the walk applies to that function.
**Fork-layer + control-plane rework of proxy health** — ships with `shaterd`
(the shater router daemon), not as an lx release tag; recorded here because the
load-bearing half lives in fork zones (`common/urltest`, `protocol/group`).
+63 -8
View File
@@ -62,16 +62,61 @@ flowchart LR
C["LAN client"] -->|"nft tproxy, mark → tproxy port"| IN["sing-box tproxy inbound (sniff SNI/Host/QUIC)"]
IN --> R{"route: rule match — src / dst / list / geo / client"}
R -->|"proxied"| OUT["outbound / selector (balancer, chain)"]
R -->|"direct"| DIR["direct (flow-offload on)"]
R -->|"direct"| DIR["direct (out the normal route, untunnelled)"]
R -->|"blocked"| BLK["block"]
OUT --> NET["exit — VLESS/Reality/AmneziaWG2/Hysteria2/…"]
```
Reliability (ported from v0.1): own nft table `inet shater` + own marks/tables
(never touch fw4); atomic validate→stage→swap; commit-confirm rollback;
idempotent reconcile under flock; management-bypass always; fail-closed
(never touch fw4); atomic validate→stage→swap; commit-confirm rollback (opt-in —
see §5); idempotent reconcile under flock; management-bypass always; fail-closed
kill-switch (dead group → block, not a silent direct leak).
Only TCP and UDP reach that path — TPROXY carries nothing else. What happens to
the rest is §3a.
### 3a. L3 ingress and kernel egress — what TPROXY cannot carry
Two opt-in globals cover the protocols the tproxy plane leaves on the floor.
Both are off in a stock config, and both are configured through UCI only (the
panel does not expose them).
**`globals.l3_tunnel` — LAN ICMP through the tunnel.** The generator adds a
synthetic `tun` inbound tagged `l3-in` (gVisor stack, `auto_route` **off**, MTU
65535, `shater/generate/inbound.go`), so ICMP is routed by the engine's own rules
instead of being dropped or answered by a forged local reply. The device is not
one fixed name: the generator emits a stable placeholder (so a no-op reconcile
still hashes identical and does not rebuild the engine once a minute), and
`shater/engine/l3slot.go` substitutes one of the two slots `shater-l3a` /
`shater-l3b` (`netplane/l3.go`) just before `box.New` — a new generation must
never reopen the name the outgoing one still holds
(`TUNSETIFF: device or resource busy` took the whole LAN down once). The routing half is scoped and lives entirely outside
the main table: our nft prerouting chain stamps LAN `icmp`/`ipv6-icmp` with
`L3Mark` (`fwmark_base + 0x80`), and `netplane.addL3Routing` binds that mark to
`L3Table` (`table_base + 8`), whose only content is a default route out the live
slot. Because the daemon creates the device at runtime, netifd never learns about
it and fw4 would reject the forward on its own account — so `30_shater-core`
seeds a **`shater_l3` zone in the user's `/etc/config/firewall`**, matching
`list device 'shater-l3*'` (a string match that is valid before the TUN exists
and covers both slots). Ceiling: ICMP echo only, and only for L3-capable
egresses; see `DECISIONS.md` D25 for what is still not covered.
**`globals.untunnelable_egress` — everything else, carried by the kernel.** It
names an existing interface/tunnel egress. Whatever the L3 block above did not
claim — ESP/AH, GRE, IGMP, SCTP, and ICMP too when `l3_tunnel` is off — is
stamped in prerouting with **that egress's own mark** (`netplane/nft.go`,
`UntunnelableEgressBinding`) and accepted; the `fwmark → table` pair
`addEgressRouting` already installed for the egress then routes it out the
egress's device. No new mark, no new table, and the engine never sees a byte —
which is why any IP protocol works here while the L3 TUN is narrow. Order is
load-bearing: this sweep runs **after** the L3 marking (first match wins) and
**after** the local-plane accepts, so LAN-to-LAN, router-addressed traffic and
IPv6 neighbour discovery never leave through an uplink. With `ipv6=0` the mark
is scoped to `nfproto ipv4`, because `addEgressRouting` installs the `-6`
rule/table pair only when IPv6 is on and marked v6 without it would fall through
to the main table past the kill-switch. `globals.untunnelable` (block | icmp |
direct) stays in charge of whatever neither mechanism carries.
## 4. DNS + filtering + stats
```mermaid
@@ -79,7 +124,7 @@ flowchart LR
C["client :53"] -->|"hijack"| DNS["sing-box DNS (in-process)"]
DNS --> FILT{"shater filter: blocklists + allowlist + per-device policy"}
FILT -->|"blocked"| NX["NXDOMAIN / 0.0.0.0"]
FILT -->|"allowed"| RES["resolvers (DoH/DoT/plain/FakeIP) + nftset for routing"]
FILT -->|"allowed"| RES["resolvers (DoH/DoT/plain/local/FakeIP), per-rule detour"]
DNS -->|"query events (engine observability)"| AGG["shater stats aggregator"]
AGG --> PANEL["panel: top domains · per-device · allowed/blocked · timeline"]
```
@@ -87,9 +132,10 @@ flowchart LR
Because the engine's DNS runs **in our process**, every query (domain, client,
verdict, latency) is available to the stats aggregator without log-scraping —
this is the payoff of embedding. Blocklist matching uses an efficient compiled
matcher, not dnsmasq megalists (see `DECISIONS.md` D5). Per-device blocking =
engine route/DNS rule keyed by client, or nftset(device) × nftset(blocked-domain)
→ drop.
matcher, not dnsmasq megalists (see `DECISIONS.md` D5). Per-device blocking is an
engine route/DNS rule keyed by client. Routing decisions come from in-engine
rule-sets: the v0.1 mechanism where dnsmasq populated nft sets does not exist in
v0.2 (`generate/dns.go`).
## 5. Config & apply flow
@@ -100,11 +146,20 @@ stateDiagram-v2
Render --> Validate: engine config check + nft -c
Validate --> KeepOld: fail
Validate --> Apply: ok (atomic swap: engine reload + nft/route reconcile)
Apply --> ConfirmWindow
Apply --> Committed: confirm_timeout = 0 (SHIPPED DEFAULT — nothing armed)
Apply --> ConfirmWindow: confirm_timeout > 0
ConfirmWindow --> Committed: confirmed
ConfirmWindow --> Rollback: timeout
Rollback --> LastGood
```
**The confirm window is opt-in and ships closed.** `model.DefaultGlobals()` leaves
`ConfirmTimeout` at zero, the shipped `/etc/config/shater` says
`option confirm_timeout '0'`, and `apply.ArmRollback` returns immediately on a
non-positive timeout — so on a stock install every apply takes the left edge above
and there is no net under it. `shaterd apply` reports which edge it took
(`reason: commit-confirm-off` vs an armed window). Set
`globals.confirm_timeout` to arm it.
## 6. Roadmap tiers
See `ROADMAP.md` for the phased plan and `FEATURES.md` for the full feature list.
+29 -16
View File
@@ -51,6 +51,9 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
MASQUE/WARP, and gRPC observability (DNS queries / rules / outbounds). Upstream
sing-box brings VLESS/VMess/Trojan/Shadowsocks/WireGuard/Reality + Hysteria2/
TUIC. It is library-first (`libbox`) and **GPL-3.0** (compatible with us).
That list is what the FORK can build, not what shater ships: `shater/registry`
registers only what `shater/generate` can emit, and MASQUE is one of the types
deliberately left out (~6 MB of binary and resident RAM). See `FEATURES.md`.
- We **fork it** (not just depend on it) so we can embed literally everything —
control-plane, admin panel, DNS filter — and integrate tightly with the
engine internals (DNS, routing, stats). This is a deliberate, decided
@@ -84,13 +87,13 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
## Repository model
- **`shater` `main` = our fork of sing-box-lx.** After Phase 1 it contains the
full sing-box-lx tree PLUS our additive overlay (`shater/`, `panel/`,
`openwrt/`, `docs-shater/`). Upstream is tracked via a git remote and merged by tag.
- **`shater` `main` = our fork of sing-box-lx.** It contains the full sing-box-lx
tree PLUS our additive overlay (`shater/`, `panel/`, `openwrt/`, `docs-shater/`,
`scripts/`, `ci/`). Upstream is tracked via a git remote and merged by tag.
Phase 1 merged the engine in on 2026-07-14 (`v1.14.0-lx.3`); `main` has not been
a docs-only seed since.
- **`shater` branch `v0.1`** = the standalone xray-based version (frozen, ported
from).
- Until Phase 1 merges the engine in, `main` is the docs-first overlay seed you
are reading now (LICENSE, README, `docs-shater/`, the feed signing key).
## What to port from v0.1 (don't rewrite these ideas)
@@ -119,20 +122,30 @@ filter/stats engine wired into sing-box's DNS.
v0.2 fork; branch `v0.1` = the working xray-based version.
- **Upstream to track:** `https://github.com/Leadaxe/sing-box-lx` (which tracks
`https://github.com/SagerNet/sing-box`).
- **CI:** Gitea Actions (act_runner + Docker). v0.1's workflow was removed from
`main`; new CI is added when the v0.2 build exists.
- **CI:** Gitea Actions (act_runner + Docker), `.gitea/workflows/release.yml` —
builds the four packages through the ImmortalWrt 25.12.1 SDK and publishes the
signed per-arch apk repo. The opkg/`.ipk` lane was deleted, not disabled (D22).
- **Test gate:** `bash scripts/run-tests.sh` — the whole suite under the SHIPPED
build tags, on linux (in Docker from a non-linux host), with `-race`, and with
three anti-silent-skip checks. Not optional reading before touching `shater/`.
- **Feed signing:** EC (prime256v1) key for the apk index; secret in the repo
secret `KEY_APK`; public key `dist/shater-apk.pem`, installed on routers as
`/etc/apk/keys/shater-apk.pem`. Never regenerate it (D22).
- **Test VM:** OpenWrt 24.10.3 x86_64 in Docker (`docker ps --filter
name=openwrt-vm`). SSH via the ssh-manager MCP server `local_openwrt`
(localhost:2222, root/openwrt). LuCI at `http://127.0.0.1:8080` (root/openwrt),
drivable with the Playwright MCP.
- **Test VM:** **ImmortalWrt 25.12.1** (`r37978-cd0a06bfd3fd`) x86_64 in Docker
(`docker ps --filter name=openwrt-vm`), apk-tools 3.0.5 — deliberately the same
revision as `mini_router`, and required: the only package format we publish is
`.apk`, which does not install on 24.10 at all. SSH via the ssh-manager MCP
server `local_openwrt` (localhost:2222, root/openwrt). LuCI at
`http://127.0.0.1:8080` (root/openwrt), drivable with the Playwright MCP.
- **Routers:** `mini_router` (BPi-R3 Mini, ImmortalWrt 25.12.1) carries the real
home traffic; `main_router` (BPi-R4, OpenWrt 25.12.0). Both `aarch64_cortex-a53`,
both apk-tools 3.0.5 — see the table in D22.
## Current status
Repo reset done: v0.1 preserved on its branch; `main` cleaned to this docs-first
scaffold. Next is Phase 1 in `ROADMAP.md` — fork sing-box-lx into `main`
(add upstream remote, merge a pinned tag), stand up the embedding prototype
(prove AmneziaWG 2.0, measure binary size with feature-trim + `-s -w` + UPX)
before building the control plane and panel.
**v0.2 is feature-complete and running on real hardware.** ROADMAP Phases 0–8 are
done and VM-verified; the product ships as a signed apk feed and is installed on
`mini_router`. Read `ROADMAP.md` for what each phase delivered, `FEATURES.md` for
the honest MVP/T1/T2 state of each feature (including what is declared but not
shipped), and `DECISIONS.md` for why. Work since Phase 8 has been correctness and
honesty passes rather than new phases.
+718
View File
@@ -142,6 +142,12 @@ with a modest one-time decompress-into-RAM cost at start. Ship compressed; keep
uncompressed artifact for debugging.
## D13 — External DPI-bypass tool = **ByeDPI** (a SOCKS egress), NOT zapret
> **PARTLY REVERSED by [D29](#d29--byedpi-is-removed-the-presets-it-replaced-were-not-weak-they-were-broken) (2026-07-27).** The
> comparison below still stands and zapret is still rejected. What did not stand
> is the premise that the native presets were too weak to carry this: they were
> not weak, they were defective. ByeDPI, the `byedpi` egress kind and the
> `openwrt/byedpi` package are gone. Read D29 before acting on anything here.
Decided 2026-07-14. We evaluated exactly two external desync tools — **zapret**
(nfqws/tpws, NFQUEUE packet plane) vs **ByeDPI/ciadpi** (a local SOCKS5 desync
proxy) — and picked **one**: ByeDPI. DPI-bypass stays a **per-ruleset egress
@@ -319,6 +325,10 @@ Three values, not two, because the leaks differ in *kind*: an ICMP echo is ephem
user-initiated and reveals the address only to a host the user deliberately contacted,
whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A
single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".
*(Refined 2026-07-26 by D25: still true of TPROXY — but ICMP echo now has an
opt-in data plane of its own, the dedicated L3 TUN, so the policy no longer
speaks alone for ping; it keeps sole charge of ESP/GRE/IGMP and of the degraded
paths.)*
**Fail-open degradations must be visible in the panel, not only in `logread`.** The
audit deliberately converted many aborts into warn-and-continue (an unfetchable list,
@@ -772,3 +782,711 @@ server, which restores exactly the pre-D24 behaviour and clears the notice. Unti
that lands, an operator can get the same result by setting `endpoint_resolver` to a
direct resolver. Note the hazard is **not** created by D24 — any config with two
resolvers has it today; the default merely makes it universal.
## D25 — L3 ingress: LAN ICMP rides a dedicated TUN through the tunnel, not a policy verdict
Decided 2026-07-26. D17 made everything TPROXY cannot divert an explicit policy
(`Globals.Untunnelable` = block | icmp | direct) — and its premise still holds:
kernel TPROXY delivers a packet by handing it to a listening SOCKET, and sockets
exist for TCP and UDP only, so an ICMP echo has nothing to be handed to. But a
policy can only choose between losing the packet and leaking it with the
client's real source address; neither ever puts a ping THROUGH the tunnel. This
decision adds the data plane D17 could not have: **`globals.l3_tunnel` (opt-in,
default off; `model.Globals.L3Tunnel`) opens a second, dedicated ingress — a TUN
device — and LAN ICMP enters the engine as raw IP packets**, where the ordinary
route rules pick an outbound exactly as for any flow. The policy is refined, not
repealed: it keeps sole charge of the protocols the engine cannot ingest at all,
and of the degraded paths (both below).
**The whole mechanism is one mark, one rule, one device — the TPROXY plane is
untouched.** The nft prerouting chain stamps `L3Mark` (= fwmark_base + 0x80,
`netplane/nft.go` `l3MarkOffset`) on LAN `ip protocol icmp` / `meta l4proto
ipv6-icmp` ONLY, and only after every local plane was already accepted
(fib-local, RFC1918/link-local/multicast daddr sets) and — for v6 — after a
unicast ND/NA carve-out, because one tunnelled neighbour probe is enough to take
the LAN's v6 plane down (`renderNft`, the L3 block). `addL3Routing`
(`netplane/apply.go`) binds that mark to a table (= table_base + 0x08) whose
only content is `default dev shater-l3`; del-then-add idempotent, and a failed
rule or route is a NAMED operator warning, never an apply abort. `generate`
emits the synthetic `l3-in` TUN inbound bound to exactly `netplane.L3Device`,
MTU 65535 (the largest IP datagram there can be, so the KERNEL can never
fragment on the way in — see "the device MTU is not a tunnel budget" below),
point-to-point /30 + /126 addresses from private space,
the v6 one only when `globals.ipv6` is on — and only next to a tproxy inbound:
the ingress rides the same LAN divert plane, and without one the TUN would sit
dark while the config claims ICMP is tunnelled, so it is skipped with a warning
(`generate/inbound.go`, `appendL3TunInbound`). `shater/registry` registers the
`tun` inbound type; that costs no new build tag and no meaningful size because
`with_wireguard` already requires `with_gvisor` (D23, `scripts/router-tags.sh`).
**`auto_route: false` is load-bearing, not a default we happened to keep.**
sing-box's auto_route rewrites the router's MAIN routing table — it would drag
everything the router itself sends (WAN traffic, DNS, the tunnel's own underlay)
into this TUN. The fwmark rule + dedicated table above is deliberately the ONLY
entrance, and disabling the feature can never strand a stale default route in
main (`generate/inbound.go`; pinned by `TestL3TunnelEmitsTunInbound`).
**`stack: "gvisor"` is a deliberate choice, and the tempting reason for it is
wrong.** It is TRUE that sing-tun's system stack answers an ICMP echo LOCALLY —
`processIPv4ICMP` rewrites Echo→EchoReply in place and swaps the addresses
(sing-tun `stack_system.go:648`; the v6 twin sits right under it). It is FALSE
that this makes the system stack unusable here: `dispatchIPv4`
(`stack_system.go:355-372`) hands the packet to the SAME `ForwardDispatcher`
first and only falls through to that forger for packets addressed to the TUN
itself, exactly as the gVisor filter does (`stack_gvisor_filter.go:52-113`).
Both stacks would forward. gvisor is chosen because it is already linked —
`with_wireguard` requires `with_gvisor` (D23), so it costs no tag and no new
code path — and because it is the combination the integration test actually
exercises. Do not re-derive this as "the system stack fakes ping": it fakes ping
only where the dispatcher declined the packet.
**The ceiling is ICMP echo, and it is upstream's dispatcher — NOT the netstack.**
This distinction matters because the netstack answer is the intuitive one and it
is wrong. On the forward path a WireGuard/AWG endpoint never consults gVisor at
all: `Endpoint.WritePackets` (`transport/wireguard/port.go:21-58`) reads the IP
version and the destination address and hands the raw bytes to
`wgDevice.InputPackets` — the protocol byte is never examined — and
`returnDeviceWrapper.Write` (`:127-157`) offers every decrypted packet to
`returnPath.ReturnPackets` before the stack sees it. WireGuard would carry ESP
today if anything handed it one. What refuses is `ForwardDispatcher`: its parser
sets `hasFlow` for TCP, UDP and ICMP echo alone (`flow_parse.go`,
`parseTransport`, the echo identifier serving as the pseudo-port), and
`createFlow` NATs through a port-shaped selector (`flow_dispatch.go:325`,
`allocateSelector`) that ESP, AH and GRE do not have. So ESP/AH/GRE/IGMP/SCTP
cannot enter the engine in ANY configuration and REMAIN on the D17 policy —
or on the kernel egress of D26, which sidesteps the dispatcher entirely. The nft
plane encodes the same boundary on purpose: it marks `icmp`/`ipv6-icmp` only,
never `l4proto != { tcp, udp }`, because a marked ESP packet would enter the
device and vanish — a black hole wearing a tunnel's name — instead of receiving
the policy's honest verdict (`netplane/nft.go`, the prerouting L3 comment).
**What works and what does not, read off the upstream source.** ping v4/v6 —
yes. Windows `tracert` — yes: the gVisor return path recognises
`ICMPv4TimeExceeded` and `ICMPv4DstUnreachable` alongside EchoReply and NATs
them back to the LAN client (`stack_gvisor_icmp.go:341+`, `returnPacket`). IPv6
traceroute — intermediate hops stay invisible: the v6 branch of the same
function accepts EchoReply only, so just the final destination answers. Several
LAN clients behind the one tunnel address are already solved upstream:
`ForwardDispatcher` NATs by echo identifier and rewrites the source to the
outbound's port address (`flow_dispatch.go:325+`, `createFlow`; `icmpFlowKey`) —
we wrote no NAT of our own.
**Which outbounds can carry it.** The contract is `adapter.FlowOutbound`
(= `Outbound` + `tun.Port` + `PreMatchFlow`, `adapter/outbound.go`). In-tree
implementors: the WireGuard/AWG endpoint (`protocol/wireguard`), `direct`
(`protocol/direct`), `bridge` (`protocol/bridge`), `tailscale`
(`protocol/tailscale`). Of those, the shaterd registry can construct only
WireGuard/AWG and direct (`shater/registry/registry.go` — bridge and tailscale
are not registered). Every proxy protocol — vless/vmess/trojan/shadowsocks/
hysteria2/tuic/socks/http/shadowtls — is L4-only and cannot. Recorded as a known
gap: `masque` is L3 by nature (CONNECT-IP; it builds a userspace gVisor stack
per tunnel, `protocol/masque/outbound.go`) but implements no `tun.Port` and is
not in the shater registry, so today it cannot carry the ingress. Wiring it up
is possible future work, not a promise.
**ICMP to an L4-only outbound is DROPPED, and that took patching upstream files
(the `lx:l3-honest-drop` delta — see `docs-lx/lx-changelog.md`).** In the gVisor
stack the fallthrough verdict is a forgery: `ICMPForwarder.HandlePacket` answers
the echo ITSELF (Echo→EchoReply + address swap) whenever the flow judgment comes
back Accept (`stack_gvisor_icmp.go:120`), and upstream maps "no flow route" to
exactly that Accept — so a ping routed to vless would read as tunnelled while
the packet died on the router. Two small marked hunks make the truth observable:
`route/route.go` wraps the whole pre-match walk — the walk itself became
`preMatch`, and the exported `PreMatch` is now a FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for `N.NetworkICMP` —
and `adapter/router.go` (`JudgeFlow`, the `!isPort` branch) returns `ActionDrop`
for ICMP where it fell through to `ActionAccept` — the second line of defense,
because `FlowOutbound` and `tun.Port` are distinct interfaces and a drift
between them must not quietly re-enable the forger. TCP/UDP verdicts are
byte-identical; `route/prematch_icmp_lx_test.go` and
`adapter/judgeflow_icmp_lx_test.go` pin both directions. The operator-facing
text says the same out loud (`shater/apply/warnings.go`): proxy-routed addresses
"cannot be pinged at all — deliberately".
> **Why a funnel and not an override inside the walk.** The first version of
> this delta overrode the pre-declared `continueResult` inside `preMatchFlow`
> and claimed to cover "every exit point of the function at once". It covered
> every exit of THAT function; the walk above it has exits of its own that never
> reach it — the `prepareMatchMetadata` error return (which arrived later, with
> the shared-metadata refactor, upstream `b911fb078`), the sniff bail-outs, and
> the `default:` arm of the rule-action switch, which catches every action
> pre-match has no arm for (`hijack-dns`, `direct`, and whatever upstream adds
> next). Each of those returned `PreMatchContinue`, i.e. `tun.ActionAccept`,
> i.e. the forged reply. A guard on the single return value cannot be outgrown
> by a new exit. `PreMatchBypass` joined the drop for the same reason: sing-tun
> implements `ActionBypass` on the nfqueue plane only — the name appears nowhere
> in `flow_dispatch.go` or `stack_gvisor_icmp.go` — so on the TUN path it lands
> in the same `default:` arm as Accept and forges too. There is no honest bypass
> for a packet that is already inside the engine's TUN.
**The device MTU is NOT a tunnel budget, and pretending it was manufactured
forged replies.** `l3-in` is created with MTU **65535**, not the tunnel's 1420,
and the maximum is the whole argument. This MTU governs exactly one thing:
whether the KERNEL splits a packet on its way INTO the device. What the engine
then puts into the tunnel is sized separately and correctly, against the
OUTBOUND's MTU — `ForwardDispatcher.forwardToPort` (`flow_dispatch.go:445-481`)
measures every forwarded packet against `Port.PortMTU()` and either fragments to
it (no DF, `fragmentIPv4Packet`) or answers a well-formed `fragmentation needed`
quoting it (DF, `buildFragmentationNeeded`, source = the far host, so PMTU
discovery works end to end). That machinery was always there; it was simply
never handed a whole packet.
At 1420 it wasn't. Anything above 1392 bytes of payload was fragmented by the
kernel at this device, and a fragment is the one thing sing-tun will not judge:
`Dispatch` (`flow_dispatch.go:176-177`) returns on `parsed.fragment` BEFORE
calling `JudgeFlow` at all. The fragments fell through to the gVisor stack —
promiscuous and spoofing (`stack_gvisor.go:219-223`) — which reassembled them
and handed the echo to `ICMPForwarder.HandlePacket` (`stack_gvisor_icmp.go:105+`),
whose `installFlow` (`:233-244`) writes to the port UNMODIFIED and therefore
demands a port address that is valid **and UNSPECIFIED**. `direct` qualifies
(`IPv4Unspecified()`); a WireGuard/AWG endpoint reports its concrete interface
address (`transport/wireguard/port.go:13`) and does not. So it declined, and
`HandlePacket` fell past the switch and FORGED the reply: `SetType(EchoReply)` +
address swap. Net effect on the operator's bench: `ping -s 1392` honest,
`ping -s 1393` a lie told by the router — and the lie was, of course, only for
the outbounds this feature exists for. (Upstream applies the very same
unspecified test and answers it honestly in the cloudflared ICMP handler,
`protocol/cloudflare/inbound.go:163-167`: it drops. Only the TUN path forges.)
65535 rather than "big enough": no IP datagram can exceed it, so the kernel
CANNOT fragment at this device, for any packet, ever. Any smaller value leaves
a band of sizes open and re-opens the class. It is also sing-box's own default
TUN MTU on Linux. Pinned by `TestL3TunnelMTULeavesNothingForTheKernelToFragment`
and `TestL3TunnelMTUIsNotATunnelBudget` (`generate/l3mtu_test.go`), and — the
assertion that matters — by the integration test reading the MTU back off the
real kernel device, since a kernel that clamped it would restore the forgery
without changing a generated byte.
Memory was MEASURED, not reasoned about: three paired runs of
`TestIntegrationL3TunInboundStarts` under `-test.memprofilerate=1` (exact
accounting, not sampled) allocate 5.41 / 5.48 / 5.47 MB at 65535 against
5.76 / 5.46 / 5.70 MB at 1420, and a `-diff_base` profile attributes every
difference to netlink interface enumeration. Nothing in the read path scales
with the MTU: gVisor reads through `fdbased.BufConfig`, which sing-tun's `init`
pins to a single 65535-byte view regardless of MTU, and `fdbased` keeps `mtu`
only to return it from `MTU()`. Two adjacent facts, recorded because both are
easy to derive wrongly: (a) `protocol/tun` computes
`enableGSO = stack == gvisor && mtu < 49152`, so this MTU turns GSO off there —
and then `StartStateStart` turns it back ON unconditionally because an
`adapter.FlowOutbound` exists in the config, so the ~1.98 MB of TCP/UDP GRO
scaffolding is present at BOTH MTUs and is priced by the flow-capable outbound,
not by this number; (b) the `mtu_fix` on the `shater_l3` fw4 zone is now inert —
only ICMP is ever marked into the device — and its uci-defaults comment still
says "the tunnel MTU is 1420".
**What is still NOT covered, said plainly.**
1. **A big ping does not start WORKING — it starts FAILING HONESTLY.** Upstream's
ICMP NAT is unfragmented-only in BOTH directions: `classifyReturn`
(`flow_dispatch.go:703-710`) returns `returnPass` on `parsed.fragment` exactly
as the forward path does. So a non-DF `ping -s 2000` now genuinely leaves the
router (fragmented to the tunnel MTU by `forwardToPort`), the far host really
answers, and the reply — fragmented by the peer to fit the tunnel — is not
NAT'd back to the LAN client. The operator sees a timeout. That is the
feature's promise ("travels or fails honestly"), not a capability claim.
Carrying oversized ICMP end to end would need reassembly upstream does not
have; it is not planned.
2. **A client that puts fragments on the wire ITSELF.** The device MTU cannot
un-fragment what already arrived fragmented, so such packets still reach the
gVisor stack, still get reassembled there, and still receive a forged reply
when the outbound is WireGuard/AWG. This is the residue the planned
`ip frag-off & 0x3fff != 0` prerouting carve-out (`netplane/nft.go`) is for.
**Whoever writes that rule must first check whether it can ever match:** fw4's
ruleset uses conntrack, conntrack pulls in `nf_defrag_ipv4`/`nf_defrag_ipv6`,
and defrag REASSEMBLES in PREROUTING before our marking rules run. Where
defrag is active the case does not arise (the MTU covers it) and the rule is
dead; where it is not, the rule is the only cover. Verify on the bench with
`nft list ruleset | grep -c ct` and a fragment counter, do not assume.
3. **The DF path changed hands and is untested on hardware.** It used to be the
kernel that answered `fragmentation needed` (from the router's LAN address,
MTU 1420); it is now the engine (from the far host's address, quoting
`Port.PortMTU()`). Both are correct PMTUD; only the first has ever run on a
real router.
**fw4 has to be told about the device, and `list device` is the only spelling
that works.** nftables runs EVERY table on every packet and a drop in any one of
them wins — an accept in `inet shater` cannot override fw4, and fw4 WILL reject
this forward: netifd never learns about a device the daemon creates at runtime,
so `shater-l3` belongs to no zone and falls into fw4's zone-less defaults. Hence
a real fw4 zone `shater_l3` + a lan→shater_l3 forwarding, seeded idempotently
(NAMED sections) and unconditionally in uci-defaults
(`openwrt/shater-core/files/etc/uci-defaults/30_shater-core`, `seed_l3_zone`),
with `mtu_fix` set. That `mtu_fix` is now inert and should be read as such: it
clamps forwarded TCP MSS to the route MTU, the device MTU is 65535, and nothing
but ICMP is ever marked into this device — the uci-defaults comment still says
"the tunnel MTU is 1420" and is stale. The device is attached
via `list device`, deliberately NOT `list network`: fw4 resolves a zone's
networks through netifd, which yields an EMPTY device set for a runtime-created
TUN (a proto-none stub would have to be brought UP to contribute an l3_device,
and nothing ever brings it up), while `list device` compiles to a plain
iifname/oifname string match — valid before the TUN exists, matching from the
moment shaterd creates it. `kmod-tun` joined DEPENDS so a slimmed image cannot
lose `/dev/net/tun` (`openwrt/shater-core/Makefile`). Our own forward chain
accepts both TUN legs ahead of the fail-closed drops — accepts that speak for
OUR table only (`netplane/nft.go`, forward chain step 4).
**What the policy still owns, and the one combination that now warns.** With the
ingress on, the mark is stamped in prerouting and the ROUTING decision carries
echo into the TUN before the forward chain — where the policy's verdicts live —
is ever consulted; that holds under every `untunnelable` value. The policy
therefore governs exactly two things: the never-markable protocols above, and
the fallback when the L3 rule/route did not come up (engine down, partial apply)
— `block` turns that failure into an honest loss, `direct` into a silent leak
with the real address. That is why `l3_tunnel` + `untunnelable=direct` draws a
validation warning naming the safe choice (`model/validate.go`), why every
rule/route failure surfaces as a named panel warning rather than an abort
(`addL3Routing`), and why the D17 HOLDING plane never marks: the TUN is created
BY the engine, and the holding plane exists precisely because the engine is not
running — marking would dead-end ping in a device that does not exist
(`netplane/nft.go`, hold comment).
- **Rejected: `auto_route` / letting the engine own the routing.** It rewrites
the main table and intercepts the router's own WAN/DNS/underlay traffic; the
blast radius of a toggle meant for LAN ping would be the whole router.
- **Rejected: marking all `l4proto != { tcp, udp }` into the TUN.** ESP/AH/GRE/
IGMP/SCTP cannot be parsed into flows upstream; they would vanish inside the
device. A drop with a name (the policy's) beats a silent black hole.
- **Rejected: keeping upstream's accept-and-forge for unroutable ICMP.** A ping
that "works" without leaving the router is the inverted lie this project keeps
deleting (D17's fiction purge, D23's dead WireGuard, D24's obedient-client
leak).
**Proven, and not proven, said plainly.** The cold start is PROVEN, not assumed:
`TestIntegrationL3TunInboundStarts`
(`shater/generate/l3_integration_linux_test.go`, run as root with NET_ADMIN and
`/dev/net/tun`, PASS) drives an `l3_tunnel=1` config through the SLIM registry
(`registry.Context`, not upstream's `include.Context`) under the shipped router
tag set: `box.New` + `Start` accept it, the kernel really ends up with the
`shater-l3` device at the contract MTU 65535 — the assertion the value exists
for, since a kernel that clamped it would silently restore the forged-reply
band — and Close removes it; precisely
the "built with X, verified with Y" gap class D23 exists for (a lost
`tun.RegisterInbound` or a trimmed `with_gvisor` changes no generated byte and
would otherwise surface only on the operator's router). Each layer contract is
pinned besides (`generate` `TestL3Tunnel*`, `netplane` `TestL3Ingress*`,
`route/prematch_icmp_lx_test.go`). Exactly two things remain UNVERIFIED:
(a) the end-to-end path on live hardware — LAN client → prerouting mark →
ip rule → TUN → WireGuard peer → reply back to the client — has not been
exercised on a real router; (b) the steady-state memory cost of the second
gVisor netstack (the `l3-in` TUN beside the WireGuard endpoint's own) is
unmeasured on the target hardware. An indicative figure exists and is only
that: on x86_64 in a container, idle and carrying no flows, peak RSS of a
process that brought the same engine up went from ~26.0-26.8 MB without
`l3_tunnel` to ~28.3-28.7 MB with it over three paired runs — about +2.2 MB.
That was measured on a throwaway harness, not on aarch64, not under load, and
with an empty ICMP NAT table, so it bounds nothing on the router. Neither
item is folded into any claim above.
Consequence: a ping from the LAN either genuinely travels through the tunnel
(WireGuard/AWG, direct) or fails honestly, at every size the router itself can
put into the device — and a router that never opts in renders the pre-feature
plane byte-for-byte (`TestL3IngressOptIn` pins the off-state render). Read
"fails honestly" strictly: above the tunnel MTU a non-DF ping now leaves the
router for real and then times out, because upstream's ICMP NAT does not carry
fragments back either. The one qualifier left is item 2 above — a client that
puts fragments on the wire ITSELF, on a router where conntrack defrag is not
reassembling them first. This paragraph has been overclaimed twice already;
extend it only against a bench result, never against a reading.
## D26 — What the engine cannot carry, the kernel carries: `untunnelable_egress`
Decided 2026-07-26. D25 ended with ESP/AH/GRE/IGMP/SCTP still owned by the D17
policy — that is, with a choice between dropping them and leaking them out the
WAN, never a data plane. This decision gives them one, and deliberately NOT
ours: **`globals.untunnelable_egress` (default empty;
`model.Globals.UntunnelableEgress`) names an existing egress of type
interface/tunnel, and LAN traffic that is neither TCP nor UDP is stamped in
prerouting with that egress's own mark, so the KERNEL routes it out that
egress's device with the kernel's own NAT.** No proxy, no engine, no userspace
stack ever touches the packet — which is exactly why every protocol works.
**Where the engine's boundary actually is — recorded so nobody digs for it
twice.** It is NOT the gVisor stack, and it is not WireGuard: on the forward
path the WG/AWG endpoint never consults gVisor at all. `Endpoint.WritePackets`
(`transport/wireguard/port.go:21-58`) takes the raw IP packet bytes, reads
exactly the IP version and the destination address, and hands
`device.InputPacketRef`s to `wgDevice.InputPackets` — the protocol byte is
never read; on the way back (`port.go:127-157`) `returnDeviceWrapper.Write`
offers every decrypted packet to `returnPath.ReturnPackets` first and only the
unconsumed remainder falls through to the gVisor device. gVisor serves
`DialContext`/`ListenPacket` — traffic the ENGINE originates — while forwarded
traffic bypasses the stack in both directions, indifferent to protocol. The
real ceiling sits one step earlier, in sing-tun's `ForwardDispatcher`:
`parseTransport` (`flow_parse.go:106-153`) sets `hasFlow` for exactly TCP, UDP,
ICMPv4 Echo/EchoReply and ICMPv6 EchoRequest/EchoReply — a packet of any other
protocol is never dispatched as a flow — and `createFlow`
(`flow_dispatch.go:325`) builds its NAT through
`allocateSelector(packet.protocol, …, packet.source.Port())` (line 355), which
needs a port-like selector that ESP/AH/GRE simply do not have (SCTP has ports,
but the parser above never grants it a flow either). Tailscale documents the
same frontier for its own userspace mode — "Any IP protocol other than TCP or
UDP (such as SCTP) is not supported in userspace mode… All IP protocols are
supported" in kernel mode
(https://tailscale.com/docs/reference/kernel-vs-userspace-routers) — useful as
external corroboration of where userspace data planes generally end, though OUR
boundary is the dispatcher, not the stack. The kernel egress was therefore
chosen not because userspace "cannot" in principle, but because the kernel
delivers all protocols with zero new code on the hot path.
**The mechanism already existed; the feature is one binding and one marking
step.** `addEgressRouting` (`netplane/apply.go`) has always installed, for
every interface/tunnel egress, an `ip rule fwmark <EgressMark> lookup
<EgressTable>` plus a `default dev <device>` route in that table — per-rule
egress selection rides on it. The only missing piece was that nothing ever
marked non-TCP/UDP traffic: `untunnelable=direct` merely ACCEPTED it in the
forward chain, so it left over the main table, i.e. the WAN.
`UntunnelableEgressBinding` (`netplane/nft.go`) resolves the option to the
egress's index, its OWN mark and its OWN device — deliberately no third
mark/table pair to keep coherent — and the prerouting chain stamps that mark on
the untunnelable protocols. A name that does not resolve to an interface/tunnel
egress with a device renders nothing and is reported: the D17 policy stays in
sole charge, which is the fail-closed reading of a typo.
**Why `l4proto != { tcp, udp }` is safe here when D25 banned it.** D25 rejected
the broad filter because the receiving side was the `ForwardDispatcher`, which
classifies nothing beyond TCP/UDP/ICMP echo — a marked ESP packet would enter
the TUN and vanish, a black hole wearing a tunnel's name. Here the receiving
side is the kernel, which forwards ANY IP protocol and NATs what it has
machinery for: SCTP carries ports and NATs like TCP/UDP; GRE is NATed only
through the PPTP helper keyed on the call-id — the kernel's own comment calls
GRE "generally not very suited for NAT, as it has no protocol-specific part as
port numbers" (`net/netfilter/nf_conntrack_proto_gre.c`); ESP/AH pass as plain
routed IP. Nothing on this path can silently swallow a protocol it does not
understand, which was the entire objection.
**Order against D25: the L3 ingress claims ICMP first.** With `l3_tunnel` on,
LAN ICMP is marked into the engine's TUN before the egress carrier is consulted
— the engine path routes ping by the operator's rules, which a kernel egress
cannot do — and only the remaining protocols go to the egress. With `l3_tunnel`
off, ICMP goes to the egress with everything else. In both shapes marked
traffic is settled by ROUTING before the forward chain speaks, so the D17
policy now governs exactly the failure case — the rule or route that did not
come up — the same division D25 already established for the L3 mark.
**What the feature refuses to promise — and the operator text refuses with it
(`shater/apply/warnings.go`, the egress-carrier note).** (a) It is not a tunnel
per se: the option accepts any interface/tunnel egress, and on the target
routers a WireGuard device is the exception (`kmod-wireguard` is usually
absent) while a second WAN is routine. Through a WireGuard egress this
genuinely is a tunnel; through a second WAN it is simply another uplink, and
the destination sees that uplink's real address. No text, comment or doc line
may call it a tunnel unconditionally. (b) It does not revive IPTV: IGMP is
LAN-side multicast group management, WireGuard is L3 point-to-point and carries
no multicast, and multicast never crossed this router under any setting —
routing IGMP out an egress restores nothing, and no wording may hint otherwise.
(c) IPsec through NAT-T never needed it: RFC 3948 encapsulates ESP in UDP/4500,
so a modern IPsec client behind NAT is ordinary UDP that already follows the
routing rules; the raw-ESP case this feature carries is the no-NAT-T remainder.
- **Rejected: teaching the engine these protocols.** Extending `parseTransport`
and the selector NAT upstream would be new hot-path code in an
actively-maintained adversarial area, for protocols the kernel already
forwards for free — and for ESP/AH/GRE there is no port-like selector to NAT
by in the first place.
- **Rejected 2026-07-26: carrying them through the userspace AWG endpoint
site-to-site, with no NAT at all.** This is the alternative the "no port-like
selector" line above does NOT dispose of, and it is written down because the
obvious reading of that line — "impossible" — is wrong and would be
re-derived. The endpoint is already protocol-blind in both directions
(`transport/wireguard/port.go:21-58`, `:127-157`), so an ESP packet could be
forwarded UNTOUCHED, keeping the LAN client's own source address, and the
reply would come back addressed to that client and need only be written to the
TUN. No selector, no NAT, every protocol. It needs two things we declined to
take on: lx-owned code in the forward hot path, bypassing `ForwardDispatcher`
on both legs — precisely the surface CONSTITUTION §2 exists to keep small on
an actively-maintained upstream — and a SERVER-side prerequisite (our LAN
prefix in the peer's `AllowedIPs`, plus a route back), which turns a router
option into a deployment contract. The kernel egress above buys the same
protocols with zero hot-path code, so this stays a design on file, not a gap.
- **Rejected: a dedicated mark/table pair for the carrier.** `addEgressRouting`
already binds `EgressMark`/`EgressTable` to the device; a third pair would be
a second copy of the same route that could drift from the first.
**Not verified, said plainly.** The end-to-end path — LAN client → prerouting
mark → ip rule → egress device → far end and back — has not been exercised with
real ESP or GRE on live hardware. Nothing above claims it has.
Consequence: raw IPsec, PPTP/GRE, SCTP — and ICMP when the L3 ingress is off —
leave through an egress the operator explicitly named, under kernel routing and
kernel NAT, instead of being dropped or silently leaking out the WAN; and with
the option empty (the default) the plane renders byte-for-byte as before, with
the D17 policy in sole charge.
## D27 — The `sing-quic` pin moves forward; two use-after-release defects in the QUIC/HTTP-3 client path
Decided 2026-07-26, after an audit of `common/httpclient`, `dns/transport/quic`
and `transport/v2rayquic`. Two suspicions were put to a test rather than to a
reading. Both turned out to be real, and neither was the resource leak the
suspicion named — both are objects released while still in use.
**1. The HTTP/3 race handed back a response nobody could read
(`common/httpclient/http3_transport.go`).** `roundTripHTTP3Race` ran both racers
on one `context.WithCancel` child and its `drainRemaining()` called `cancel()`
before returning the WINNER. quic-go and net/http both reset a request's stream
when its context is cancelled, so the caller received a `*http.Response` whose
body died mid-read. Measured, not inferred:
`H3_REQUEST_CANCELLED (local) (read 2687 of 65536 bytes)`. The path is taken
whenever there is no cached HTTP/3 connection and the request is replayable —
that is, the FIRST request to every host, plus every request after an idle
close. Anything configured with `"version": 3` and no
`disable_version_fallback` was affected: subscription fetches, remote rule-set
downloads, URLTest probes.
Fixed by giving each racer a context of its own. Losers are cancelled where the
old code cancelled everything; the winner's `cancel` travels with its body and
fires on `Close`. Pinned by
`common/httpclient/http3_race_lx_test.go` — one test per winner, and the
loser-is-torn-down assertion so the fix cannot be "stop cancelling" either.
**How defect 2 is pinned, and why it takes two tests.** The success path is
visible on the wire, so `TestHTTP3ExchangeRequestBufferOutlivesRoundTrip`
compares the query the server received with the query we packed. The failure
path is NOT visible on the wire — see the `CancelWrite` note below — so
`TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory` pins the CAUSE instead of
the symptom: it tags a query with a random needle, runs an exchange that fails
(a server that never answers; a context already cancelled) and then drains the
buffer pool on the same goroutine `RoundTripOpt` ran on, demanding the needle is
not in it. Both carry a control that must produce a DIRTY result first — the
poison must be shown to reach a released buffer, the scan must be shown to find a
query that really was packed into pool memory — because a clean verdict from an
instrument that cannot produce a dirty one is not evidence. The second test
encodes the chosen design; a future guarded-pool implementation would fail it on
purpose and would have to replace it, saying so.
**2. DoH3 sent the DNS server whatever the next caller put in a recycled buffer
(`dns/transport/quic/http3.go`).** `Exchange` packed the query into a POOLED
`buf.Buffer`, handed `bytes.NewReader` over it to the request, and called
`requestBuffer.Release()` the moment `RoundTrip` returned. But http3's
`doRequest` writes the request body on a goroutine of its own and returns as
soon as the response HEADERS arrive — the body is still being read. The test
holds that window open (a 2 KB server stream window, an answer written before
the body is read) and shows the query on the wire diverging from the query we
packed **at exactly offset 8192** — `bodyCopyBufferSize`, the amount quic-go had
already copied out before the buffer went back to the pool. Everything past that
was the next pool user's memory, sent to the resolver. That is a data race and a
small memory-disclosure primitive, not a slowdown.
**The first fix for this was wrong, and the way it was wrong is the point.** It
transferred ownership to the transport: the body became a `pooledRequestBody`
whose `Close` released the buffer, justified as "`http3.Transport` closes the
request body on every path, hence the `sync.Once`". That sentence is true about
HOW MANY TIMES the body is closed and says nothing about WHEN — the exact shape
of dishonesty this document keeps having to name. Review caught it before
release. On the failure path `RoundTripOpt` (`http3/transport.go:167-173`) closes
the body the moment `doRequest` returns, and `doRequest`
(`http3/client.go:338-341`) waits only on the request-CANCELLATION watchdog —
`close(reqDone); <-done` — never on the goroutine writing the body. Nothing in
quic-go ever joins that goroutine. So `Close` is not a handoff point, and the
`sync.Once` prevented a double `Release` while doing nothing about a read after
one.
**One correction to the review's severity, for the record:** on the failure path
the damage stops at the data race. Every `ReadResponse` error branch
(`http3/stream.go:325`, `:336`, `:343`, `:363`) calls `str.CancelWrite` BEFORE
`RoundTripOpt` closes the body, so the bytes the writer reads out of the recycled
buffer are thrown at an already-cancelled stream and never reach the resolver.
The memory-disclosure primitive is the SUCCESS path only. The failure path is
"merely" a read of memory owned by somebody else — still undefined behaviour,
still a `-race` finding, still not shippable.
**Fixed instead by not sharing at all:** the query is packed with `Pack()` into
memory the request body owns outright, and no pooled buffer is involved. The
alternative on the table — a lock around `Read` and `Close` so reads after
release return an error — would also be correct, and was rejected because it
keeps a released-but-referenced object alive and leaves a live invariant for the
next person to break, which is now twice in one day that an assumption about
quic-go's internal lifetimes has been wrong.
The cost turned out to be negative, measured rather than assumed: `Pack` runs
87 ns/op at 64 B and 1 alloc against 108 ns/op at 64 B and 1 alloc for the pooled
version, because `buf.NewSize` allocates the `Buffer` struct itself — the same 64
bytes — and then adds `Get`/`Put` on top. **The pool was never saving an
allocation on this path.** The response buffer stays pooled: it is read and
unpacked before `Exchange` returns and nothing outlives it.
Both files thereby DIVERGE from upstream again, six hours after `0a6689b29`
made them byte-identical on purpose. That was the right call then and this is
the right call now; upstream carries defect 2 in `dns/transport/https.go` as
well (same shape, HTTP/1.1 and HTTP/2 write bodies asynchronously too) and that
file was left alone — it is outside the audit's scope, and it is written down
here so the next person finds it instead of rediscovering it.
**3. The pin moved: `sing-quic` v0.6.2-0.20260525051024 -> v0.6.4-0.20260709034545.**
`quic.go` — `Dial`/`DialEarly`/`CreateTransport` — is byte-identical across the
two, so the packet-conn ownership fix in `0a6689b29` is NOT duplicated by the
bump and is not made redundant by it: quic-go still does not own the socket, and
we still close it. What the newer module does carry is the OTHER half of the
same family, and we had taken only our half:
`clientConn.Close()` in `tuic/`, `hysteria/` and `hysteria2/` now sets a past
write deadline after `Stream.Close()`, word for word the fix
`transport/v2rayquic/stream.go` already had — quic-go's `Stream.Close` does not
release a write blocked on flow control. We ship tuic and hysteria2, so on the
old pin every such close could park a goroutine for the life of the process.
It also brings a QUIC-connection-death watchdog and a handshake deadline to the
hysteria clients.
**The cost, measured, not estimated:** six new indirect modules (`libp2p/go-nat`
and its UPnP/NAT-PMP/gopacket tail) for hysteria2's realm port mapping, and
**+256 KiB exactly** on the stripped aarch64 `shaterd` (26 542 242 ->
26 804 386 bytes, +0.99%), before UPX. The port-mapping path is unreachable from
anything `shater/generate` emits — `realm` is only built when the JSON names it,
and it never does — so the growth is dead weight, but it is small dead weight
next to a goroutine leak on the two QUIC protocols we actually ship. `upstream/lx`
is already on this pin, so keeping the old one would mean fighting every rebase.
**Not verified here:** the tuic/hysteria2 close fix is read from the module diff,
not exercised — those packages are outside this audit's file set and testing them
needs a live tuic/hysteria2 server. `go test ./common/... ./dns/... ./transport/...
./protocol/...` is green under the shipped tag set apart from
`common/windivert`'s `TestIntegration*`, which want Windows SCM access and fail
on any developer machine, bump or no bump.
## D28 — The L3 TUN is a reclaimable SLOT, not a name: a fixed device made every apply fatal
Decided 2026-07-26. D25 gave the L3 ingress one fixed device, `shater-l3`. On the
production router that single name made **every** configuration change with
`l3_tunnel=1` an outage, three times in a row:
```
19:10:33 reconcile failed: start inbound/tun[l3-in]: open tun: TUNSETIFF: device or resource busy
19:14:02 start instance failed and could not restore previous config; engine stopped
19:14:38 reconcile failed: TUNSETIFF: device or resource busy
```
followed by `plane: hold` — the fail-closed ruleset — i.e. the whole house
offline until somebody ran `/etc/init.d/shater restart` by hand.
**The mechanism is a collision between two GENERATIONS of the engine, and the
fatal part is where the collision lands.** An apply builds a fresh box and starts
it; with one name that box must open the device the outgoing box still holds.
That alone is a failed apply. What turned it into an outage is the recovery path:
`closeOldThenStart` answers a failed start by rebuilding the PREVIOUS config —
and that config names the same device, so **the rescue failed for exactly the
reason the rescue was needed**. A recovery path must never depend on the resource
whose contention it is recovering from; that sentence, not the device name, is
the decision here.
**Two slots (`shater-l3a` / `shater-l3b`), chosen by the ENGINE at box-build
time.** Not by `generate`, and that is load-bearing rather than incidental:
`generate` runs on every reconcile including the once-a-minute no-ops, and its
output is what `Apply` hashes to decide whether anything changed. A device name
that alternated there would change the hash every minute and rebuild the whole
engine forever; a name that tracked "whichever slot exists" would hand the BUSY
one to every real change. Only the engine knows it is building a new generation.
So `generate` emits `netplane.L3DeviceBase` as a **placeholder that is never
created**, and `engine.newBox` substitutes a slot on a COPY of the options —
after the hash, so the stored config stays canonical (`engine/l3slot.go`).
**Rotation alone is NOT the fix, and that was measured, not reasoned.** The
two-slot build survived five applies of five different kinds and then failed on
4 of 10 back-to-back changes with the original outage in full. A retired
generation does not hand its device back when its replacement is adopted:
`Box.Close` walks its subsystems under a budget and the TUN dies with the last
fd. One apply gives it time; two inside that window do not. Rotation widens the
race by one generation — a better outage, not the absence of one.
**So an occupied non-current slot is DELETED, not waited for** (`L3SlotFor`).
The running generation's slot is excluded first and never touched; every other
slot belongs to a retired generation that is not in the L3 routing table and is
carrying nothing, so taking its device away is safe and, if anything, helps the
close already in flight. There is deliberately **no bounded wait**: waiting on an
asynchronous kernel teardown is the race this design removes, and adding it back
as a "safety net" would only make the failure intermittent.
**The firewall never learns which slot is live.** Our forward accepts and the fw4
zone match by PREFIX — `iifname "shater-l3*"` / `oifname "shater-l3*"`, and
`list device 'shater-l3*'` in uci-defaults. Verified on ImmortalWrt 25.12.1
(kernel 6.12.94, nftables 1.1.6) that both forms validate AND load, and that fw4
compiles the wildcard into exactly those matches. So the rendered ruleset is
byte-identical across a swap: no nft reload, and no window in which the accept
names a device that is already gone. Routers seeded by a pre-slot build are
migrated in place (`migrate_l3_zone_wildcard`); without it an upgrade would
silently go back to fw4 dropping the forward.
**A2 — turning the feature off used to leave the plane installed.** `addL3Routing`
returned early when `l3_tunnel=0`, so after switching it off the router still had
the device, `ip rule fwmark 0x2080 lookup 8200` and table 8200. Nobody else
removes them: the engine's new config simply has no TUN inbound. The surviving
device is also the commonest way back into the EBUSY above. The disabled branch
now removes rule, table and device, symmetrically with `removeEgressRouting`, and
`TeardownRouting` sweeps every name including the legacy `shater-l3`.
**Two smaller lies found while proving the above, both measured on the stand and
both fixed here.** `ip -6 route flush table N` does not remove a non-unicast
route while the v4 flush does, so (a) the floor survived the flush and the next
`add` answered `File exists`, which was reported as a CRITICAL "this table has NO
fail-closed floor, traffic can leave over the plain WAN" — on every single apply,
about a table whose floor was sitting right there; and (b) teardown left that
floor behind. An already-present floor is now success, and teardown deletes it
explicitly. (b) cannot misroute anything — no rule points at the table — but a
teardown whose result is not "the table prints nothing" is one nobody can verify.
**Verified on the stand** (`local_openwrt`, ImmortalWrt 25.12.1, kernel 6.12.94 —
the router's revision), before-and-after with binaries built from the same tree:
the pre-fix binary reproduces the production failure on the "edit a node URI and
apply" round (engine stopped, `plane: hold`) and leaves device + rule behind on
`l3_tunnel=0`; the fixed binary survives all five apply kinds, 12 back-to-back
changes, and leaves nothing at all — no device, no rule on either family, both
tables printing empty.
**Known fragility, deliberately NOT addressed here.** Our `ip rule` priorities
are whatever the kernel hands out (the L3 rule lands at 32764, counting down from
32765), so the ORDER of our rules between reboots depends on insertion sequence.
It is not a live defect — the marks are disjoint, each rule catches its own, and
all of them end up above `main` — but it is luck, not design. Moving to explicit
`pref` values touches every existing rule and needs a migration for rules already
installed on implicit numbers; that is its own piece of work, not a rider on this
one.
## D29 — ByeDPI is removed: the presets it replaced were not weak, they were broken
Decided 2026-07-27. **This reverses the ByeDPI half of [D13](#d13--external-dpi-bypass-tool--byedpi-a-socks-egress-not-zapret).** The
`byedpi` egress kind, the `openwrt/byedpi` package (`ciadpi`), the readiness
endpoint `GET /api/byedpi` and the panel's readiness plate are all deleted. The
zapret half of D13 is untouched: zapret was rejected for reasons that have
nothing to do with this and stays rejected.
**Why it was added.** D13 read: the native route-action presets
(`tls_fragment` / `tls_record_fragment` / `tls_spoof`) "cover the light 'just
fragment the ClientHello' case", and ByeDPI "is what we add for the stronger
methods the engine lacks". That sentence was written from observed behaviour —
the native presets were tried against a real ISP and did not get through — and
the conclusion drawn from it was that the METHOD was too weak.
**Why it is removed.** The method was never tried. `common/tlsfragment/conn.go`
chose the split point by dropping a number of labels equal to the number of DOTS
in the whole name — and a name always has one more label than it has dots. So
the cut always landed inside the **first** label. For `www.youtube.com` it split
`www` and handed `youtube` to the wire in one intact piece, which is precisely
the word the DPI matches on. Measured against the live ISP: of six blocked
names, exactly one got through — `youtube.com`, the one whose first label IS the
blocked word. That is not a weak desync, it is a desync aimed at the wrong three
characters, and every conclusion drawn from its failure rate was a conclusion
about our own defect.
The fix is one expression. With it, the built-in presets do the job that the
external process was brought in to do, and the external process is 100 KB of
binary, a second procd service, a second UCI file, a port that agreed with our
egress by hand-written comment only, a readiness prober, a five-state service
model, a per-egress port cross-check and a panel plate — all to work around a
bug in fifteen lines of our own code.
**So this is not "ByeDPI turned out to be bad".** It is a good tool. It turned
out not to be needed, and the reason we thought it was needed was ours.
**What happens to a config that still says `type 'byedpi'`.** Nothing is
migrated and nothing is rewritten. The kind stays unbuildable, which means it
stays **fail-closed**: no outbound, no mark, no `ip rule`, no routing table, so
everything bound to that egress is blocked rather than released onto the plain
WAN. The one thing that changes is what the operator is TOLD. `byedpi` is
recorded in `model.RetiredEgressTypes` — a closed, positive table read by both
`ValidateEgresses` and the generator — and draws a sentence naming the removal,
the replacement (`direct`/`interface` with `dpi 'record'`), the honest caveat
that which preset defeats a given ISP is not something we can promise, and how
to take the dead package off the router.
A migration rewriting `byedpi` → `direct` was considered and rejected. It is the
only rewrite that leaves the egress routing at all, and it would silently
convert a blocked egress into a live plain-WAN path with the router's real
address — the exact leak class this codebase refuses everywhere else, performed
by an upgrade, on a config nobody touched. `CurrentSchemaVersion` is therefore
**not** bumped either: no stored field changes meaning, nothing is migrated, and
a bump would only make configs written by this build unreadable to an older
daemon (Migrate refuses a newer schema) for no gain.
**The package is not uninstalled by this change.** Dropping it from the feed does
not remove it from a router it is already on. `apk del byedpi` does, it is safe,
and it is written down in `INSTALL.md` beside the update command — which loses
its fourth name: `apk upgrade shaterd shater-core luci-app-shater`.
**When ByeDPI would win again.** If a fixed `tls_fragment`/`tls_record_fragment`
still fails against an ISP that fake/disorder/oob/autottl defeats, the argument
in D13 for choosing ByeDPI over zapret is still the right argument and this
decision is the one to revisit — with a measurement of the FIXED presets first,
which is the step that was skipped last time.
+39 -3
View File
@@ -6,12 +6,44 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
## Proxy engine & protocols (from the sing-box fork)
- **[MVP]** VLESS, VMess, Trojan, Shadowsocks, WireGuard, Reality/XTLS.
- **[MVP]** **AmneziaWG 2.0** (I1–I5 CPS decoy packets) — a driving requirement.
- **[T1]** Hysteria2, TUIC, ShadowTLS, XHTTP, MASQUE/CONNECT-IP (Cloudflare WARP).
- **[MVP]** Hysteria2, TUIC (`hysteria2://`/`hy2://`/`tuic://`, `shater/parse`),
XHTTP transport — all shipped: the router tag set carries `with_quic` and
`with_xhttp` and `shater/registry` registers them (`scripts/router-tags.sh`,
`buildtags.Features`).
- **[T1]** ShadowTLS — half-built: `shater/generate` emits it and `shater/registry`
registers it, but no parser produces one (there is no `shadowtls://` share link
and no subscription path), so a config cannot reach it today.
- **NOT SHIPPED** MASQUE/CONNECT-IP (Cloudflare WARP). `masque` appears nowhere in
`shater/parse`, `shater/generate` or `shater/model`, and `shater/registry` names
it among the upstream types it deliberately does not register (~6 MB of binary
and resident RAM). The engine fork can build it; this product does not.
- **[MVP]** Transports: TCP/WS/gRPC/HTTPUpgrade/H2/QUIC as upstream provides.
## Transparent proxying & routing
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** **L3 ingress for ICMP** (`globals.l3_tunnel`, opt-in, default off):
LAN ping travels THROUGH the tunnel instead of being dropped or answered by a
forged local reply. The engine opens a dedicated TUN (`shater-l3`, gVisor
stack, `auto_route` off); nft marks LAN icmp/icmpv6 only and a scoped
`ip rule` routes it in — the TPROXY plane and the main routing table stay
untouched (D25). Carried only by L3-capable egresses (WireGuard/AmneziaWG,
direct); ICMP routed to vless/vmess/… is honestly dropped, never faked.
Ceiling is upstream sing-tun's: ICMP echo only — Windows tracert works, IPv6
traceroute shows just the destination; ESP/AH/GRE/IGMP stay with the
`untunnelable` policy (D17) unless `untunnelable_egress` carries them (D26).
- **[MVP]** **Kernel egress for untunnelable protocols**
(`globals.untunnelable_egress`, opt-in, default empty): names an existing
interface/tunnel egress, and IPsec (ESP/AH), PPTP/GRE, SCTP — everything that
is neither TCP nor UDP, plus ICMP when the L3 ingress is off — is routed out
that egress's device by the KERNEL with kernel NAT, reusing the egress's own
fwmark/table from `addEgressRouting`; the proxy never sees a byte, which is
why every protocol works (D26). What that buys depends on the device: a
WireGuard interface really is a tunnel, a second WAN is just another uplink
whose real address the destination sees. It does not revive multicast IPTV,
and UDP-based VPNs (WireGuard, OpenVPN-UDP, IPsec NAT-T) never needed it —
they follow the routing rules as before. The `untunnelable` policy (D17)
keeps only the failure case: a route that did not come up.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
@@ -87,8 +119,12 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
## Reliability ("железно")
- **[MVP]** Fail-closed kill-switch (dead group → block, never silent direct leak);
IPv6 dropped when disabled.
- **[MVP]** Atomic apply with engine + `nft -c` validation; commit-confirm
auto-rollback to last-good.
- **[MVP]** Atomic apply with engine + `nft -c` validation. Commit-confirm
auto-rollback to last-good is built and works, but it is **opt-in and ships
OFF**: `DefaultGlobals()` leaves `ConfirmTimeout` at 0, the shipped
`/etc/config/shater` says `confirm_timeout '0'`, and `apply.ArmRollback` returns
at once on a non-positive timeout. Until an operator sets a window, an apply on
a stock box has no net under it — and `shaterd apply` says so.
- **[MVP]** Idempotent reconcile from hotplug/boot under flock; restart engine only
on real config change; management-bypass (SSH/LuCI/LAN) always exempt.
- **[MVP]** Own nft table `inet shater` + own marks/tables; never touch fw4.
+190 -22
View File
@@ -83,14 +83,13 @@ probe case in `shater/generate/shipped_tags_linux_test.go`.
## 2. Packages
Four OpenWrt packages live under `openwrt/`:
Three OpenWrt packages live under `openwrt/`:
| Package | Arch | What it ships |
|--------------------|-----------|---------------|
| `shaterd` | per-arch | **Prebuilt** static `shaterd` binary → `/usr/bin/shaterd` (this is the ship artifact from step 1). |
| `shater-core` | all | procd init (supervises `shaterd run`), cron, hotplug, sysctl, inert default UCI. `DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full`. |
| `shater-core` | all | procd init (supervises `shaterd run`), the boot armor (§4), cron, hotplug, sysctl, inert default UCI. `DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle`. |
| `luci-app-shater` | all | Thin LuCI launcher: mini dashboard + token-handoff "Open panel" button. `DEPENDS:=+shater-core +rpcd`. |
| `byedpi` | per-arch | *Optional* ByeDPI (`ciadpi`) local desync SOCKS proxy for a `type='byedpi'` egress. |
### Why `shaterd` is a prebuilt-binary package
@@ -138,11 +137,9 @@ builds. `ci/sdk-build-apk.sh` then **asserts** the produced `.apk` really carrie
it, so a lost variable fails the build instead of shipping a stale version. The
release job asserts the same version again on the published rolling repo (§5.1).
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
installed. Stamping our tag on it would also be a downgrade: every comparator
reads `0.2.7 < 0.17.3` (component-wise, `2 < 17`). Bump its `PKG_RELEASE` by hand
when our packaging of it changes.
All three are versioned this way. There used to be a fourth package carrying its
upstream's own version and therefore exempt from the assertion above; it is gone
(D29), and with it the exception nobody could be expected to remember.
## 3. Install on a router
@@ -160,7 +157,6 @@ apk filenames carry no architecture, so make sure you copied the `.apk` built fo
apk add --allow-untrusted ./shaterd-<ver>.apk
apk add --allow-untrusted ./shater-core-<ver>.apk
apk add --allow-untrusted ./luci-app-shater-<ver>.apk
apk add --allow-untrusted ./byedpi-0.17.3-r1.apk # optional: ByeDPI egress
```
From the repo instead (§5 sets it up once), deps pull the rest in:
@@ -176,11 +172,27 @@ install. Configure nodes/rules (via the LuCI panel or `uci`), then enable and ap
```sh
uci set shater.globals.enabled=1
# The safety net is NOT on by default — see below. 120 s is a window wide enough
# to re-open SSH/LuCI and decide whether the new config is any good.
uci set shater.globals.confirm_timeout=120
uci commit shater
shaterd apply # apply + arm commit-confirm on the running daemon
shaterd confirm # confirm (cancels the auto-rollback)
shaterd apply # apply + arm the auto-rollback for 120 s
shaterd confirm # confirm inside that window (cancels the auto-rollback)
```
> **Commit-confirm ships OFF.** `model.DefaultGlobals()` does not seed
> `ConfirmTimeout`, the shipped `/etc/config/shater` carries
> `option confirm_timeout '0'`, and `apply.ArmRollback` returns immediately on a
> non-positive timeout — so on a stock box `shaterd apply` arms **nothing** and an
> apply that costs you SSH/LuCI access simply stays. The daemon says so rather
> than implying otherwise: the `commit-confirm-off` outcome of `shaterd apply`
> prints *"globals.confirm_timeout is 0, so commit-confirm is switched OFF: this
> apply armed NO automatic rollback"*, and the panel's Overview reads
> `confirm: no auto-rollback`. Set a window (UCI as above, or Settings in the
> panel) if you want the net. Non-obvious detail: the option is written back only
> when non-zero, so an explicit `0` disappears from `/etc/config/shater` on the
> first write — absent and `0` mean the same thing.
`/etc/init.d/shater enable && /etc/init.d/shater start` brings up the procd-supervised
daemon (`shaterd run`), which owns the engine, the `inet shater` data plane, policy
routing, in-process DNS, and the admin panel (default `:8088`). The LuCI app's
@@ -218,6 +230,55 @@ Your `0` is kept: `/etc/config/shater` is a conffile (upgrades never replace it)
the daemon always writes the option back explicitly, so it is never re-enabled by a
default.
### The boot-time fail-closed armor
`shater-core` installs a **third** init script, `/etc/init.d/shater-armor`, and
`30_shater-core` enables it at install time. It exists because `/etc/init.d/shater`
is `START=99`: by then fw4 (19) has loaded `lan -> wan ACCEPT` and netifd (20) has
brought the LAN bridge up, so between link-up and the daemon's first apply the
router forwards LAN traffic to the WAN in the clear — on router hardware with a
UPX-packed binary that is the seconds in which Wi-Fi associates and every client
reconnects. `kill_switch=closed` covered none of it, because the protection lived
inside a process that had not started.
**How it works.** On every apply the daemon persists a copy of its fail-closed
*holding plane* — the same ruleset it installs when the engine is down — to
`/etc/shater/boot.nft`. `shater-armor` runs at `START=21` (after fw4 and netifd),
validates that file with `nft -c` and loads it. When the daemon comes up it
replaces the table atomically, so there is never a moment with no table. Its
`stop()` is deliberately a no-op.
**LAN forwarding is blocked until the daemon applies — management access is not.**
The chain hooks `forward` only, so SSH, LuCI and the admin panel (all `input` hook,
to the router's own addresses) stay reachable **on purpose**: a kill switch you
cannot switch off is a brick. If you see the syslog line
```
fail-closed plane armed from /etc/shater/boot.nft: LAN->WAN forwarding is BLOCKED
until shaterd applies. SSH, LuCI and the admin panel stay reachable.
```
that is the mechanism working, not a fault.
**When it refuses to arm** — each is a state check made at boot, never a record of
something that happened on the way down:
| Condition | Behaviour |
|---|---|
| `/etc/shater/boot.nft` absent | Nothing to do, silent. The file exists only while the last applied config was **both** `enabled=1` **and** `kill_switch=closed`; either being off removes it at the next apply, and an operator-typed `/etc/init.d/shater stop` removes it there and then. Powering off does **not** — and neither does the `stop` a package upgrade issues while the service stays enabled, so being replaced cannot leave the next boot unprotected. |
| the file is empty, or fails `nft -c` | Refuses, logs an error — the LAN is unprotected until `shaterd` starts. |
| `/usr/bin/shaterd` missing, or no `S??shater` symlink in `/etc/rc.d` | Refuses: nothing would ever come along to replace the block with a working data plane. This is what makes an uninstalled or disabled product safe regardless of what the file says. |
| UCI is readable **and** says `globals.enabled` is not `1` | Removes `boot.nft` and does not arm. An **unreadable** UCI is not a refusal — that case is exactly why the armor is a file rather than a query. |
| `nft` not installed | Refuses, logs an error. |
**Turning it off.** The durable off-states are the two the script itself asks
about — `uci set shater.globals.enabled=0 && uci commit shater && shaterd apply`
(the next apply removes `boot.nft`), or `/etc/init.d/shater disable`. A bare
`/etc/init.d/shater stop` typed at the shell also removes the file, but it is not
durable: `S99shater` is still linked, so procd starts the daemon again on the next
boot. To remove just the armor and keep the stack: `/etc/init.d/shater-armor
disable`.
## 5. The signed apk repo (the normal install path)
OpenWrt/ImmortalWrt **25.12** packages with Alpine's **apk**: `.apk` files, a
@@ -281,7 +342,6 @@ echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/a
# 3) refresh + install (shaterd pulled in as a dependency).
apk update
apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # optional: ByeDPI desync egress
```
### 5.3 Updating
@@ -293,7 +353,7 @@ unrelated system packages. Always name ours:
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
apk upgrade shaterd shater-core luci-app-shater
```
apk-tools 3 documents exactly this behaviour for `apk upgrade`: *"When no
@@ -303,16 +363,118 @@ dependencies."* The equivalent form, which additionally re-pins the packages in
`world`, is:
```sh
apk add -u shaterd shater-core luci-app-shater byedpi # -u = --upgrade
apk add -u shaterd shater-core luci-app-shater # -u = --upgrade
```
Drop `byedpi` from either list if you never installed it. Check what you are on
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
Check what you are on
with `apk list -I shaterd shater-core luci-app-shater` — the version reads
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
as `0.2.0-r3` and `apk update` offered nothing). Rolling vs pinned repo URL —
§5.1.
#### If this router still has `byedpi` installed
Older releases shipped a fourth, optional package — `byedpi` (the `ciadpi` local
desync proxy) — behind an egress of `type='byedpi'`. Both are **removed from the
product** ([D29](DECISIONS.md#d29--byedpi-is-removed-the-presets-it-replaced-were-not-weak-they-were-broken)):
the desync it provided is done by the engine's own `dpi` presets, which were
failing for a defect of ours that is fixed.
Dropping it from the feed does **not** take it off a router it is already on —
nothing here uninstalls anything. Remove it by hand:
```sh
apk del byedpi
```
That is safe. No shater package depends on it, nothing in `shaterd` looks for the
`ciadpi` binary any more, and it takes `/etc/init.d/byedpi`, `/usr/bin/ciadpi`
and (unless you edited it) `/etc/config/byedpi` with it. Leaving it installed is
also harmless — it is then simply a service nothing routes to.
If an egress in `/etc/config/shater` still says `option type 'byedpi'`, it is
**blocked, not leaking**: the daemon builds no outbound for the kind, so every
node, group and rule bound to that egress stops rather than going out over the
plain WAN. The panel and the apply warnings name it and say what to change it to
— `direct` (or `interface`) with `dpi 'record'`. Nothing is migrated for you on
purpose: the only automatic rewrite that would keep the egress routing is the one
that would silently put that traffic on the plain WAN.
### 5.4 Downgrading — going back to an older build
Per-version releases live forever (`apk-vX.Y.Z-<arch>`, §5.1), so the way back is
always open. It is a different command from updating, because **`apk upgrade`
never downgrades** — that is not a policy of ours, it is what the solver does.
**Read this first: an older build refuses to write a newer config.** The UCI
schema version is stored in `globals.schema_version`, and a build that finds a
schema NEWER than it understands refuses every config write — the panel, the
6-hourly subscription refresh and the profile watcher all stop persisting, with a
message naming both versions. That refusal is the recoverable outcome: your
`/etc/config/shater` is untouched, and installing the newer package again brings
everything back. The alternative would have been an older build quietly rewriting
the file in its own, poorer form. The file as it stood before the first change of
any build is also kept at `/etc/shater/config.pre-v<schema>.bak`.
So: downgrade across a schema bump only as a temporary measure, and expect the
box to hold the config it has rather than accept edits.
```sh
# 0) note where you are, and keep it.
apk list -I shaterd shater-core luci-app-shater
# 1) repoint the repo file at the PINNED per-version release you want.
echo "https://git.qomar.pw/omar/shater/releases/download/apk-v0.2.9-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
# 2) refresh, and read the exact version strings that feed offers.
apk update
apk list shaterd shater-core luci-app-shater
# 3) install them BY NAME with an explicit version. `=` is what downgrades.
apk add shaterd=0.2.9-r1 shater-core=0.2.9-r1 luci-app-shater=0.2.9-r1
```
Going back up afterwards is two commands, not one — the `=` form leaves a
**pinned constraint** in `/etc/apk/world` (`shaterd=0.2.9-r1`), and a pin outranks
an upgrade:
```sh
# repoint /etc/apk/repositories.d/shater.list back (rolling, or the newer tag)
apk update
apk add shaterd shater-core luci-app-shater # drops the =version pin
apk upgrade shaterd shater-core luci-app-shater # moves the packages
```
**Measured, not inferred** — on the testbed VM (ImmortalWrt 25.12.1
`r37978-cd0a06bfd3fd`, apk-tools 3.0.5, x86_64), against the real
`apk-v0.2.9-x86_64` and `apk-v0.2.10-x86_64` feeds, in an isolated `--root`
sandbox so nothing on the box moved:
| Command, with only v0.2.9 in the shater feed | What apk actually did |
|---|---|
| `apk upgrade shaterd` | nothing — stayed on `0.2.10-r1` |
| `apk add shaterd=0.2.9-r1` | `Downgrading shaterd (0.2.10-r1 -> 0.2.9-r1)`, and `world` became `shaterd=0.2.9-r1` |
| `apk add shaterd` (after the pin) | pin cleared; the installed version did **not** move |
| `apk upgrade shaterd` (pin cleared, feed back at v0.2.10) | `Upgrading shaterd (0.2.9-r1 -> 0.2.10-r1)` |
| `apk add shaterd=0.2.9-r1` while the feed carries only 0.2.10 | `ERROR: unable to select packages: shaterd-0.2.10-r1: breaks: world[shaterd=0.2.9-r1]` — nothing installed. Repoint the repo FIRST. |
**`apk upgrade -a` is not the way to do this**, even though it does downgrade.
`--available` reconciles against the repositories rather than against the
installed versions, and naming our packages does **not** keep it to them: the same
run on the testbed reported
```
(22/27) Downgrading shaterd (0.2.10-r1 -> 0.2.9-r1)
(23/27) Downgrading shater-core (0.2.10-r1 -> 0.2.9-r1)
(24/27) Downgrading luci-app-shater (0.2.10-r1 -> 0.2.9-r1)
```
together with `luci-app-attendedsysupgrade`, `luci-i18n-firewall-zh-cn` and two
more unrelated packages rolled back to whatever the distfeed snapshot holds. Use
the `=version` form, which touched exactly the three packages named.
### BananaWRT `25.12-mtk-vendor` compatibility
The mtk-vendor channel (base: `SuperKali/immortalwrt-mt798x-rebase`, branch
@@ -320,9 +482,15 @@ The mtk-vendor channel (base: `SuperKali/immortalwrt-mt798x-rebase`, branch
**`aarch64_cortex-a53`**, and its images even point their distfeeds at vanilla
`downloads.immortalwrt.org/releases/25.12-SNAPSHOT` — so packages built with the
vanilla ImmortalWrt 25.12 filogic SDK install cleanly; no SuperKali-special SDK
is needed. We ship **no kmods** (shaterd is a static Go binary, byedpi plain C),
so the vendor 6.6 kernel is irrelevant to our packages; the kmod *dependencies*
of shater-core (`kmod-nft-tproxy`, `kmod-nft-socket`, plus `ip-full`) are
already **baked into the BananaWRT mtk-vendor image** (verified in its
`config.buildinfo`). On a self-built 25.12 image, make sure those kmods come
from the image's own kernel build.
is needed. We ship **no kmods** (shaterd is a static Go binary),
so the vendor 6.6 kernel is irrelevant to our packages.
What was actually checked in the BananaWRT mtk-vendor `config.buildinfo` is
`kmod-nft-tproxy`, `kmod-nft-socket` and `ip-full` — those three are baked into
the image. `shater-core` also depends on `kmod-tun`, `nftables-json` and
`ca-bundle` (added later; see the annotated `DEPENDS` in
`openwrt/shater-core/Makefile`), and **those were not part of that check**. They
are ordinarily present on a stock image — apk will pull whatever is missing from
the distfeeds — but if you install offline or from a slimmed image, verify them
yourself. On a self-built 25.12 image, make sure the kmods come from the image's
own kernel build.
+88 -14
View File
@@ -28,6 +28,14 @@ openwrt/
luci-app-shater/ # thin LuCI launcher [Phase 3]
```
> That block is the Phase-2 **plan**, kept because the wave assignment below reads
> from it. The tree that shipped is flatter — `shater/` holds `alert apply
> buildtags cmd devices engine generate logsink model netplane panel parse
> registry stats subscribe` — with rulesets, schedules, profiles, backup and
> migration living inside `model/` and `generate/` rather than as packages of
> their own, and with **no `preset/`**: v0.2 has no preset subsystem at all (see
> the `config preset` note in the schema section).
Build order / waves (parallel agents must own DISJOINT dirs, build only their own
package, and never edit `go.mod`):
- **Wave 1 (contract):** `model/` (+ uci reader + migrate). Everything imports it.
@@ -76,8 +84,7 @@ flock. Commit-confirm/rollback and read verbs can follow, but Teardown must be h
`procd_set_param file` watch); `reload_service`→start/stop; `service_triggers`
reload-trigger "shater"; `stop` sends SIGTERM (honest teardown). `init.d/shater-cron`
(sub/ruleset/schedule due + reconcile). `uci-defaults/30_shater-core` (seed rt_tables
8192, enable inits, seed disabled presets, `model.Migrate`, sysctl from
`netplane.SysctlConf`). `hotplug.d/iface/99-shater` (debounced `shaterd reconcile`).
8192, enable inits, `model.Migrate`, sysctl from `netplane.SysctlConf`). `hotplug.d/iface/99-shater` (debounced `shaterd reconcile`).
`sysctl.d/99-shater.conf` (from `netplane`). Default `/etc/config/shater` conffile.
---
@@ -136,7 +143,7 @@ reload-trigger "shater"; `stop` sends SIGTERM (honest teardown). `init.d/shater-
- **conns.go** — `ConnsJSON`, `parseConntrack` (live flows, proxied/direct).
- **flock_unix.go** — real blocking cross-process flock; `lockPath=/var/lock/xrayctl.lock`.
- **geodata.go** — `geoAssetPresent`, `geoStrip`, `GeodataStatus/Download/Remove` (strip geo matchers when dat absent).
- **migrate.go** — `Migrate`, `CurrentSchemaVersion=1`, `uciRunner` (UCI schema migration; refuses newer).
- **migrate.go** — `Migrate`, `uciRunner`, and on the v0.1 branch `CurrentSchemaVersion = 1` with a single step `{0 -> 1}`. **That 1 is v0.1's number and nothing else's.** v0.2's `shater/model/migrate.go` is at `CurrentSchemaVersion = 2` with `{0 -> 1, 1 -> 2}` — `migrate1to2` is the one that removed `dst_domain`/`dst_ip` from `config rule` (see the schema subsection below, which is the live document). Both versions refuse a config NEWER than the build; in v0.2 that refusal also covers every config WRITE (`model.ErrSchemaTooNew`), so a downgraded package cannot quietly rewrite a newer config into the older form.
- **nodeops.go** — `NodeSetEnabled/NodeDelete/NodeAssignGroup/NodeQR`.
- **observatory.go** — read xray live state via gRPC API inbound (127.0.0.1:10853). (rewrite → lx command server)
- **preset.go** — `presetRules`, `presetDef`, `builtinPresets` (curated rule packs → synthetic Rules).
@@ -251,6 +258,20 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
> v0.2: "restart engine only on change" → config-hash gate + Close+New box (no reload).
### uci.go — `/etc/config/shater` schema
> This subsection alone describes the **CURRENT v0.2 schema**, not v0.1 — the
> shipped `/etc/config/shater` points its reader here by name, so it is kept in
> step with `shater/model/uci.go` (read) and `render.go` (write), which are the
> only two places a section type or option name exists. Everything else in PART A
> is the v0.1 survey the port was planned from and is deliberately frozen.
>
> An option not listed below is not "undocumented" — it is IGNORED: the parser's
> type switch drops an unknown section type whole, and an unknown option inside a
> known section is never read. That is deliberate (`TestUnknownSectionAndOptionIgnored`
> pins that such a config still parses — the daemon has to come up on whatever it
> finds), and it is also why a dead knob here is SILENT: setting one changes the
> file and nothing else, with no error anywhere to say so.
- `config globals` — the full option set, with the value used when the option is
ABSENT (the `model.DefaultGlobals` seed). Booleans are always written back as
`'1'`/`'0'` by `render.go`, so an explicit value never decays into the seed:
@@ -273,10 +294,12 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
| `block_doh` | `0` | NXDOMAIN the known public DoH hostnames + the Firefox canary and reject `:443` to their IPs, so clients fall back to `:53` (which the engine catches) |
| `group_health` | `1` | OUR background group probing (the observatory). Does not touch sing-box's own urltest inside a group |
| `untunnelable` | `block` | policy for what TPROXY cannot carry (ICMP/IGMP/ESP/AH/GRE/SCTP): `block` \| `icmp` (echo out, rest dropped) \| `direct` (all out, bypassing the tunnel) |
| `l3_tunnel` | **`1`** | Opens the synthetic `l3-in` TUN so LAN ICMP is routed by the engine instead of dropped/forged; nft marks LAN `icmp`/`ipv6-icmp` with `fwmark_base+0x80` and a scoped `ip rule` sends it to table `table_base+8`. ON by default since the flip (`model/l3tunnel_default_test.go`): with it off a LAN ping is decided by `untunnelable` alone, whose every rung either drops the echo or lets it out of the WAN with the client's real address. An ABSENT option therefore comes back ON; only an explicit `0` closes it, and that opt-out survives the write→read round-trip. Set through UCI — no panel control writes it (the Networks page reads it to explain what `untunnelable` still decides). See D25 and `ARCHITECTURE.md` §3a |
| `untunnelable_egress` | unset | **opt-in**, UCI-only. Names a `config egress`; everything the L3 block did not claim (ESP/AH, GRE, IGMP, SCTP, and ICMP when `l3_tunnel=0`) is stamped with that egress's OWN mark and routed out its device by the kernel — no new mark, no new table, engine not in the path. Empty ⇒ `untunnelable` above stays in sole charge (D26) |
| `geo_provider` | unset = auto | `sagernet` \| `loyalsoldier` \| `metacubex` \| `custom`; auto = country codes from SagerNet, everything else from Loyalsoldier |
| `geosite_url` / `geoip_url` | unset | `{category}` templates, honoured only when `geo_provider=custom` |
| `geosite_index_url` / `geoip_index_url` | unset | git-trees URLs used to SUGGEST categories in the panel; empty = no suggestions |
| `stats_backend` | `memory` | `off` (no aggregation at all) \| `memory` (RAM, lost on restart) \| `sqlite` (aggregates in RAM + query/connection log on disk) |
| `stats_backend` | `memory` | `off` (no aggregation at all) \| `memory` (RAM, lost on restart) \| `sqlite` (aggregates in RAM + query/connection log on disk). The value NAME is historical: the on-disk store is **bbolt**, not SQLite, since the migration — a leftover sqlite-era `stats.db` is detected by its file magic and replaced (`shater/stats/boltring.go`) |
| `stats_ring_size` / `stats_timeline_minutes` / `stats_max_domains` | `200` / `60` / `5000` | live-log length, sparkline minutes, domain-map cap. **`0` = UNLIMITED** (grows with traffic), which is why these three are always emitted |
| `stats_disk_limit_mb` | `64` | on-disk cap of `stats.db`; only meaningful for `stats_backend=sqlite`; `0` = unlimited |
| `stats_retention_disabled` | `0` | master switch that turns OFF all trimming/pruning — every aggregate then grows unbounded |
@@ -285,19 +308,66 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
Deleted options still parse (unknown keys are ignored) and drain out on the next
render: `dns_mode` (D17 — fake-IP is a resolver TYPE), `sweep_interval` (D19).
- `config inbound`: name, enabled, type, network, tproxy_port(12345), listen, port, auth, user, pass, target_addr, target_port, target_network, tcp, udp, sniff.
- `config subscription`: name, enabled, url, update_interval, fetch_via(direct|proxy), ua, hwid, device_os, ver_os, device_model, list header, format, list include/exclude/filter_proto/filter_country, dedup, expire_alert_days.
- `config node`: name, enabled, uri, mux, mux_concurrency, xudp_concurrency, xudp_udp443, sockopt_mark, tcp_fast_open, tcp_keepalive_idle.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, probe_url, probe_interval.
- `config chain`: name, list hop. `config egress`: name, type, interface, target.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url|geosite|geoip), url, path, format, update_interval, list category, list entry.
- `config inbound`: name, enabled, type, network, tproxy_port(12345), listen, port, auth, user, pass, target_addr, target_port, target_network, tcp, udp.
**No `sniff`.** Since sing-box 1.11 sniffing is a leading route ACTION rule with no
inbound matcher, so every inbound is sniffed always; the flag was read by nothing but
its own UCI round-trip. Re-adding it would be a regression, not a restored feature —
the hijack-dns rule matches the SNIFFED `dns` protocol, so a per-inbound toggle is a
DNS-leak switch wearing a performance label (`model.go`, `Inbound`).
- `config subscription`: name, enabled, url, update_interval, fetch_via(`direct`|`proxy`), fetch_detour, ua, hwid, device_os, ver_os, device_model, list header, format, list include/exclude/filter_proto/filter_country, dedup, expire_alert_days — plus the persisted `Subscription-Userinfo` state the daemon writes back itself: user_upload, user_download, user_total, user_expire, userinfo_at (absent = `0` = "not reported", on both the read and the write side).
`fetch_detour` is consulted only when `fetch_via=proxy`; it names the outbound the fetch dials through — `group:X` \| `node:X` \| `egress:X` \| `chain:X` \| `direct`, empty = direct (`apply.HTTPClient`/`resolveVia`).
- `config node`: name, enabled, uri, mux, mux_concurrency, sockopt_mark, tcp_fast_open, tcp_keepalive_idle, egress — the last binds THIS node's own upstream to a `config egress` (multi-WAN), so its exit connection leaves over the chosen device.
**No `xudp_concurrency`/`xudp_udp443`**: xudp was an xray packet-encoding knob with no sing-box counterpart, and the fields went with the generator rewrite.
`from_sub` / `fingerprint` / `stale` are READ but never written. Subscription nodes live in per-subscription cache files (`model/subcache.go`); the options survive only so an old config's cached nodes are imported once on first read, after which they drain out of UCI.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, egress — the last binds EVERY member's dialer to that egress (a node's own `egress` is more specific and wins).
**No `probe_url`/`probe_interval`**: the per-group overrides were deleted; probing is configured once, in `globals` (D20). Old configs carrying them still parse — the options are ignored and drain out on the next render.
- `config chain`: name, list hop (`group:<n>` \| `node:<n>`, L1..Ln, Ln = exit).
- `config egress`: name, type, interface, dpi.
type is `interface` \| `direct`; `tunnel` is an accepted ALIAS of `interface` and an empty value means `direct` — both folded to the canonical spelling once, at the config boundary (`Model.NormalizeEgressTypes`, called by `ReadUCI`), so the engine half and the data-plane half cannot disagree about a type name. An unrecognised type stays unrecognised (reported by `ValidateEgresses`, every binding to it fail-closed). A type this product REMOVED is a third case with the same fail-closed behaviour and a different sentence — `model.RetiredEgressTypes` is the closed table both `ValidateEgresses` and the generator read it from, so an operator whose config was correct for an older build is told what happened rather than that their value is a typo (D29).
`interface` names the device and is meaningful for the `interface` type only; `dpi` is the native DPI-bypass preset — `off`|`fragment`|`record`|`spoof` (D13).
**No `port`**: no surviving egress kind dials anything, so the option is not parsed and drains out on the next render. It belonged to the removed SOCKS-hop kind (D29).
**No `target`**: v0.1's `proxy`/`block` egress kinds are gone — where traffic goes is a rule's `target`, what device it leaves by is an egress.
- `config ruleset`: name, type(`domain`|`ipcidr`, default `domain`), source(`inline`|`file`|`url`|`geosite`|`geoip`, default `inline`), url, path, format, update_interval, list category, list entry.
`format` names the ENGINE rule-set format and has exactly two real values, `binary` (a compiled `.srs`) and `source` (a sing-box rule-set `.json`); empty — and the v0.1 leftover `plain`, and `auto` — mean "infer from the file name", which is sing-box's own behaviour. Ignored for inline/geosite/geoip.
`list category` is the canonical spelling (one chip per geosite category or geoip country code, each materialised as its own remote `.srs`); a legacy single `option category` is still accepted on read and re-emitted as a list.
- `config rule`: name, enabled, order, list src, list dst_ruleset, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end, sched_utc_offset.
v0.1 carried `dst_domain`/`dst_ip` on the rule itself; **schema v2 removed both** — a
destination is a `config ruleset` and nothing else. `shaterd migrate` folds each legacy
list into a generated `rule-<name>` (and `rule-<name>-ip`) inline ruleset; see
`DECISIONS.md` D21 for the entry-by-entry conversion table.
- `config preset`: name, enabled, order, target. `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config resolver`: name, type, address, detour, pool. `config dns_rule`: order, list match_domain/match_src, resolver.
- `config profile`: name, enabled, priority, list match_iface, list enable_rule, list disable_rule, endpoint_resolver.
The condition is `match_iface` and nothing else: the active default-route device must be in that set (`cmd/shaterd/profilewatch.go`). The overrides are the two rule lists plus an optional per-profile `endpoint_resolver`, which overrides `globals.endpoint_resolver` while the profile is active (`generate/dns.go`; the point is keying the bootstrap resolver to the WAN in use).
**No `probe_url`/`probe_mode`**, and they were deleted rather than documented: they promised "activate while a probe succeeds/fails", nothing ever probed, AND the selector treated a profile carrying a `probe_url` as having an unsatisfiable condition and skipped it — so adding a probe to a working profile silently switched that profile off.
**No `sched_*`, no `default_target`/`default_egress`** either — a profile enables and disables named rules; it does not carry a routing default of its own.
- `config resolver`: name, type, address, detour, pool.
type is `doh` (synonym `https`) \| `dot` (synonym `tls`) \| `plain` (synonyms `udp` and the empty default) \| `tcp` \| `local` \| `fakeip`. Anything else is NOT built — the resolver simply does not exist, and `generate` says so by name.
Each type consumes a different subset, and the rest is thrown away (`generate.warnIgnoredResolverFields` warns per field rather than dropping it silently): `doh`/`dot`/`plain`/`tcp` use address + detour, ignore pool; `local` uses detour only (it reads the router's `/etc/resolv.conf`); `fakeip` uses pool only (default `198.18.0.0/15`) — it mints answers locally, so it has nothing to dial and ignores both address and detour.
`detour` is per-SERVER, not global: every sing-box DNS server carries its own.
- `config dns_rule`: order, list match_domain, list match_src, resolver. It has no `name` — a DNS rule is identified by its order and selectors.
- `config blocklist`: name, enabled, source(`inline`|`file`|`url`|`geosite`), url, path, list category, list entry, response, update_interval.
`response` has exactly two values: `nxdomain` (default — Rcode 3, empty answer) and `zero` (NOERROR + `A 0.0.0.0`, and `AAAA ::` when `globals.ipv6`). There is no third; neither is sing-box's `action:reject`, which answers REFUSED.
- `config allowlist`: the same minus `response` (an allowlist has no verdict to render); it overrides every blocklist, being emitted at higher priority.
- `config device`: name, mac, ip, enabled, list block, list allow — where `block`/`allow` are DOMAINS, not targets: a device entry is per-device DNS filtering (block regardless of the global `dns_filter` switch, allow overriding every blocklist). Per-device ROUTING is not a device concern — it is an ordinary `config rule` whose `src` names the device.
- `config alert`: name, enabled, type(`telegram`|`webhook`), token + chat_id (telegram), url (webhook), list event, via, fallback.
- **`config preset` is NOT a section type, and there is no preset subsystem in
v0.2.** `ParseUCIExport`'s type switch has no `preset` branch, so such a section
is parsed by nothing, reaches no part of the model, and setting `enabled=1` on
one changes the file and nothing about the router.
It used to appear anyway: `30_shater-core` seeded three (`block_ads`,
`ru_bypass`, `private`) "so the LuCI Rules page renders their toggles", and
v0.2's LuCI app is a launcher with no Rules page. Worse than inert — the panel's
first save wiped them, because `writeUCIWith` replaces the whole package
(`uci delete shater` + `uci import`), so the placeholder deleted itself and read
as a breakage. **The seeding is gone, and the script now DELETES any `preset`
section an older release left behind** (safe by construction: the type is read
by no consumer, so there is no setting to lose).
This is not a feature waiting to be re-enabled. v0.1's packs were xray
`geosite:`/`geoip:` matcher lists materialised into synthetic rules
(`xrayctl/preset.go` on the v0.1 branch); under schema v2 a rule's destination
IS a `config ruleset`, so the same pack is an ordinary ruleset + rule — which
the panel's Routing page builds today, geosite/geoip sources included. Anything
richer needs a section type the parser knows, and that has to land in
`shater/model` first.
### Subscriptions & HAPP fetch
Schemes: `vless:// vmess:// trojan:// ss:// wireguard:// wg://`. Body formats (`DetectSubFormat`): clash-YAML, xray-JSON, singbox-JSON, base64/plain link list. All converge to URIs re-parsed by `ParseShareLink`. HAPP fetch: UA default `Happ/3.13.0`; headers `x-hwid` (auto UUIDv4/sub), `x-device-os`, `x-ver-os`, `x-device-model`, custom. `fetch_via=proxy` dials local socks. Quota/expiry from `Subscription-Userinfo` (`upload;download;total;expire`). Reconcile by `Fingerprint` (sha256 of proto|addr|port|id|net|sec|sni|path) → new/keep/stale (drop after 3 stale refreshes).
@@ -305,10 +375,14 @@ Schemes: `vless:// vmess:// trojan:// ss:// wireguard:// wg://`. Body formats (`
---
## PART B — v0.1 packaging (`shater-core/`, branch `v0.1`)
Pure scripts+config, `PKGARCH:=all`. v0.1 DEPENDS: `+xrayctl +xray-core +dnsmasq-full +kmod-nft-tproxy +kmod-nft-socket +ip-full`. → **v0.2 deps: `+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full`** (engine does DNS in-process, so dnsmasq-full may be droppable — confirm the :53 listener is our engine). `/etc/config/shater` is a conffile.
Pure scripts+config, `PKGARCH:=all`. v0.1 DEPENDS: `+xrayctl +xray-core +dnsmasq-full +kmod-nft-tproxy +kmod-nft-socket +ip-full`. → **v0.2 deps (authoritative: `openwrt/shater-core/Makefile`, which annotates each one): `+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle`** — `dnsmasq-full` is gone (the engine owns the `:53` hijack listener); `kmod-tun` is `/dev/net/tun` for the L3 ingress, `nftables-json` is the `nft -j` output `netplane/stats.go` parses, `ca-bundle` is the cert store a `CGO_ENABLED=0` binary has no host fallback for. `/etc/config/shater` is a conffile.
- **init.d/shater** (procd, START=99/STOP=10): v0.1 supervised `xray run -c /etc/xray/run.json`; → v0.2 supervises `shaterd`. `respawn 3600 5 0` (infinite). **No `procd_set_param file` watch** (would bounce tunnel on commit). Inert unless `globals.enabled=1`. `ACTIVE_FLAG=/var/run/shater.active` gates hotplug/cron. `stop` clears flag + tears down nft table + reserved routing tables. `reload_service`→start/stop. trigger `procd_add_reload_trigger "shater"`.
- **init.d/shater-cron** (START=96): supervised `loop`; per-item due-check, runs sub/ruleset update + reconcile + schedule due; watchdog: engine dead 5 ticks ⇒ kill_switch=open stops stack (fail-open), closed logs crit.
- **uci-defaults/30_shater-core**: seed `rt_tables` (8192 shater), enable both inits, seed preset packs (disabled), run migrate, apply sysctl.
- **init.d/shater-armor** (START=21/STOP=89, v0.2-only — no v0.1 counterpart): the fail-closed plane BEFORE the daemon exists. `/etc/init.d/shater` is START=99, so from netifd's `ifup` until the daemon's first apply the router forwarded LAN→WAN in the clear. The daemon persists its holding plane to `/etc/shater/boot.nft` on every apply; this loads it after fw4 (19) and netifd (20), `nft -c`-validated. Four state checks refuse to arm (no/empty/invalid file, missing `shaterd`, no `S??shater` rc-link, readable UCI saying `enabled≠1`) — asked ON THE WAY UP, deliberately not recorded on the way down. Hooks `forward` only, so SSH/LuCI/panel stay reachable. `stop()` is a NO-OP. Operator-facing writeup: `INSTALL.md` §4.
- **uci-defaults/30_shater-core**: seed `rt_tables` (8192 shater), `mkdir /etc/shater`, DELETE any leftover `config preset` section (a type nothing parses — see the schema note above; the seeding of three of them is gone), seed the `shater_l3` fw4 zone + a `<zone>→shater_l3` forwarding for every zone (named sections, `list device 'shater-l3*'`) and migrate a legacy exact-name entry to the wildcard, run `shaterd migrate` (whose result is **classified and reported**, not discarded — see below), apply sysctl, then a DETACHED bring-up (enable+restart `shater`/`shater-cron`, enable `shater-armor`, conditional `firewall reload`) — detached because an inline init call inside an apk/opkg transaction deadlocks on procd's flock.
- **`shaterd migrate` reporting** (both call sites: `uci-defaults/30_shater-core` and `init.d/shater`'s `start_service`). The verb exits 1 for every failure, so the shell classifies the outcome itself, with a CLOSED positive list — `ok` / `downgrade` / `unreadable` / `failed` (`shater_migrate_class`, duplicated in the two scripts because the package installs no shell library they could share; `TestMigrateClassifiersAgree` fails if they ever diverge). `downgrade` is recognised by the substring `newer than this build`, which both `model.migrateWith`'s refusal and `model.ErrSchemaTooNew` contain — a contract pinned by `TestMigrateDowngradeSignatureIsAContract`. An unrecognised failure lands on `failed`, which says so and quotes the binary verbatim, rather than being reported as one of the causes we can name.
Failures reach the operator on **two channels that are not syslog**, because `globals.log_syslog=0` is a deliberate setting about the syslog stream and not a request to be left uninformed: the script's own **stderr** (the operator's terminal on a hand-typed `restart`; the package manager's output inside `apk add`/`opkg install`), and **`/etc/shater/migrate-failed`** on flash, written on failure and REMOVED on the first success — its absence is the all-clear. syslog gets the same line too when `log_syslog` allows it. A migration that SUCCEEDED is routine and stays on the syslog channel only.
`30_shater-core` still **exits 0** after a failed migration, deliberately: a uci-defaults script that exits non-zero is kept and re-run at every boot, and this one re-runs a detached enable+restart of `shater`/`shater-cron` plus a firewall reload — so one recoverable failure would become permanent boot-time churn, to carry a status nothing reads. The retry that matters already exists: `init.d/shater` runs the migration on every start.
- **hotplug.d/iface/99-shater**: ifup/ifdown → debounced (2s) `reconcile` (netifd wipes ip rules on reload). Guarded by enabled + ACTIVE_FLAG.
- **sysctl.d/99-shater.conf**: `ip_forward=1`, `rp_filter=0` (all+default), `lo.route_localnet=1`, `lo.accept_local=1`, `all.src_valid_mark=1`, `ipv6.all.forwarding=1`.
+16 -20
View File
@@ -49,25 +49,21 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
the gate: fail-closed forward drop (4f618140), engine apply-swap close-first
fallback (9b6b9406), DNS hijack-dns per D14 (86194ce6).
## Phase 2b — DPI-bypass egress = ByeDPI (D13)
- The one external desync tool is **ByeDPI (ciadpi)** — chosen over zapret because
it *is* a SOCKS egress (fits shater's "routing picks the egress" model with zero
packet-plane conflict); zapret is explicitly rejected (see D13).
- Add egress `dpi` value `byedpi`: a supervised local `ciadpi` SOCKS5 instance +
a `socks` outbound pointed at it. New `openwrt/` procd package + musl-static
cross-build of ciadpi (~100 KB); `model` reserves the egress kind, `generate`
wires the `socks` outbound.
- QUIC gap closed by routing (drop `udp/443` for desync-domains → TCP+TLS
fallback), not by adopting a packet plane.
- The **`byedpi` package** (`openwrt/byedpi/`, SEPARATE & optional) provides the
`ciadpi` process behind a `type='byedpi'` egress: it cross-compiles ciadpi via
the SDK toolchain and ships a procd init that supervises one `ciadpi` SOCKS5
desync instance per enabled `config instance` in `/etc/config/byedpi`
(127.0.0.1:`<port>`). Install it only when you want a byedpi egress; a
`type='byedpi'` egress with no matching `ciadpi` listener simply has nothing to
dial. `shater-core` does NOT depend on it (opt-in).
- **Gate:** a DPI-blocked domain (that plain `fragment` can't crack) loads via the
`byedpi` egress on the VM, direct (no tunnel), kill-switch still honest.
## Phase 2b — DPI-bypass egress ✅ DONE, then REVERSED (D29, 2026-07-27)
- Shipped as **ByeDPI (ciadpi)**: a supervised local SOCKS5 desync process in its
own optional `openwrt/byedpi/` package, reached through a `type='byedpi'`
egress. Chosen over zapret because it *is* an egress and needed no second
packet plane (D13); zapret stays rejected.
- **Removed in full on 2026-07-27** — package, egress kind, readiness endpoint and
panel plate. It was adopted because the engine's own `tls_fragment` /
`tls_record_fragment` presets did not get through; the cause was a defect in
our fragmentation (the split always landed inside the first label of the name),
not a limit of the method. With that fixed the built-in presets carry this, and
the external process is weight without a job. Full argument: **D29**.
- What survives from this phase: the native `dpi` presets `fragment` / `record` /
`spoof` on a `direct` or `interface` egress, which is what the feature is now.
- A config still naming the removed kind is fail-closed and told so by name — see
`model.RetiredEgressTypes`.
## Phase 3 — Admin panel MVP + thin LuCI launcher ✅ DONE (2026-07-15)
- `panel/`: embedded web server on its own port + session store; token-mint ubus
@@ -86,7 +82,7 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
well-known lists (StevenBlack/OISD/AdGuard).
- **Gate:** ad/tracker domains blocked network-wide; big list loads fast; RAM sane.
## Phase 5 — Statistics (per-domain / client / device)
## Phase 5 — Statistics (per-domain / client / device) ✅ DONE
- Stats aggregator consuming the engine's DNS/routing/stats observability + nft
counters: top domains, allowed vs blocked, per-device breakdown, timelines,
per-node/per-rule traffic, live query log with one-click block.
+7 -1
View File
@@ -46,7 +46,7 @@ require (
github.com/sagernet/sing v0.8.12-0.20260702081104-2ded2af32d3d
github.com/sagernet/sing-cloudflared v0.1.3-0.20260706062323-d9787e794aa3
github.com/sagernet/sing-mux v0.3.5
github.com/sagernet/sing-quic v0.6.2-0.20260525051024-9467ede27fb7
github.com/sagernet/sing-quic v0.6.4-0.20260709034545-e23afe1172dc
github.com/sagernet/sing-shadowsocks v0.2.8
github.com/sagernet/sing-shadowsocks2 v0.2.1
github.com/sagernet/sing-shadowtls v0.2.1
@@ -104,14 +104,20 @@ require (
github.com/google/btree v1.1.3 // indirect
github.com/google/go-cmp v0.7.0 // indirect
github.com/google/go-querystring v1.1.0 // indirect
github.com/google/gopacket v1.1.19 // indirect
github.com/google/nftables v0.2.1-0.20240414091927-5e242ec57806 // indirect
github.com/google/uuid v1.6.0 // indirect
github.com/hashicorp/yamux v0.1.2 // indirect
github.com/hdevalence/ed25519consensus v0.2.0 // indirect
github.com/huin/goupnp v1.2.0 // indirect
github.com/inconshreveable/mousetrap v1.1.0 // indirect
github.com/jackpal/go-nat-pmp v1.0.2 // indirect
github.com/klauspost/compress v1.18.0 // indirect
github.com/klauspost/cpuid/v2 v2.3.0 // indirect
github.com/koron/go-ssdp v0.0.4 // indirect
github.com/kr/fs v0.1.0 // indirect
github.com/libp2p/go-nat v1.0.1-0.20250821073202-01afc089f138 // indirect
github.com/libp2p/go-netroute v0.2.1 // indirect
github.com/mdlayher/socket v0.5.1 // indirect
github.com/mitchellh/go-ps v1.0.0 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
+26 -2
View File
@@ -101,6 +101,8 @@ github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/google/go-querystring v1.1.0 h1:AnCroh3fv4ZBgVIf1Iwtovgjaw/GiKJo8M8yD/fhyJ8=
github.com/google/go-querystring v1.1.0/go.mod h1:Kcdr2DB4koayq7X8pmAG4sNG59So17icRSOU623lUBU=
github.com/google/gopacket v1.1.19 h1:ves8RnFZPGiFnTS0uPQStjwru6uO6h+nlr9j6fL7kF8=
github.com/google/gopacket v1.1.19/go.mod h1:iJ8V8n6KS+z2U1A8pUwu8bW5SyEMkXJB8Yo/Vo+TKTo=
github.com/google/nftables v0.2.1-0.20240414091927-5e242ec57806 h1:wG8RYIyctLhdFk6Vl1yPGtSRtwGpVkWyZww1OCil2MI=
github.com/google/nftables v0.2.1-0.20240414091927-5e242ec57806/go.mod h1:Beg6V6zZ3oEn0JuiUQ4wqwuyqqzasOltcoXPtgLbFp4=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
@@ -109,10 +111,14 @@ github.com/hashicorp/yamux v0.1.2 h1:XtB8kyFOyHXYVFnwT5C3+Bdo8gArse7j2AQ0DA0Uey8
github.com/hashicorp/yamux v0.1.2/go.mod h1:C+zze2n6e/7wshOZep2A70/aQU6QBRWJO/G6FT1wIns=
github.com/hdevalence/ed25519consensus v0.2.0 h1:37ICyZqdyj0lAZ8P4D1d1id3HqbbG1N3iBb1Tb4rdcU=
github.com/hdevalence/ed25519consensus v0.2.0/go.mod h1:w3BHWjwJbFU29IRHL1Iqkw3sus+7FctEyM4RqDxYNzo=
github.com/huin/goupnp v1.2.0 h1:uOKW26NG1hsSSbXIZ1IR7XP9Gjd1U8pnLaCMgntmkmY=
github.com/huin/goupnp v1.2.0/go.mod h1:gnGPsThkYa7bFi/KWmEysQRf48l2dvR5bxr2OFckNX8=
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw=
github.com/insomniacslk/dhcp v0.0.0-20260220084031-5adc3eb26f91 h1:u9i04mGE3iliBh0EFuWaKsmcwrLacqGmq1G3XoaM7gY=
github.com/insomniacslk/dhcp v0.0.0-20260220084031-5adc3eb26f91/go.mod h1:qfvBmyDNp+/liLEYWRvqny/PEz9hGe2Dz833eXILSmo=
github.com/jackpal/go-nat-pmp v1.0.2 h1:KzKSgb7qkJvOUTqYl9/Hg/me3pWgBmERKrTGD7BdWus=
github.com/jackpal/go-nat-pmp v1.0.2/go.mod h1:QPH045xvCAeXUZOxsnwmrtiCoxIr9eob+4orBN1SBKc=
github.com/jessevdk/go-flags v1.4.0/go.mod h1:4FA24M0QyGHXBuZZK/XkWh8h0e1EYbRYJSGM75WSRxI=
github.com/jsimonetti/rtnetlink v1.4.0 h1:Z1BF0fRgcETPEa0Kt0MRk3yV5+kF1FWTni6KUFKrq2I=
github.com/jsimonetti/rtnetlink v1.4.0/go.mod h1:5W1jDvWdnthFJ7fxYX1GMK07BUpI4oskfOqvPteYS6E=
@@ -122,6 +128,8 @@ github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zt
github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ=
github.com/klauspost/cpuid/v2 v2.3.0 h1:S4CRMLnYUhGeDFDqkGriYKdfoFlDnMtqTiI/sFzhA9Y=
github.com/klauspost/cpuid/v2 v2.3.0/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
github.com/koron/go-ssdp v0.0.4 h1:1IDwrghSKYM7yLf7XCzbByg2sJ/JcNOZRXS2jczTwz0=
github.com/koron/go-ssdp v0.0.4/go.mod h1:oDXq+E5IL5q0U8uSBcoAXzTzInwy5lEgC91HoKtbmZk=
github.com/kr/fs v0.1.0 h1:Jskdu9ieNAYnjxsi0LbQp1ulIKZV1LAFgK1tWhpZgl8=
github.com/kr/fs v0.1.0/go.mod h1:FFnZGqtBN9Gxj7eW1uZ42v5BccTP0vu6NEaFoC2HwRg=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
@@ -138,6 +146,10 @@ github.com/libdns/cloudflare v0.2.2 h1:XWHv+C1dDcApqazlh08Q6pjytYLgR2a+Y3xrXFu0v
github.com/libdns/cloudflare v0.2.2/go.mod h1:w9uTmRCDlAoafAsTPnn2nJ0XHK/eaUMh86DUk8BWi60=
github.com/libdns/libdns v1.1.1 h1:wPrHrXILoSHKWJKGd0EiAVmiJbFShguILTg9leS/P/U=
github.com/libdns/libdns v1.1.1/go.mod h1:4Bj9+5CQiNMVGf87wjX4CY3HQJypUHRuLvlsfsZqLWQ=
github.com/libp2p/go-nat v1.0.1-0.20250821073202-01afc089f138 h1:YohuNPT/1k3VcThCQlBZ43PCPWPfMRS1zcxWBF2SLK8=
github.com/libp2p/go-nat v1.0.1-0.20250821073202-01afc089f138/go.mod h1:TXQg5tfSy+bUjnhT5728j5j/MBj7keIYqqZ1+8k/ui8=
github.com/libp2p/go-netroute v0.2.1 h1:V8kVrpD8GK0Riv15/7VN6RbUQ3URNZVosw7H2v9tksU=
github.com/libp2p/go-netroute v0.2.1/go.mod h1:hraioZr0fhBjG0ZRXJJ6Zj2IVEVNx6tDTFQfSmcq7mQ=
github.com/logrusorgru/aurora v2.0.3+incompatible h1:tOpm7WcpBTn4fjmVfgpQq0EfczGlG91VSDkswnjF5A8=
github.com/logrusorgru/aurora v2.0.3+incompatible/go.mod h1:7rIyQOR62GCctdiQpZ/zOJlFyk6y+94wXzv6RNZgaR4=
github.com/mdlayher/netlink v1.9.0 h1:G8+GLq2x3v4D4MVIqDdNUhTUC7TKiCy/6MDkmItfKco=
@@ -264,8 +276,8 @@ github.com/sagernet/sing-cloudflared v0.1.3-0.20260706062323-d9787e794aa3 h1:3y6
github.com/sagernet/sing-cloudflared v0.1.3-0.20260706062323-d9787e794aa3/go.mod h1:XEqEDYRCAYLaoPjZ1ifVWJg5iWAJHL2gOAXe/PM28Cg=
github.com/sagernet/sing-mux v0.3.5 h1:RHnhVEc+SFqkrK4xMygYjDwwLhzp2Bj3lztSukONfhI=
github.com/sagernet/sing-mux v0.3.5/go.mod h1:QvlKMyNBNrQoyX4x+gq028uPbLM2XeRpWtDsWBJbFSk=
github.com/sagernet/sing-quic v0.6.2-0.20260525051024-9467ede27fb7 h1:hFLPJ21uNZSbRnzhOKz4Zv0b4F93mpDorWyN93BeRcM=
github.com/sagernet/sing-quic v0.6.2-0.20260525051024-9467ede27fb7/go.mod h1:+oqD54aHel4ALKkp1hVXWCgLU/EjLojvm6AUzDfvj0I=
github.com/sagernet/sing-quic v0.6.4-0.20260709034545-e23afe1172dc h1:zdc0fj4JdAdgAmQIoh7ZF+B/wPTEF2X75lYDqTmvlaw=
github.com/sagernet/sing-quic v0.6.4-0.20260709034545-e23afe1172dc/go.mod h1:9k+dzGsWMttUGldBzq3dU792YHXzW6NgfbOGltnXq+0=
github.com/sagernet/sing-shadowsocks v0.2.8 h1:PURj5PRoAkqeHh2ZW205RWzN9E9RtKCVCzByXruQWfE=
github.com/sagernet/sing-shadowsocks v0.2.8/go.mod h1:lo7TWEMDcN5/h5B8S0ew+r78ZODn6SwVaFhvB6H+PTI=
github.com/sagernet/sing-shadowsocks2 v0.2.1 h1:dWV9OXCeFPuYGHb6IRqlSptVnSzOelnqqs2gQ2/Qioo=
@@ -360,6 +372,8 @@ go4.org/mem v0.0.0-20240501181205-ae6ca9944745 h1:Tl++JLUCe4sxGu8cTpDzRLd3tN7US4
go4.org/mem v0.0.0-20240501181205-ae6ca9944745/go.mod h1:reUoABIJ9ikfM5sgtSF3Wushcza7+WeD01VB9Lirh3g=
go4.org/netipx v0.0.0-20231129151722-fdeea329fbba h1:0b9z3AuHCjxk0x/opv64kcgZLBseWJUpBw5I82+2U4M=
go4.org/netipx v0.0.0-20231129151722-fdeea329fbba/go.mod h1:PLyyIXexvUFg3Owu6p/WfdlivPbZJsZdgWZlrGope/Y=
golang.org/x/crypto v0.0.0-20190308221718-c2843e01d9a2/go.mod h1:djNgcEr1/C05ACkg1iLfiJU5Ep61QUkGW8qpdssI0+w=
golang.org/x/crypto v0.0.0-20191011191535-87dc89f01550/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
golang.org/x/crypto v0.0.0-20210513164829-c07d793c2f9a/go.mod h1:P+XmwS30IXTQdn5tA2iutPOUgjI07+tq3H3K9MVA1s8=
golang.org/x/crypto v0.48.0 h1:/VRzVqiRSggnhY7gNRxPauEQ5Drw9haKdM0jqfcCFts=
golang.org/x/crypto v0.48.0/go.mod h1:r0kV5h3qnFPlQnBSrULhlsRfryS2pmewsg+XfMgkVos=
@@ -367,17 +381,24 @@ golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93 h1:fQsdNF2N+/YewlRZiricy4P1i
golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93/go.mod h1:EPRbTFwzwjXj9NpYyyrvenVh9Y+GFeEvMNh7Xuz7xgU=
golang.org/x/image v0.27.0 h1:C8gA4oWU/tKkdCfYT6T2u4faJu3MeNS5O8UPWlPF61w=
golang.org/x/image v0.27.0/go.mod h1:xbdrClrAUway1MUTEZDq9mz/UpRwYAkFFNUslZtcB+g=
golang.org/x/lint v0.0.0-20200302205851-738671d3881b/go.mod h1:3xt1FjdF8hUf6vQPIChWIBhFzV8gjjsPE/fR3IyQdNY=
golang.org/x/mod v0.1.1-0.20191105210325-c90efee705ee/go.mod h1:QqPTAvyqsEbceGzBzNggFXnrqF1CaUcvgkdR5Ot7KZg=
golang.org/x/mod v0.33.0 h1:tHFzIWbBifEmbwtGz65eaWyGiGZatSrT9prnU8DbVL8=
golang.org/x/mod v0.33.0/go.mod h1:swjeQEj+6r7fODbD2cqrnje9PnziFuw4bmLbBZFrQ5w=
golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
golang.org/x/net v0.0.0-20210226172049-e18ecbb05110/go.mod h1:m0MpNAwzfU5UDzcl9v0D8zg8gWTRqZa9RBIspLL5mdg=
golang.org/x/net v0.0.0-20210525063256-abc453219eb5/go.mod h1:9nx3DQGgdP8bBQD5qxJ1jj9UTztislL4KSBs9R2vV5Y=
golang.org/x/net v0.50.0 h1:ucWh9eiCGyDR3vtzso0WMQinm2Dnt8cFMuQa9K33J60=
golang.org/x/net v0.50.0/go.mod h1:UgoSli3F/pBgdJBHCTc+tp3gmrU4XswgGRgtnwWTfyM=
golang.org/x/oauth2 v0.34.0 h1:hqK/t4AKgbqWkdkcAeI8XLmbK+4m4G5YeQRrmiotGlw=
golang.org/x/oauth2 v0.34.0/go.mod h1:lzm5WQJQwKZ3nwavOZ3IS5Aulzxi68dUSgRHujetwEA=
golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20210220032951-036812b2e83c/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.19.0 h1:vV+1eWNmZ5geRlYjzm2adRgW2/mcpevXNg50YZtPCE4=
golang.org/x/sync v0.19.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI=
golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20200217220822-9197077df867/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20200728102440-3e129f6d46b1/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20201119102817-f84b799fce68/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
@@ -389,6 +410,7 @@ golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.40.0 h1:36e4zGLqU4yhjlmxEaagx2KuYbJq3EwY8K943ZsHcvg=
golang.org/x/term v0.40.0/go.mod h1:w2P8uVp06p2iyKKuvXIm7N/y0UCRt3UfJTfZ7oOpglM=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.34.0 h1:oL/Qq0Kdaqxa1KbNeMKwQq0reLCCaFtqu2eNuSeNHbk=
@@ -396,8 +418,10 @@ golang.org/x/text v0.34.0/go.mod h1:homfLqTYRFyVYemLBFl5GgL/DWEiH5wcsQ5gSh1yziA=
golang.org/x/time v0.11.0 h1:/bpjEDfN9tkoN/ryeYHnv5hcMlc8ncjMcM4XBk5NWV0=
golang.org/x/time v0.11.0/go.mod h1:CDIdPxbZBQxdj6cxyCIdrNogrJKMJ7pr37NYpMcMDSg=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.0.0-20200130002326-2f3ba24bd6e7/go.mod h1:TB2adYChydJhpapKDTa4BR/hXlZSLoq2Wpct/0txZ28=
golang.org/x/tools v0.42.0 h1:uNgphsn75Tdz5Ji2q36v/nsFSfR/9BRFvqhGBaJGd5k=
golang.org/x/tools v0.42.0/go.mod h1:Ma6lCIwGZvHK6XtgbswSoWroEkhugApmsXyrUmBhfr0=
golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20200804184101-5ec99f83aff1 h1:go1bK/D/BFZV2I8cIQd1NKEZ+0owSTG1fDTci4IqFcE=
golang.org/x/xerrors v0.0.0-20200804184101-5ec99f83aff1/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
-99
View File
@@ -1,99 +0,0 @@
#
# byedpi — ByeDPI (ciadpi), a tiny portable-C SOCKS5/HTTP desync proxy.
#
# This is the process behind a Shater egress of `type='byedpi'`: shaterd's
# `generate` emits a SOCKS5 outbound `egress-<name>` -> 127.0.0.1:<port>, and a
# `ciadpi` instance supervised by this package listens on that port, applies
# TCP/TLS desync to the connections passing through it, and goes DIRECT to the
# target (no tunnel). Kept as a SEPARATE, optional package: a byedpi egress is
# opt-in — install this only when you want the external desync engine.
#
# Compiled C (musl, per target) => NOT PKGARCH:=all. The OpenWrt SDK toolchain
# cross-compiles ciadpi via its own plain Makefile.
#
include $(TOPDIR)/rules.mk
PKG_NAME:=byedpi
# DELIBERATELY NOT auto-versioned from our git tag (unlike shaterd/shater-core/
# luci-app-shater, which take SHATER_PKG_VERSION/SHATER_PKG_RELEASE from
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
# actually installed. Stamping our tag on it would be both a lie and a
# regression: our tags are 0.2.x, and the version comparator (apk-tools 3,
# verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
PKG_VERSION:=0.17.3
PKG_RELEASE:=1
# Pinned upstream release tag v0.17.3 (commit
# 7efde1b1296eaaa187b70e951894dde17527489c). codeload emits a stable tarball
# per tag; PKG_HASH is the sha256 of that tarball (build fails on mismatch).
PKG_SOURCE:=$(PKG_NAME)-$(PKG_VERSION).tar.gz
PKG_SOURCE_URL:=https://codeload.github.com/hufrea/byedpi/tar.gz/refs/tags/v$(PKG_VERSION)?
PKG_HASH:=0a9cb8585554c68c3e2be88c33c9bf6f99f8e8c7f54b362285adab99e262566c
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=MIT
PKG_LICENSE_FILES:=LICENSE
include $(INCLUDE_DIR)/package.mk
define Package/byedpi
SECTION:=net
CATEGORY:=Network
TITLE:=ByeDPI (ciadpi) local SOCKS5/HTTP desync proxy
URL:=https://github.com/hufrea/byedpi
# Pure C against musl; every target has a C toolchain, so no arch-depends.
# No runtime library deps beyond libc (static-ish tiny binary).
DEPENDS:=
endef
define Package/byedpi/description
ByeDPI is a small local SOCKS5/HTTP proxy that applies TCP/TLS desynchronization
(split, disorder, fake packets, TLS-record splitting) to the connections passing
through it and then connects DIRECTLY to the destination — no upstream tunnel.
Its binary is `ciadpi`. In the Shater stack it is the process behind an egress of
`type='byedpi'`: shaterd routes selected traffic to a local SOCKS5 outbound
pointed at ciadpi's 127.0.0.1:<port>. Multi-instance, driven by /etc/config/byedpi.
endef
# ciadpi's upstream Makefile appends its own required flags with `CFLAGS +=`.
# A CFLAGS set on the make command line CLOBBERS that `+=` (GNU make: a
# command-line assignment overrides the makefile's append), so we must re-supply
# ciadpi's own needed flags (-I. -std=c99 and its warning set) alongside
# $(TARGET_CFLAGS). CPPFLAGS (-D_DEFAULT_SOURCE) is left untouched by not
# overriding it. The default target `all` builds the `ciadpi` binary; its link
# rule is `$(CC) -o ciadpi $(OBJ) $(LDFLAGS)`, so $(TARGET_LDFLAGS) reaches the
# link. Kernel headers (linux/netfilter_ipv4.h) come from the SDK sysroot.
define Build/Compile
+$(MAKE) -C $(PKG_BUILD_DIR) \
CC="$(TARGET_CC)" \
CFLAGS="$(TARGET_CFLAGS) -I. -std=c99 -Wall -Wno-unused -Wextra -Wno-unused-parameter" \
LDFLAGS="$(TARGET_LDFLAGS)" \
all
endef
define Package/byedpi/install
$(INSTALL_DIR) $(1)/usr/bin
$(INSTALL_BIN) $(PKG_BUILD_DIR)/ciadpi $(1)/usr/bin/ciadpi
$(INSTALL_DIR) $(1)/etc/init.d
$(INSTALL_BIN) ./files/etc/init.d/byedpi $(1)/etc/init.d/byedpi
$(INSTALL_DIR) $(1)/etc/config
$(INSTALL_CONF) ./files/etc/config/byedpi $(1)/etc/config/byedpi
$(INSTALL_DIR) $(1)/etc/uci-defaults
$(INSTALL_BIN) ./files/etc/uci-defaults/40_byedpi $(1)/etc/uci-defaults/40_byedpi
endef
# /etc/config/byedpi is user-editable desired state -> preserve on upgrade.
define Package/byedpi/conffiles
/etc/config/byedpi
endef
$(eval $(call BuildPackage,byedpi))
-33
View File
@@ -1,33 +0,0 @@
#
# ByeDPI (ciadpi) desync SOCKS proxies (/etc/config/byedpi).
#
# Each `config instance` is one supervised `ciadpi` process bound to
# 127.0.0.1:<port>. Point a Shater egress at it:
#
# config egress 'bd'
# option name 'bd'
# option type 'byedpi'
# option port '1080' # must match an enabled instance's `port`
#
# then a rule with `option target 'egress:bd'` routes selected traffic through
# ciadpi, which desyncs it and connects DIRECTLY to the target (no tunnel).
#
# This shipped default is INERT (enabled='0'): installing the package opens no
# listener. Set enabled='1' and apply to bring the proxy up.
#
# This file is a conffile — your edits survive package upgrades.
#
config instance 'default'
option enabled '0'
option port '1080'
# Desync preset (documented ByeDPI example — general RKN/YouTube-busting):
# --disorder 1 : split the first segment at offset 1 and send the two
# parts in REVERSE order (TCP desync; defeats naive
# stream reassembly in the DPI).
# --auto=torst : auto mode — if the connection is reset (TCP RST), retry
# with the desync params instead of failing.
# --tlsrec 1+s : re-frame the TLS record boundary at the SNI offset +1,
# so the ClientHello SNI is split across TLS records and
# SNI-based DPI can't match the hostname.
option args '--disorder 1 --auto=torst --tlsrec 1+s'
-79
View File
@@ -1,79 +0,0 @@
#!/bin/sh /etc/rc.common
# /etc/init.d/byedpi — procd supervisor for ByeDPI (ciadpi) desync SOCKS proxies.
#
# One supervised `ciadpi` process per ENABLED `config instance` in
# /etc/config/byedpi. Each instance is a local SOCKS5 desync proxy bound to
# 127.0.0.1:<port>; a Shater egress of type='byedpi' with the matching `port`
# routes traffic to it (shaterd emits a SOCKS5 outbound to that port). ciadpi
# desyncs the connection and goes DIRECT to the target — no tunnel.
#
# Design:
# * INERT by default: the shipped instance has enabled='0', and config_foreach
# starts nothing unless an instance is explicitly enabled. Installing this
# package can never, by itself, open a listener or affect connectivity.
# * FOREGROUND: ciadpi stays in the foreground unless -D/--daemon is given (we
# never pass it), so procd supervises the real process. respawn on crash.
# * Instances are named after their UCI section, so `reload` diffs per-section
# and restarts only the instances whose config actually changed.
# * busybox ash only — no bashisms.
USE_PROCD=1
START=90 # before shater (START=99): the SOCKS egress should be up
STOP=11 # after shater (STOP=10) — higher STOP runs LATER on shutdown,
# so the desync proxy outlives the data plane it serves.
PROG=/usr/bin/ciadpi
CONF=byedpi
# Validate one `instance` section. Datatypes per openwrt-uci:
# enabled : bool (default 0 — inert)
# port : port (default 1080 — matches shater's byedpi egress default)
# args : free-form desync flag string (passed verbatim to ciadpi)
validate_instance_section() {
uci_load_validate "$CONF" instance "$1" "$2" \
'enabled:bool:0' \
'port:port:1080' \
'args:string:'
}
start_instance() {
# $1 = section name, $2 = validation return code
local cfg="$1"
[ "$2" = 0 ] || { echo "byedpi: validation failed for '$cfg'"; return 1; }
[ "$enabled" -eq 1 ] || return 0
# Never claim to run without the binary (half-removed/failed upgrade must
# degrade to "off", not to a phantom respawn loop).
[ -x "$PROG" ] || { echo "byedpi: $PROG missing, skipping '$cfg'"; return 1; }
procd_open_instance "$cfg"
# Bind loopback only: this proxy is reachable solely by the local engine.
procd_set_param command "$PROG" -i 127.0.0.1 -p "$port"
# Desync flag list. Unquoted on purpose: word-split $args into separate argv
# tokens (e.g. "--disorder 1 --auto=torst --tlsrec 1+s" -> 5 arguments).
# shellcheck disable=SC2086
[ -n "$args" ] && procd_append_param command $args
procd_set_param respawn
# Restart this instance when its config changes (checksum-watched).
procd_set_param file /etc/config/$CONF
procd_set_param stdout 1
procd_set_param stderr 1
procd_close_instance
}
start_service() {
config_load "$CONF"
config_foreach validate_instance_section instance start_instance
}
reload_service() {
start_service
}
service_triggers() {
# Reload (not full restart) when /etc/config/byedpi changes via a
# config.change event (LuCI Save&Apply / `reload_config`).
procd_add_reload_trigger "$CONF"
procd_add_validation validate_instance_section
}
@@ -1,12 +0,0 @@
#!/bin/sh
# /etc/uci-defaults/40_byedpi
#
# Idempotent first-boot setup for the byedpi package. Runs once on first boot
# (and via the default postinst on a live opkg/apk install); must exit 0 so it
# is cleared and not retried. Enabling the init is safe on a fresh box: the init
# and the shipped config are INERT (the 'default' instance has enabled='0'), so
# nothing listens until an instance is explicitly enabled.
[ -x /etc/init.d/byedpi ] && /etc/init.d/byedpi enable
exit 0
@@ -6,11 +6,24 @@
'require ui';
// The admin panel (shaterd's own web server) listens on its own port. shaterd
// reports the configured port in status.panel_port (globals.panel_port, default
// reports the CONFIGURED port in status.panel_port (globals.panel_port, default
// 8088); the launcher builds the redirect from that, falling back to the default
// when status is unavailable. `panelPort` tracks the latest reported value and is
// refreshed from each status poll. If the panel is ever fronted by TLS, the http
// scheme below must follow.
//
// CONFIGURED IS NOT BOUND, and nothing on this page can turn one into the other:
//
// - shaterd starts the panel server in a goroutine and treats a bind error as
// log-and-continue ("panel server unavailable (daemon continues)",
// cmd/shaterd/main.go). After a restart that races the old listener, the
// daemon is healthy, `panel_port` still names the port, and nothing is
// listening on it.
// - SHATER_PANEL_ADDR can disable the panel outright (panelAddr()), while
// apply.Status still reports effectivePanelPort() — 8088 when unset.
//
// There is no "panel is listening" field to read, so this page must not imply
// one. The hint and the tooltip say the port is configured, not checked.
var DEFAULT_PANEL_PORT = 8088;
var panelPort = DEFAULT_PANEL_PORT;
@@ -26,15 +39,32 @@ var callMintToken = rpc.declare({
expect: { '': {} }
});
// led renders a small status dot: state is 'good' | 'warn' | 'bad'.
// LED palette. 'unknown' is an UNLIT socket — never amber and never green.
// Amber is this page's "degraded", and there is nothing to be degraded about
// when no reading has arrived; green on a missing reading is how the panel used
// to claim health it had not measured (see panel/src/planeState.ts, which says
// the same thing and is the wording this page is kept in step with).
var LED_COLORS = {
good: '#37b24d',
warn: '#f59f00',
bad: '#e03131',
unknown: '#6b6b6b'
};
// led renders a small status dot: state is 'good' | 'warn' | 'bad' | 'unknown'.
// The dot is decorative — every row states its condition in words beside it — so
// it is hidden from assistive tech rather than being the only carrier of meaning.
function led(state) {
var color = state === 'good' ? '#37b24d'
: state === 'warn' ? '#f59f00'
: '#e03131';
// Closed positive list. An unrecognised state resolves to UNKNOWN, never to
// green: an open default here is exactly how a state nobody thought about
// ends up painted healthy.
var color = Object.prototype.hasOwnProperty.call(LED_COLORS, state)
? LED_COLORS[state] : LED_COLORS.unknown;
var glow = (color === LED_COLORS.unknown) ? '' : ';box-shadow:0 0 5px ' + color;
return E('span', {
'aria-hidden': 'true',
'style': 'display:inline-block;width:.72em;height:.72em;border-radius:50%;' +
'margin-right:.6em;vertical-align:-.05em;background:' + color +
';box-shadow:0 0 5px ' + color
'margin-right:.6em;vertical-align:-.05em;background:' + color + glow
});
}
@@ -49,63 +79,439 @@ function row(state, label, value) {
]);
}
// statusRows maps the shaterd status object to LED rows. An empty object (the
// ubus call failed / daemon down) degrades every row to a "down" reading.
function statusRows(st) {
st = st || {};
var down = (st.running !== true);
// ---------------------------------------------------------------------------
// Pure state derivation — no DOM below this line until statusRows().
//
// statusReadout() maps a shaterd status object to a list of
// { state, label, value } descriptors. It is deliberately free of E()/DOM so it
// can be run against recorded fixtures offline; tests/status-readout.test.js
// does exactly that for the four cases this page has to tell apart — engine up,
// engine down with the daemon answering, no daemon at all, and a daemon that
// cannot read the configuration — plus the degenerate and old-daemon answers.
// The gate runs it as step [7/7].
// ---------------------------------------------------------------------------
// PLANES is the closed set of values a LIVE Applier.Status() can put in `plane`
// (shater/apply/apply.go: "full" | "hold" | "none"). It is the FALLBACK proof of
// daemon liveness for a shaterd that predates daemon_answered — see daemonState().
var PLANES = { full: true, hold: true, none: true };
// daemonState — is the shaterd PROCESS answering?
//
// 'up' — a live Applier produced this status.
// 'down' — proven not: `shaterd status` printed its OFFLINE STUB.
// 'unknown' — no usable answer, or an answer too old to say either way.
//
// THE FIELD, THEN THE FALLBACK.
//
// `daemon_answered` (cmd/shaterd/main.go, statusDaemonAnsweredKey) is the
// CONTRACT: true means a running daemon answered over the control socket and
// every other field is that daemon's own Applier.Status(); false means the object
// is the offline stub — the apply.Status zero value plus a UCI and kernel read —
// and is not a status report at all. It is a positive, closed, two-valued
// statement about where the object came from, which is exactly what this page
// needs and what it never had.
//
// The `plane` test below is what this page used BEFORE that field existed, and it
// is kept only for the non-atomic-update window: `apk upgrade` can leave a new
// luci-app-shater beside an old shaterd, and that shaterd emits no
// daemon_answered. It works because the stub leaves `plane` at "" while a live
// Status() always assigns one of the three words — a side effect, not a promise,
// which is precisely why it is now second and not first. If the two ever
// disagree, the explicit field wins: a stub that somehow carried a plane word
// must still read as "no daemon answered".
//
// `running` MUST NOT be used for this. It changed meaning on 2026-07-26
// (a8970b8ac): it used to be a hardcoded true, and is now the ENGINE's liveness
// (apply.go `Running: engineUp`). A daemon that is perfectly alive with a dead
// engine reports running=false — and this page used to answer that with a red
// "Daemon: not running", the advice "start the Shater service first", and a
// DISABLED button to the one place the config can be fixed. The holding plane
// keeps management reachable on purpose (shater/netplane/nft.go); LuCI was the
// only thing taking that guarantee away.
//
// Nor is "the ubus call returned" sufficient: the rpcd plugin shells out to
// `shaterd status`, which prints a parseable object on BOTH branches (the exit
// code is what differs, and command substitution in the plugin drops it).
//
// Everything else is unknown and is painted as unknown: {} from a failed ubus
// call, {"error":...} from the plugin (which is ALSO what a live-but-wedged
// daemon produces — cmdStatus prints nothing and exits 1 on a control-socket
// timeout, so "wedged" must not be reported as "dead"), and a status from a
// daemon predating both fields.
function daemonState(st) {
if (!st || typeof st !== 'object')
return 'unknown';
// Closed positive list on the contract field. Anything that is not exactly
// `true` or exactly `false` is not a verdict — it falls through rather than
// being coerced, because a truthy string is not an answer.
if (st.daemon_answered === true)
return 'up';
if (st.daemon_answered === false)
return 'down';
// Compatibility fallback: an old shaterd under a new LuCI.
if (typeof st.plane === 'string' && PLANES[st.plane] === true)
return 'up';
if (st.plane === '')
return 'down';
return 'unknown';
}
// configState — could the daemon READ the router's configuration when it took
// this status?
//
// 'ok' — config_readable=true: enabled, kill_switch and panel_port below
// are readings.
// 'failed' — config_readable=false: those three are ZERO VALUES AND MEAN
// NOTHING (apply.go Status.ConfigReadable). Rendering `enabled`
// false as "switched off" here is the documented defect: the read
// fails when /overlay is full or a `uci commit` was interrupted,
// which is exactly when the fail-closed plane has the LAN cut off —
// and telling the owner they switched it off themselves sends them
// to a settings page backed by the same unreadable file.
// 'unknown' — nobody said. Two ways to get here, and neither may be read as
// 'failed':
// * a daemon predating the field (non-atomic package update);
// * NO LIVE DAEMON AT ALL. The offline stub in cmdStatus reads
// UCI directly and never sets ConfigReadable, so it emits
// config_readable=false while its enabled/kill_switch/
// panel_port ARE genuine reads. Taking that at face value would
// put "the configuration could not be read" on screen for a
// perfectly readable configuration, and would throw away the
// only facts a dead-daemon status does carry. So the field is
// only consulted when a daemon actually answered.
function configState(st) {
if (daemonState(st) !== 'up')
return 'unknown';
if (st.config_readable === true)
return 'ok';
if (st.config_readable === false)
return 'failed';
return 'unknown';
}
// killSwitchSetting — the CONFIGURED kill-switch policy, or null when it cannot
// be known. It is one of the three config-sourced fields, so an unreadable
// configuration leaves it "" — and "" must never normalise to "closed", which is
// how a router nobody could read printed a green "fail-closed" row (the panel's
// killSwitchReadout carries the same note).
function killSwitchSetting(st) {
if (configState(st) === 'failed')
return null;
if (st.kill_switch === 'closed')
return 'closed';
if (st.kill_switch === 'open')
return 'open';
return null;
}
// panelTarget — where the launcher will point, and whether the port came from the
// router or from this file's built-in default. It never claims the panel answers
// there; see the DEFAULT_PANEL_PORT comment at the top.
function panelTarget(st) {
st = (st && typeof st === 'object') ? st : {};
// The config gate is belt-and-braces: an unreadable configuration already
// leaves panel_port at 0, which the > 0 test rejects. It is spelled out so the
// rule ("only meaningful with config_readable=true") is visible at the point
// of use rather than inferred from a zero value.
if (configState(st) !== 'failed' &&
typeof st.panel_port === 'number' && st.panel_port > 0)
return { port: st.panel_port, known: true };
return { port: DEFAULT_PANEL_PORT, known: false };
}
// engineState — is a sing-box instance actually started?
//
// The engine lives INSIDE the shaterd process, so a dead daemon is a dead engine
// and this page may say so without guessing. With the daemon up, `engine_running`
// is the self-documenting field and `running` carries the same fact by
// construction; either may prove a NEGATIVE, and a negative always wins. Neither
// asserting anything leaves 'unknown'.
function engineState(st) {
var d = daemonState(st);
if (d === 'down')
return 'down';
if (d === 'unknown')
return 'unknown';
if (st.running === false || st.engine_running === false)
return 'down';
if (st.running === true || st.engine_running === true)
return 'up';
return 'unknown';
}
function mk(state, label, value) {
return { state: state, label: label, value: value };
}
function statusReadout(st) {
st = (st && typeof st === 'object') ? st : {};
var dstate = daemonState(st);
var estate = engineState(st);
var cstate = configState(st);
var kswitch = killSwitchSetting(st);
// WHICH FIELDS SURVIVE A DEAD DAEMON. The offline stub's contract
// (cmd/shaterd/main.go statusDaemonAnsweredKey) names them: enabled, active,
// table, kill_switch and panel_port are read on the spot from UCI and the
// kernel and are real; running, engine_running, plane, traffic, hash, warnings
// and uptime "are placeholders, not measurements". Today the stub happens to
// leave all of them at their zero values, so reading them would look harmless
// — but "it happens to be zero" is the same side-effect reasoning that
// daemon_answered was added to replace. They are dropped on the stated
// contract instead, so a stub that ever grew a value cannot paint this page
// green.
var live = (dstate !== 'down');
var planeWord = live ? st.plane : '';
var traffic = (live && st.traffic && typeof st.traffic === 'object') ? st.traffic : {};
var rows = [];
// Daemon process itself.
rows.push(row(
down ? 'bad' : 'good',
_('Daemon (shaterd)'),
down ? _('not running') : _('running')
));
// --- The shaterd process itself. ------------------------------------------
// Its own liveness is not a field; it is whether a live daemon answered.
if (dstate === 'up')
rows.push(mk('good', _('Daemon (shaterd)'), _('responding')));
else if (dstate === 'down')
rows.push(mk('bad', _('Daemon (shaterd)'),
_('not responding — start the Shater service')));
else
rows.push(mk('unknown', _('Daemon (shaterd)'),
_('no usable answer — state unknown')));
// Desired state: globals.enabled in UCI.
rows.push(row(
st.enabled ? 'good' : 'warn',
_('Service enabled'),
st.enabled ? _('enabled') : _('inert (disabled)')
));
// --- The engine (sing-box) inside it. -------------------------------------
if (estate === 'up')
rows.push(mk('good', _('Engine (sing-box)'), _('running')));
else if (estate === 'down' && dstate === 'down')
rows.push(mk('bad', _('Engine (sing-box)'),
_('stopped — it runs inside shaterd, which is not answering')));
else if (estate === 'down')
rows.push(mk('bad', _('Engine (sing-box)'),
_('stopped — the daemon is up but no instance is running')));
else
rows.push(mk('unknown', _('Engine (sing-box)'), _('not reported')));
// Interception raised (ACTIVE_FLAG present after a successful enabled apply).
rows.push(row(
st.active ? 'good' : (st.enabled ? 'warn' : 'bad'),
_('Interception'),
st.active ? _('active') : _('inactive')
));
// --- Could the configuration be read at all? ------------------------------
// Placed ABOVE the three rows it qualifies, because it decides what they mean.
if (cstate === 'ok')
rows.push(mk('good', _('Configuration'), _('readable')));
else if (cstate === 'failed')
rows.push(mk('bad', _('Configuration'),
_('COULD NOT BE READ — the service/kill-switch/panel-port rows below say ' +
'"not known" because the daemon has no reading, NOT because anything is ' +
'switched off. If traffic is blocked that is the fail-closed plane; do not ' +
'turn anything off to fix it.') +
(typeof st.config_error === 'string' && st.config_error !== ''
? ' [' + st.config_error + ']' : '')));
else
rows.push(mk('unknown', _('Configuration'),
dstate === 'up'
? _('not reported by this daemon')
: _('not reported — no daemon answered')));
// Data plane: the `inet shater` nft table is loaded.
rows.push(row(
st.table ? 'good' : (st.enabled ? 'warn' : 'bad'),
_('Data plane'),
st.table ? _('nft table inet shater loaded') : _('not loaded')
));
// --- Desired state: globals.enabled in UCI. -------------------------------
// Read through cstate first: `enabled` is sourced from the configuration, so
// with config_readable=false its `false` is a zero value, not a choice.
if (cstate === 'failed')
rows.push(mk('unknown', _('Service enabled'),
_('not known — the configuration could not be read')));
else if (st.enabled === true)
rows.push(mk('good', _('Service enabled'), _('enabled')));
else if (st.enabled === false)
rows.push(mk('warn', _('Service enabled'), _('inert (disabled)')));
else
rows.push(mk('unknown', _('Service enabled'), _('not reported')));
// Kill-switch: fail-closed ("closed") is the safe posture; "open" leaks
// LAN→WAN if the engine goes down. Unknown (older daemon) degrades to warn.
var ks = st.kill_switch;
rows.push(row(
ks === 'closed' ? 'good' : 'warn',
_('Kill-switch'),
ks === 'closed' ? _('closed (fail-closed)')
: ks === 'open' ? _('open (leaky)')
: _('unknown')
));
// --- The ACTIVE_FLAG latch. -----------------------------------------------
// NOT a health signal, and this row must never read as one. apply.go states
// the contract: it is the "the service is meant to be running" latch that
// gates hotplug and cron; it is raised by a successful enabled apply and
// cleared only by teardown, so it STAYS UP while the engine is down and the
// fail-closed holding plane is blocking the LAN — deliberately, because
// clearing it would switch off the very cron reconcile that brings the engine
// back. This page used to render it as "Interception: active", in green, over
// a dead engine and a blocked LAN. The lamp now reports only whether the
// latch AGREES with globals.enabled.
//
// The latch itself is a filesystem fact and stays readable when the
// configuration does not; what an unreadable configuration takes away is the
// COMPARISON, since the lamp only reports whether the latch agrees with
// globals.enabled. So the words stay and the lamp goes out.
if (typeof st.active !== 'boolean')
rows.push(mk('unknown', _('Service latch'), _('not reported')));
else if (cstate === 'failed')
rows.push(mk('unknown', _('Service latch'), st.active
? _('raised — the service is meant to be running; whether that matches the ' +
'setting is not known, the configuration could not be read')
: _('cleared — the service is torn down; whether that matches the setting is ' +
'not known, the configuration could not be read')));
else if (st.active)
rows.push(mk(st.enabled === true ? 'good' : 'warn', _('Service latch'),
_('raised — the service is meant to be running')));
else
rows.push(mk(st.enabled === false ? 'good' : 'warn', _('Service latch'),
_('cleared — the service is torn down')));
// Running engine config hash ("" when the engine is not started).
rows.push(row(
st.hash ? 'good' : 'warn',
_('Config hash'),
st.hash ? st.hash : '—'
));
// --- What is loaded in the kernel right now. ------------------------------
// The row this page was missing. `plane` distinguishes the working ruleset
// from the FAIL-CLOSED HOLDING PLANE, which `table` cannot: `table` is a bare
// existence check, so a held LAN and a working one look identical through it.
switch (planeWord) {
case 'full':
// Deliberately mechanical wording. "full" means the table, the policy
// routing and the engine are all in place — it does NOT mean traffic is
// tunnelled. That claim belongs to the traffic verdict below.
rows.push(mk('good', _('Traffic plane'),
_('full — ruleset, routing and engine are all installed')));
break;
case 'hold':
rows.push(mk('bad', _('Traffic plane'),
_('hold — the engine is down and LAN→WAN forwarding is BLOCKED')));
break;
case 'none':
// The alarming wording is earned by a KNOWN fail-closed setting. With an
// unreadable configuration kswitch is null, and the row states the fact it
// has (nothing is installed) without the claim it does not.
rows.push(mk(kswitch === 'closed' ? 'bad' : 'warn', _('Traffic plane'),
kswitch === 'closed'
? _('none — nothing is installed; traffic reaches the WAN unprotected')
: _('none — no data plane is installed')));
break;
default:
rows.push(mk('unknown', _('Traffic plane'),
dstate === 'down'
? _('not reported — no daemon answered')
: _('not reported by this daemon')));
break;
}
// --- Where the traffic goes under the running config. ---------------------
// Separate from the plane on purpose: a router with one `default -> direct`
// rule has a fully installed plane and sends every packet out the plain WAN
// with its real address.
switch (traffic.verdict) {
case 'tunnel':
rows.push(mk('good', _('Traffic verdict'), _('tunnel — unmatched traffic is proxied')));
break;
case 'split':
rows.push(mk('warn', _('Traffic verdict'),
_('split — the default leaves directly; only matched rules are tunnelled')));
break;
case 'direct':
rows.push(mk('warn', _('Traffic verdict'),
_('direct — nothing is tunnelled; traffic leaves over the plain WAN')));
break;
case 'blocked':
rows.push(mk('warn', _('Traffic verdict'), _('blocked — unmatched traffic is dropped')));
break;
default:
rows.push(mk('unknown', _('Traffic verdict'), _('not reported')));
break;
}
// --- The nft table, as a bare presence check. -----------------------------
// Kept because the offline stub still reads it straight from the kernel, so
// it is the one plane fact available when no daemon answers. It says nothing
// about WHICH ruleset is loaded — that is the Traffic plane row.
//
// "Not loaded" is only the calm amber when the service is KNOWN to be switched
// off. With an unreadable configuration that is not known, and an absent
// firewall table with no explanation is the alarming side, not the calm one.
if (st.table === true)
rows.push(mk('good', _('nft table'), _('inet shater is loaded')));
else if (st.table === false)
rows.push(mk(cstate !== 'failed' && st.enabled === false ? 'warn' : 'bad',
_('nft table'), _('not loaded')));
else
rows.push(mk('unknown', _('nft table'), _('not reported')));
// --- Kill-switch: the configured policy, and whether it is in force. ------
// "closed" with no plane installed is a setting that is not in effect, which
// is worse news than "open" and must not share its amber lamp.
if (cstate === 'failed')
rows.push(mk('unknown', _('Kill-switch'),
_('not known — the configuration could not be read')));
else if (kswitch === 'closed' && planeWord === 'none')
rows.push(mk('bad', _('Kill-switch'),
_('closed, but NOT in effect — no data plane is installed')));
else if (kswitch === 'closed' && dstate === 'up')
rows.push(mk('good', _('Kill-switch'), _('closed (fail-closed)')));
else if (kswitch === 'closed')
rows.push(mk('warn', _('Kill-switch'),
_('configured closed; whether it is installed is not known')));
else if (kswitch === 'open')
rows.push(mk('warn', _('Kill-switch'), _('open (leaky)')));
else
rows.push(mk('unknown', _('Kill-switch'), _('not reported')));
// --- Running engine config hash ("" when the engine is not started). ------
if (live && typeof st.hash === 'string' && st.hash !== '')
rows.push(mk('good', _('Config hash'), st.hash));
else if (estate === 'down')
rows.push(mk('unknown', _('Config hash'), _('none — the engine is not started')));
else
rows.push(mk('unknown', _('Config hash'), _('not reported')));
// --- Where the "Open panel" button will point. ----------------------------
// The lamp is UNLIT even on a perfectly healthy router, and that is the point:
// the lamps on this page report a condition, and the condition an operator
// cares about here — "will the panel answer on that port" — is one nothing in
// this status measures. shaterd's panel server is started in a goroutine whose
// bind error is only logged, so a taken port leaves a healthy daemon reporting
// a port nothing is listening on; SHATER_PANEL_ADDR can switch the server off
// entirely and the port is still reported. A green lamp here would be a
// promise made out of a configuration value.
var target = panelTarget(st);
if (target.known)
rows.push(mk('unknown', _('Panel port'),
_('%PORT% — configured; nothing here reports whether the panel is listening on it')
.replace('%PORT%', String(target.port))));
else if (cstate === 'failed')
rows.push(mk('unknown', _('Panel port'),
_('not known — the configuration could not be read; the button will try the built-in default %PORT%')
.replace('%PORT%', String(target.port))));
else
rows.push(mk('unknown', _('Panel port'),
_('not reported — the button will try the built-in default %PORT%')
.replace('%PORT%', String(target.port))));
return rows;
}
// statusRows turns the readout into LED table rows.
function statusRows(st) {
return statusReadout(st).map(function(r) {
return row(r.state, r.label, r.value);
});
}
// panelTitle describes the button's target and what is known about it. It never
// promises the panel is up — only where the launcher will point.
function panelTitle(st) {
var dstate = daemonState(st);
var target = panelTarget(st);
// Second sentence, on every branch: the port is a configuration value that
// nothing on the wire confirms. A page that says "opens the panel" and lands
// on a connection refused has made a claim it had no field to support.
var port = target.known
? _('The port is the configured one; a status cannot say whether the panel is listening on it, so a connection error here means the port, not your token.')
: _('No panel port was reported, so this falls back to the built-in default and may well be the wrong port.');
if (dstate === 'up')
return _('Mint a session token and open the admin panel.') + ' ' + port;
if (dstate === 'down')
return _('shaterd is not answering, so this will probably fail — but the panel is served by the daemon, not by the engine, so it is worth trying: any failure is reported here.') + ' ' + port;
return _('The daemon state is not known. Try it — a failure is reported here rather than hidden.') + ' ' + port;
}
// panelHint is the grey line beside the button. Same rule as the tooltip: it
// states what the button DOES, and marks the port as configured rather than
// checked.
function panelHint(hostname, target) {
return (target.known
? _('opens http://%HOST%:%PORT%/ with a single-use session token — %PORT% is the configured port, not a checked one')
: _('opens http://%HOST%:%PORT%/ with a single-use session token — no port was reported, so %PORT% is this page\'s built-in default'))
.replace('%HOST%', hostname)
.replace(/%PORT%/g, String(target.port));
}
// handleOpenPanel mints a single-use token and hands it to the panel via the
// ARCHITECTURE §2 browser bridge: GET http://<router>:<port>/?t=<token>. The panel
// validates+consumes the token and drops a session cookie.
@@ -151,13 +557,21 @@ return view.extend({
handleSave: null,
handleReset: null,
// Exposed so the offline fixture harness (tests/status-readout.test.js) can
// exercise the state derivation without a browser, a router, or a DOM.
statusReadout: statusReadout,
daemonState: daemonState,
engineState: engineState,
configState: configState,
panelTarget: panelTarget,
panelTitle: panelTitle,
panelHint: panelHint,
load: function() {
return L.resolveDefault(callStatus(), {});
},
render: function(st) {
var self = this;
var table = E('table', { 'class': 'table' }, statusRows(st));
var openBtn = E('button', {
@@ -165,34 +579,37 @@ return view.extend({
'click': ui.createHandlerFn(this, handleOpenPanel)
}, [ _('Open panel') ]);
function hintText() {
return _('opens http://%s:%d/ with a single-use session token')
.format(window.location.hostname, panelPort);
}
var hint = E('span', {
'style': 'margin-left:1em;color:#888;font-size:90%'
}, hintText());
}, panelHint(window.location.hostname, panelTarget(st)));
// Reflect daemon reachability on the button up front, then keep the whole
// dashboard live. Also track the panel port reported in status so the
// launcher redirect and hint follow globals.panel_port.
// Track the panel port reported in status so the launcher redirect and the
// hint follow globals.panel_port.
//
// THE BUTTON IS NEVER DISABLED. It used to be locked whenever
// `running !== true`, which after a8970b8ac means "the engine is down" —
// precisely the situation the panel exists to get you out of, and one in
// which the daemon and its web server are still up and still minting
// tokens (cmd/shaterd/main.go starts the panel server independently of the
// engine). Locking it on a guess is the failure; a mint that fails already
// reports itself through ui.addNotification, which is the recoverable
// direction for an unknown state.
function reflect(state) {
state = state || {};
panelPort = state.panel_port || DEFAULT_PANEL_PORT;
hint.textContent = hintText();
var down = (state.running !== true);
openBtn.disabled = down;
openBtn.title = down
? _('shaterd is not running — start the Shater service first')
: _('Mint a session token and open the admin panel');
state = (state && typeof state === 'object') ? state : {};
// One source for the port: the same panelTarget() the "Panel port" row
// renders, so the row, the hint, the tooltip and the redirect cannot
// disagree about where the button goes.
var target = panelTarget(state);
panelPort = target.port;
hint.textContent = panelHint(window.location.hostname, target);
openBtn.title = panelTitle(state);
}
reflect(st || {});
reflect(st);
poll.add(function() {
return L.resolveDefault(callStatus(), {}).then(function(s) {
dom.content(table, statusRows(s));
reflect(s || {});
reflect(s);
});
}, 5);
@@ -209,7 +626,7 @@ return view.extend({
E('div', { 'class': 'cbi-section' }, [
E('h3', {}, _('Admin panel')),
E('p', { 'class': 'cbi-value-description' },
_('The rich admin panel is served by shaterd on its own port. LuCI mints a short-lived, single-use token for your browser — the panel has no separate login.')),
_('The rich admin panel is served by shaterd on its own port — by the daemon, not by the engine, so it stays reachable while the engine is down. LuCI mints a short-lived, single-use token for your browser; the panel has no separate login.')),
E('div', {}, [ openBtn, hint ])
])
]);
@@ -4,9 +4,41 @@
# Registers the ubus object "shater" (object name == this file's name) with two
# read-side methods the thin LuCI launcher calls over ubus:
#
# status -> passthrough of `shaterd status` ({running,enabled,active,table,hash})
# status -> passthrough of `shaterd status`
# mint_token -> passthrough of `shaterd mint-token` ({"token":"..."} | {"error":"..."})
#
# The status object is whatever apply.Status marshals (shater/apply/apply.go is the
# only definition; this script never parses or reshapes it), plus the one field
# `shaterd status` splices in itself. As of 2026-07-27 that is:
#
# daemon_answered, running, engine_running, enabled, active, table, plane,
# traffic, hash, kill_switch, panel_port, config_readable, config_error,
# can_rollback, warnings, started_unix, uptime_seconds
#
# Three of those are load-bearing for the caller and easy to misread:
#
# daemon_answered — WHERE THE OBJECT CAME FROM, and the only field that says so by
# contract (cmd/shaterd/main.go, statusDaemonAnsweredKey). true: a running
# daemon answered over the control socket. false: this is the OFFLINE STUB —
# the apply.Status zero value plus a UCI and kernel read — and the fields only
# a live daemon can know (running/engine_running/plane/traffic/hash/warnings/
# uptime) are placeholders. dashboard.js keys "daemon: down" off this.
# running / engine_running — the ENGINE's liveness, not the daemon's. `running`
# was a hardcoded true until a8970b8ac (2026-07-26) and is now `engineUp`, so
# a healthy daemon with a dead engine reports running=false. The daemon's own
# liveness is not a field of apply.Status at all.
# config_readable — whether the daemon could READ the configuration. When false,
# enabled/kill_switch/panel_port are zero values and mean NOTHING. Note the
# stub above never sets it, so its false is not a failed read either — a
# consumer must check daemon_answered first. dashboard.js does.
#
# EXIT CODE: `shaterd status` now exits 1 when no daemon answered, and this script
# deliberately ignores that — command substitution below keeps only stdout. The stub
# is still worth relaying (it carries the real UCI and nft-table readings, and it
# says what it is), and dropping it would blank the LuCI page instead of degrading
# it. The wedged case is the one where the exit code matters, and it reports itself
# by printing NOTHING: the case below then emits the error object.
#
# Why shell out to shaterd instead of talking to /var/run/shaterd.ctl directly:
# a reliable AF_UNIX client is NOT guaranteed on stock OpenWrt (busybox `nc` is
# usually built without `-U`; socat/ucode-socket aren't in the base image). shaterd
@@ -0,0 +1,521 @@
#!/usr/bin/env node
/*
* Offline harness for the dashboard's state derivation.
*
* Run: node openwrt/luci-app-shater/tests/status-readout.test.js
* (the gate runs it as step [7/7]; see scripts/run-tests.sh NONGO_TEST_RUNNERS)
*
* Why this exists: the LuCI page is the ONE screen an operator reaches when the
* engine is down and the fail-closed holding plane is blocking the LAN. What it
* says there is a claim about the router's behaviour, and until now nothing
* checked those claims. There is no browser and no router in this loop — the view
* exposes statusReadout/daemonState/engineState/configState/panelTarget/
* panelTitle/panelHint as plain functions, and this file feeds them recorded
* status objects.
*
* The fixtures are not invented. Each is what the wire actually carries:
*
* ENGINE_UP — apply.Status() from a live daemon with a started engine,
* marked daemon_answered:true by the CLI.
* ENGINE_DOWN — the same daemon with a dead engine; the holding plane is
* installed and the LAN is blocked.
* DAEMON_DOWN — the OFFLINE STUB `shaterd status` prints when the daemon is
* unreachable (cmd/shaterd/main.go cmdStatus): apply.Status zero
* value + a UCI and kernel read, marked daemon_answered:false.
* CONFIG_BAD — a LIVE daemon that could not read the configuration
* (config_readable:false): enabled/kill_switch/panel_port are
* zero values that mean NOTHING.
* NO_ANSWER — {} , what L.resolveDefault hands render() when the ubus call
* fails outright.
* PLUGIN_ERROR — {"error":...} from the rpcd plugin, which is ALSO what a
* live-but-wedged daemon produces.
* OLD_* — a shaterd predating daemon_answered/config_readable, under a
* new luci-app-shater. Packages do not update atomically, so
* this combination WILL exist in the field.
* LEGACY — a daemon predating plane/engine_running as well.
*
* Mutation check (each has been run; the named assertions in brackets fail):
* - delete the daemon_answered branches from daemonState() [H/*, C/*]
* - delete the `plane` fallback from daemonState() [OLD/*]
* - read config_readable without the daemon gate [C/config-*]
* - paint config_readable:false rows from their zero values [K/*]
* Reproduce by copying dashboard.js, breaking the copy, and pointing
* DASHBOARD_JS at it.
*/
'use strict';
var fs = require('fs');
var path = require('path');
// --- Load the view module with LuCI's globals stubbed. ----------------------
// The view file is a module body LuCI wraps in a function, so it ends in a
// top-level `return` and cannot be require()d. Wrapping it in new Function is the
// same thing LuCI's loader does. The 'require x' lines are bare string literals
// and evaluate to nothing.
// DASHBOARD_JS points the harness at a copy of the view. It exists so the
// mutation check is repeatable: copy dashboard.js, reintroduce the defect in the
// copy, run this file against it, and watch the named assertions fail. A test
// that cannot be shown to fail on the broken code is decoration.
var SRC = process.env.DASHBOARD_JS || path.join(__dirname, '..', 'htdocs',
'luci-static', 'resources', 'view', 'shater', 'dashboard.js');
function loadView() {
var src = fs.readFileSync(SRC, 'utf8');
var factory = new Function('view', 'dom', 'poll', 'rpc', 'ui', 'E', '_', 'L',
'window', src);
return factory(
{ extend: function(o) { return o; } }, // view
{ content: function() {} }, // dom
{ add: function() {} }, // poll
{ declare: function() { return function() {}; } }, // rpc
{ createHandlerFn: function() { return function() {}; }, addNotification: function() {} },
function() { return {}; }, // E
function(s) { return s; }, // _ (identity)
{ resolveDefault: function(p, d) { return Promise.resolve(d); } },
{ location: { hostname: 'router' }, open: function() { return null; } }
);
}
var page = loadView();
// --- Fixtures ---------------------------------------------------------------
var ENGINE_UP = {
daemon_answered: true,
running: true, engine_running: true, enabled: true, active: true, table: true,
config_readable: true, config_error: '',
plane: 'full', traffic: { verdict: 'tunnel', default: 'proxy', tunnel_rules: 3 },
hash: 'a1b2c3d4', kill_switch: 'closed', panel_port: 8088, can_rollback: true,
warnings: [], started_unix: 1753500000, uptime_seconds: 3600
};
var ENGINE_DOWN = {
daemon_answered: true,
running: false, engine_running: false, enabled: true, active: true, table: true,
config_readable: true, config_error: '',
plane: 'hold', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: 'closed', panel_port: 8088, can_rollback: true,
warnings: [], started_unix: 1753500000, uptime_seconds: 3600
};
// Exactly what Status.JSON() + markStatusOrigin(false) emit for the cmdStatus
// offline stub. NOTE config_readable:false: the stub reads UCI directly (so
// enabled/kill_switch/panel_port below ARE real readings) but never sets the
// field, so it ships the zero value. Anything keying on config_readable without
// first checking that a daemon answered will call this configuration unreadable.
var DAEMON_DOWN = {
daemon_answered: false,
running: false, engine_running: false, enabled: true, active: true, table: true,
config_readable: false, config_error: '',
plane: '', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: 'closed', panel_port: 8088, can_rollback: false,
warnings: null, started_unix: 0, uptime_seconds: 0
};
// A LIVE daemon whose readConfig() failed: /overlay full, or a `uci commit`
// interrupted. apply.go then leaves enabled=false, kill_switch="" and
// panel_port=0 as PLACEHOLDERS and raises config_readable=false + config_error.
// The engine cannot be generated, so the fail-closed holding plane is what is
// installed — which is the whole trap: the LAN is cut off and the three
// placeholders spell out a calm "the owner switched it off".
var CONFIG_BAD = {
daemon_answered: true,
running: false, engine_running: false, enabled: false, active: true, table: true,
config_readable: false,
config_error: 'uci: read /etc/config/shater: no space left on device',
plane: 'hold', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: '', panel_port: 0, can_rollback: false,
warnings: [], started_unix: 1753500000, uptime_seconds: 90
};
var NO_ANSWER = {};
var PLUGIN_ERROR = { error: 'shaterd unavailable' };
// Non-atomic package update: new luci-app-shater, old shaterd. No
// daemon_answered, no config_readable — the `plane` fallback is all there is.
var OLD_DAEMON_UP = {
running: false, engine_running: false, enabled: true, active: true, table: true,
plane: 'hold', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: 'closed', panel_port: 8088, can_rollback: true,
warnings: [], started_unix: 1753500000, uptime_seconds: 3600
};
var OLD_DAEMON_DOWN = {
running: false, engine_running: false, enabled: true, active: true, table: true,
plane: '', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: 'closed', panel_port: 8088, can_rollback: false,
warnings: null, started_unix: 0, uptime_seconds: 0
};
// Older still: no plane, no engine_running either.
var LEGACY = {
running: true, enabled: true, active: true, table: true, hash: 'deadbeef',
kill_switch: 'closed', panel_port: 8088
};
// A configured panel port that is NOT the built-in default, so "reported" and
// "fell back" are distinguishable in the readout and in the button hint.
var CUSTOM_PORT = Object.assign({}, ENGINE_UP, { panel_port: 9090 });
// --- Assertions -------------------------------------------------------------
var failures = [];
function check(name, cond, detail) {
if (cond) return;
failures.push(name + (detail ? ': ' + detail : ''));
}
// readout indexes the rows by label. A row that is NOT emitted must fail by name
// rather than by throwing on `undefined.state`: a harness that dies mid-run stops
// reporting the assertions after it, which is the silent-skip failure this
// project has been bitten by. Missing rows come back as a loud sentinel instead.
var MISSING = { state: '<row absent>', value: '<row absent>', missing: true };
function readout(st) {
var out = {};
page.statusReadout(st).forEach(function(r) { out[r.label] = r; });
return new Proxy(out, {
get: function(t, k) {
if (typeof k !== 'string' || k in t) return t[k];
return MISSING;
},
has: function(t, k) { return k in t; }
});
}
function show(title, st) {
process.stdout.write('\n=== ' + title + ' ===\n');
process.stdout.write(' daemon=' + page.daemonState(st) +
' engine=' + page.engineState(st) +
' config=' + page.configState(st) + '\n');
page.statusReadout(st).forEach(function(r) {
process.stdout.write(' [' + r.state.padEnd(7) + '] ' +
r.label.padEnd(18) + ' ' + r.value + '\n');
});
}
function lamps(st) {
return page.statusReadout(st).map(function(r) { return r.state; });
}
// 1. Live daemon, engine up.
show('A. daemon alive, engine running', ENGINE_UP);
check('A/daemon', page.daemonState(ENGINE_UP) === 'up');
check('A/engine', page.engineState(ENGINE_UP) === 'up');
check('A/config', page.configState(ENGINE_UP) === 'ok');
check('A/no-red', lamps(ENGINE_UP).indexOf('bad') === -1,
'a fully healthy router must show no red lamp');
// Closed positive list of the rows that a fully-reporting healthy router MUST
// have a verdict for. It replaces a blanket "no unknown anywhere", which stopped
// being the right assertion when the Panel port row was added: that row is unlit
// even here, on purpose, because nothing in a status measures whether the panel
// is listening. Naming the rows says which readings are owed, and a row that
// silently disappears fails here rather than passing as "no unknowns".
var MUST_BE_KNOWN = ['Daemon (shaterd)', 'Engine (sing-box)', 'Configuration',
'Service enabled', 'Service latch', 'Traffic plane', 'Traffic verdict',
'nft table', 'Kill-switch', 'Config hash'];
var a = readout(ENGINE_UP);
MUST_BE_KNOWN.forEach(function(label) {
check('A/known:' + label, a[label].state !== 'unknown' && !a[label].missing,
'every field is present, so this row may not read as unknown (got ' +
a[label].state + ')');
});
check('A/panel-port-unlit', a['Panel port'].state === 'unknown',
'the panel port is a CONFIGURED value; a lit lamp would promise a listener ' +
'that nothing in this status measures');
// 2. THE DEFECT. Live daemon, dead engine, LAN held.
show('B. daemon alive, engine DOWN, holding plane', ENGINE_DOWN);
check('B/daemon-up', page.daemonState(ENGINE_DOWN) === 'up',
'the daemon is answering; calling it dead is the bug being fixed');
check('B/engine-down', page.engineState(ENGINE_DOWN) === 'down');
var b = readout(ENGINE_DOWN);
check('B/daemon-row-green', b['Daemon (shaterd)'].state === 'good',
'got ' + b['Daemon (shaterd)'].state + ' / ' + b['Daemon (shaterd)'].value);
check('B/daemon-row-no-start-advice',
b['Daemon (shaterd)'].value.indexOf('start') === -1,
'must not tell the operator to start a service that is already running');
check('B/plane-red', b['Traffic plane'].state === 'bad');
check('B/plane-says-blocked', /BLOCKED/.test(b['Traffic plane'].value));
check('B/latch-not-called-interception',
!b['Service latch'].missing && b['Interception'].missing === true,
'`active` is the run latch, not a "we are proxying" signal — apply.go: ' +
'"Never render it as \'we are proxying\'"');
check('B/latch-value-is-a-latch', /meant to be running/.test(b['Service latch'].value),
'the latch row must state the latch, not claim traffic is being proxied');
check('B/latch-not-active-word', !/^active$/.test(b['Service latch'].value));
check('B/verdict-unknown', b['Traffic verdict'].state === 'unknown',
'no verdict was published; it must not be painted as tunnel');
check('B/some-red', lamps(ENGINE_DOWN).indexOf('bad') !== -1,
'a blocked LAN must not be an all-green screen');
// The button is a property of render(), not of the pure readout, so it is
// guarded at the source level: nothing may ever set `disabled` on the launcher.
// Locking the way into the panel while the engine is down is the defect this
// whole file exists for, and it must not come back by a different route.
check('B/button-never-disabled',
!/openBtn\s*\.\s*disabled/.test(fs.readFileSync(SRC, 'utf8')),
'dashboard.js assigns openBtn.disabled — the launcher must never be locked');
// 3. Daemon not answering at all — the offline stub, marked daemon_answered:false.
show('C. daemon NOT running (offline stub)', DAEMON_DOWN);
check('C/daemon-down', page.daemonState(DAEMON_DOWN) === 'down',
'daemon_answered:false is the contract; got ' + page.daemonState(DAEMON_DOWN));
check('C/engine-down', page.engineState(DAEMON_DOWN) === 'down');
var c = readout(DAEMON_DOWN);
check('C/daemon-row-red', c['Daemon (shaterd)'].state === 'bad');
check('C/plane-unknown', c['Traffic plane'].state === 'unknown',
'the stub reports no plane; got ' + c['Traffic plane'].value);
check('C/kill-switch-not-green', c['Kill-switch'].state !== 'good',
'"closed" from a dead daemon proves nothing is installed to enforce it');
check('C/distinct-from-B',
c['Daemon (shaterd)'].value !== b['Daemon (shaterd)'].value,
'engine-down and daemon-down must not render identically');
// The stub ships config_readable:false without ever having tried a config read
// (cmd/shaterd/main.go never sets it), while its enabled/kill_switch/panel_port
// ARE genuine UCI reads. Believing the field here would put a red "COULD NOT BE
// READ" on screen for a perfectly readable file and throw away the only facts a
// dead-daemon status carries.
check('C/config-not-claimed-unreadable', page.configState(DAEMON_DOWN) === 'unknown',
'config_readable from the offline stub is a zero value, not a reading; got ' +
page.configState(DAEMON_DOWN));
check('C/config-row-unknown', c['Configuration'].state === 'unknown',
'got ' + c['Configuration'].state + ' / ' + c['Configuration'].value);
check('C/config-row-blames-the-daemon', /no daemon answered/.test(c['Configuration'].value));
check('C/enabled-still-read', c['Service enabled'].state === 'good',
'the stub read globals.enabled from UCI; that reading must survive');
check('C/panel-port-still-read', page.panelTarget(DAEMON_DOWN).known === true,
'the stub read globals.panel_port from UCI; the button must use it');
// 4/5/6/7. Degenerate answers must degrade to unknown, never to healthy.
[['D. ubus call failed ({})', NO_ANSWER],
['E. rpcd plugin error / wedged daemon', PLUGIN_ERROR],
['F. legacy daemon (no plane, no engine_running)', LEGACY]].forEach(function(p) {
show(p[0], p[1]);
var st = p[1];
check(p[0] + '/daemon-unknown', page.daemonState(st) === 'unknown');
check(p[0] + '/engine-unknown', page.engineState(st) === 'unknown');
var r = readout(st);
check(p[0] + '/plane-unknown', r['Traffic plane'].state === 'unknown');
check(p[0] + '/daemon-row-unknown', r['Daemon (shaterd)'].state === 'unknown');
check(p[0] + '/config-unknown', r['Configuration'].state === 'unknown');
check(p[0] + '/no-false-green-plane', r['Traffic plane'].state !== 'good');
});
// The legacy fixture additionally must not crash and must not lose the fields it
// DOES carry — a non-atomic package update must degrade, not black out.
var f = readout(LEGACY);
check('F/enabled-still-read', f['Service enabled'].state === 'good');
check('F/hash-still-read', f['Config hash'].value === 'deadbeef');
check('F/table-still-read', f['nft table'].state === 'good');
// --- G. the compatibility fallback: an OLD shaterd under this LuCI ----------
// `apk upgrade shaterd shater-core luci-app-shater` is not atomic, so a new page
// will meet a daemon that emits no daemon_answered at all. `plane` must keep
// carrying the verdict on its own.
show('G1. OLD shaterd, alive (plane fallback)', OLD_DAEMON_UP);
check('G1/daemon-up', page.daemonState(OLD_DAEMON_UP) === 'up',
'plane:"hold" is a word only a live Applier.Status() writes; got ' +
page.daemonState(OLD_DAEMON_UP));
var g1 = readout(OLD_DAEMON_UP);
check('G1/daemon-row-green', g1['Daemon (shaterd)'].state === 'good');
check('G1/plane-red', g1['Traffic plane'].state === 'bad',
'the holding plane must still be reported as blocking');
check('G1/config-unknown', g1['Configuration'].state === 'unknown',
'an old daemon says nothing about config_readable; absence is not failure');
check('G1/enabled-still-read', g1['Service enabled'].state === 'good',
'a missing config_readable must NOT suppress an enabled that means what it says');
show('G2. OLD shaterd, dead (plane:"" fallback)', OLD_DAEMON_DOWN);
check('G2/daemon-down', page.daemonState(OLD_DAEMON_DOWN) === 'down',
'plane:"" is the pre-daemon_answered stub signature; got ' +
page.daemonState(OLD_DAEMON_DOWN));
check('G2/daemon-row-red', readout(OLD_DAEMON_DOWN)['Daemon (shaterd)'].state === 'bad');
// --- H. the field beats the side effect -------------------------------------
// If the two signals ever disagree, the explicit contract wins. A stub that
// somehow carried a plane word (a future change to cmdStatus, a merged object)
// must still read as "no daemon answered" — that is the entire reason the field
// was added, and keying off `plane` first would quietly restore the old defect.
var STUB_WITH_PLANE = Object.assign({}, DAEMON_DOWN, { plane: 'full' });
show('H. daemon_answered:false but plane:"full"', STUB_WITH_PLANE);
check('H/field-wins', page.daemonState(STUB_WITH_PLANE) === 'down',
'daemon_answered:false must outrank a plane word; got ' +
page.daemonState(STUB_WITH_PLANE));
var h = readout(STUB_WITH_PLANE);
check('H/daemon-row-red', h['Daemon (shaterd)'].state === 'bad');
// And the fields the stub CANNOT know are dropped rather than rendered. The
// stub's contract lists plane/traffic/hash among "placeholders, not
// measurements"; a green "full — ruleset, routing and engine are all installed"
// over a daemon that never answered is the original defect in a new costume.
check('H/plane-row-unknown', h['Traffic plane'].state === 'unknown',
'a plane word from an object marked daemon_answered:false is a placeholder; ' +
'got ' + h['Traffic plane'].state + ' / ' + h['Traffic plane'].value);
var STUB_WITH_EVERYTHING = Object.assign({}, DAEMON_DOWN, {
plane: 'full', hash: 'cafebabe',
traffic: { verdict: 'tunnel', default: 'proxy', tunnel_rules: 3 }
});
var he = readout(STUB_WITH_EVERYTHING);
check('H/verdict-row-unknown', he['Traffic verdict'].state === 'unknown',
'got ' + he['Traffic verdict'].value);
check('H/hash-row-not-shown', he['Config hash'].value.indexOf('cafebabe') === -1,
'a hash nobody measured must not be printed as the running config: ' +
he['Config hash'].value);
check('H/no-green-plane-anywhere', lamps(STUB_WITH_EVERYTHING).filter(function(s, i) {
return s === 'good' && ['Traffic plane', 'Traffic verdict', 'Config hash']
.indexOf(page.statusReadout(STUB_WITH_EVERYTHING)[i].label) !== -1;
}).length === 0,
'no engine-side row may be green while daemon_answered is false');
// Control for the three above: the SAME values from a daemon that did answer are
// rendered in full. Without this, "dropped" would also pass on a page that
// dropped them unconditionally.
var ANSWERED_WITH_EVERYTHING = Object.assign({}, STUB_WITH_EVERYTHING,
{ daemon_answered: true });
var ha = readout(ANSWERED_WITH_EVERYTHING);
check('H/control-plane-shown', ha['Traffic plane'].state === 'good');
check('H/control-verdict-shown', ha['Traffic verdict'].state === 'good');
check('H/control-hash-shown', ha['Config hash'].value === 'cafebabe');
// A non-boolean daemon_answered is not a verdict and must not be coerced: it
// falls through to the fallback, which here proves the daemon live by itself.
var WEIRD_FLAG = Object.assign({}, ENGINE_UP, { daemon_answered: 'true' });
check('H/non-boolean-not-coerced', page.daemonState(WEIRD_FLAG) === 'up',
'a string is not the field; the plane fallback must decide instead');
var WEIRD_FLAG_NO_PLANE = { daemon_answered: 'yes' };
check('H/non-boolean-alone-is-unknown',
page.daemonState(WEIRD_FLAG_NO_PLANE) === 'unknown',
'with nothing else to go on, a non-boolean flag is not an answer');
// --- I/J. unrecognised plane words ------------------------------------------
// Closed positive list, recoverable default — a word this build does not know
// lands in unknown, never in the last-listed branch.
var FUTURE_PLANE = Object.assign({}, ENGINE_UP, { plane: 'partial' });
check('I/daemon-still-up', page.daemonState(FUTURE_PLANE) === 'up',
'the daemon SAID it answered; an unknown plane word does not unsay it');
check('I/plane-row-unknown', readout(FUTURE_PLANE)['Traffic plane'].state === 'unknown');
var OLD_FUTURE_PLANE = Object.assign({}, OLD_DAEMON_UP, { plane: 'partial' });
check('J/fallback-list-is-closed', page.daemonState(OLD_FUTURE_PLANE) === 'unknown',
'with no contract field, an unrecognised plane value proves nothing');
check('J/plane-row-unknown', readout(OLD_FUTURE_PLANE)['Traffic plane'].state === 'unknown');
// --- K. THE CONFIGURATION CANNOT BE READ ------------------------------------
// A live daemon with config_readable:false. enabled/kill_switch/panel_port are
// zero values; rendering them as readings is the inverted lie the panel already
// fixed on its side (planeState.ts protectionState / killSwitchReadout).
show('K. daemon alive, CONFIGURATION UNREADABLE', CONFIG_BAD);
check('K/daemon-up', page.daemonState(CONFIG_BAD) === 'up');
check('K/config-failed', page.configState(CONFIG_BAD) === 'failed');
var k = readout(CONFIG_BAD);
check('K/config-row-red', k['Configuration'].state === 'bad',
'got ' + k['Configuration'].state);
check('K/config-row-carries-the-error',
k['Configuration'].value.indexOf('no space left on device') !== -1,
'the daemon said WHY; dropping it makes the operator guess');
check('K/config-row-says-dont-switch-off', /do not turn anything off/i.test(k['Configuration'].value),
'the instinct here is to switch things off, and that is the one action that ' +
'makes it worse');
check('K/enabled-not-called-disabled', k['Service enabled'].state === 'unknown',
'enabled=false with config_readable=false is a zero value, not the owner\'s ' +
'choice; got ' + k['Service enabled'].state + ' / ' + k['Service enabled'].value);
check('K/enabled-says-not-known', /not known/.test(k['Service enabled'].value));
check('K/kill-switch-not-green', k['Kill-switch'].state === 'unknown',
'kill_switch:"" must not normalise to a green "fail-closed"; got ' +
k['Kill-switch'].state + ' / ' + k['Kill-switch'].value);
check('K/kill-switch-says-not-known', /not known/.test(k['Kill-switch'].value));
check('K/latch-lamp-out', k['Service latch'].state === 'unknown',
'the latch is readable but the comparison against globals.enabled is not');
check('K/latch-still-states-the-latch', /raised/.test(k['Service latch'].value),
'the fact survives even though the verdict does not');
check('K/plane-red', k['Traffic plane'].state === 'bad',
'the holding plane is what is installed, and it is blocking');
check('K/panel-port-unknown', page.panelTarget(CONFIG_BAD).known === false,
'panel_port:0 is a placeholder, not a port');
check('K/panel-port-falls-back-to-default', page.panelTarget(CONFIG_BAD).port === 8088);
check('K/panel-port-row-says-so', /could not be read/.test(k['Panel port'].value),
'got ' + k['Panel port'].value);
// The same fixture WITHOUT a table loaded: "not loaded" may only be the calm
// amber when the service is KNOWN to be off. With no readable configuration it
// is not known, so the alarming lamp is the correct one.
var CONFIG_BAD_NO_TABLE = Object.assign({}, CONFIG_BAD, { table: false });
check('K/table-absent-is-red',
readout(CONFIG_BAD_NO_TABLE)['nft table'].state === 'bad',
'enabled=false is a placeholder here and must not soften an absent firewall ' +
'table to amber');
// Control for the line above: with a READABLE configuration that says disabled,
// the same absent table IS the calm amber. Without this, "red" proves nothing.
var DISABLED_NO_TABLE = Object.assign({}, ENGINE_UP,
{ enabled: false, table: false, plane: 'none', kill_switch: 'open' });
check('K/table-absent-is-amber-when-really-disabled',
readout(DISABLED_NO_TABLE)['nft table'].state === 'warn',
'a service the owner switched off has no table on purpose');
// A plane:"none" whose kill-switch setting is unknown must not carry the
// "traffic reaches the WAN unprotected" claim — that sentence is earned by a
// KNOWN fail-closed setting.
var CONFIG_BAD_NO_PLANE = Object.assign({}, CONFIG_BAD, { plane: 'none' });
check('K/none-plane-claims-nothing-about-protection',
!/unprotected/.test(readout(CONFIG_BAD_NO_PLANE)['Traffic plane'].value),
'with an unreadable configuration the kill-switch setting is not known, so ' +
'"unprotected" is a claim this page cannot make');
// Control: with a readable configuration that says closed, the claim IS made.
var CLOSED_NO_PLANE = Object.assign({}, ENGINE_UP,
{ plane: 'none', kill_switch: 'closed' });
check('K/none-plane-does-claim-it-when-known',
/unprotected/.test(readout(CLOSED_NO_PLANE)['Traffic plane'].value) &&
readout(CLOSED_NO_PLANE)['Traffic plane'].state === 'bad',
'a known fail-closed setting with nothing installed IS the dangerous state');
// --- L. the launcher: a configured port is not a listening one ---------------
// shaterd starts the panel server in a goroutine and only LOGS a bind failure
// (cmd/shaterd/main.go: "panel server unavailable (daemon continues)"), and
// SHATER_PANEL_ADDR can disable it outright while apply.Status keeps reporting a
// port. Nothing on the wire says the panel is listening, so nothing on this page
// may say it either.
check('L/custom-port-used', page.panelTarget(CUSTOM_PORT).port === 9090 &&
page.panelTarget(CUSTOM_PORT).known === true);
check('L/no-answer-falls-back', page.panelTarget(NO_ANSWER).port === 8088 &&
page.panelTarget(NO_ANSWER).known === false);
check('L/garbage-falls-back', page.panelTarget(null).port === 8088);
var hintKnown = page.panelHint('router', page.panelTarget(CUSTOM_PORT));
process.stdout.write('\n=== L. launcher hint / tooltip ===\n');
process.stdout.write(' hint(known) ' + hintKnown + '\n');
check('L/hint-names-the-url', hintKnown.indexOf('http://router:9090/') !== -1,
'got ' + hintKnown);
check('L/hint-marks-the-port-unchecked', /not a checked one/.test(hintKnown),
'the hint must not read as "the panel is there": ' + hintKnown);
check('L/hint-has-no-placeholders-left', hintKnown.indexOf('%') === -1,
'unsubstituted placeholder in the hint: ' + hintKnown);
var hintUnknown = page.panelHint('router', page.panelTarget(CONFIG_BAD));
process.stdout.write(' hint(unknown) ' + hintUnknown + '\n');
check('L/hint-admits-the-guess', /built-in default/.test(hintUnknown),
'a port nobody reported must be named as this page\'s guess: ' + hintUnknown);
check('L/hint-unknown-has-no-placeholders-left', hintUnknown.indexOf('%') === -1,
'unsubstituted placeholder in the hint: ' + hintUnknown);
// Every tooltip, on every state, carries the port caveat. Closed list of the
// states, so a new branch added without the caveat fails here.
[['up', ENGINE_UP], ['engine-down', ENGINE_DOWN], ['daemon-down', DAEMON_DOWN],
['config-bad', CONFIG_BAD], ['no-answer', NO_ANSWER]].forEach(function(p) {
var t = page.panelTitle(p[1]);
process.stdout.write(' title(' + p[0] + ') ' + t + '\n');
check('L/title-caveat:' + p[0],
/configured one/.test(t) || /built-in default/.test(t),
'the tooltip promises a panel without saying the port is unverified: ' + t);
check('L/title-not-empty:' + p[0], typeof t === 'string' && t.length > 20);
});
// --- Report -----------------------------------------------------------------
process.stdout.write('\n');
if (failures.length) {
process.stdout.write('FAIL (' + failures.length + ')\n');
failures.forEach(function(f) { process.stdout.write(' - ' + f + '\n'); });
process.exit(1);
}
process.stdout.write('OK — all cases distinguished\n');
+34 -1
View File
@@ -40,6 +40,10 @@ define Package/shater-core
# shaterd : the daemon our init supervises (`shaterd run`)
# kmod-nft-tproxy : kernel TPROXY (shaterd emits the `inet shater` rules)
# kmod-nft-socket : socket match used by the tproxy divert chain
# kmod-tun : /dev/net/tun — the daemon opens the `shater-l3` TUN
# for L3 ingress (globals.l3_tunnel); usually built-in
# on stock images, but a slimmed image without it would
# make the option fail with a cryptic open() error.
# ip-full : `ip rule`/`ip route`/rt_tables for policy routing
# nftables-json : shaterd shells out to `nft`, and netplane/stats.go
# parses `nft -j list ...` — the JSON output only exists
@@ -49,7 +53,7 @@ define Package/shater-core
# ca-bundle : the daemon is CGO_ENABLED=0, so crypto/x509 has no
# host cert fallback — without /etc/ssl/certs every
# HTTPS subscription / .srs ruleset fetch fails.
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full +nftables-json +ca-bundle
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle
PKGARCH:=all
endef
@@ -80,6 +84,10 @@ define Package/shater-core/install
$(INSTALL_DIR) $(1)/etc/init.d
$(INSTALL_BIN) ./files/etc/init.d/shater $(1)/etc/init.d/shater
$(INSTALL_BIN) ./files/etc/init.d/shater-cron $(1)/etc/init.d/shater-cron
# START=21 one-shot that loads the persisted fail-closed plane before fw4's
# `lan -> wan ACCEPT` can be the only thing on the box (the main init is
# START=99, i.e. seconds of plaintext forwarding on every boot).
$(INSTALL_BIN) ./files/etc/init.d/shater-armor $(1)/etc/init.d/shater-armor
$(INSTALL_DIR) $(1)/etc/hotplug.d/iface
$(INSTALL_BIN) ./files/etc/hotplug.d/iface/99-shater $(1)/etc/hotplug.d/iface/99-shater
@@ -90,8 +98,33 @@ define Package/shater-core/install
$(INSTALL_DIR) $(1)/etc/config
$(INSTALL_CONF) ./files/etc/config/shater $(1)/etc/config/shater
# THE SAME FILE AGAIN, READ-ONLY, AS DOCUMENTATION. /etc/config/shater is 271
# lines of which 248 are comment, and on the router it is the only description
# of the schema there is (PORTING.md does not ship). Being a conffile keeps an
# upgrade from replacing it, but it does NOT keep the daemon from rewriting it:
# the config write path replaces the whole package (`uci delete shater` + `uci
# import`), which drops every comment — and it runs without an operator, from
# the panel, the 6-hourly subscription refresh and the 25-second profile
# watcher. So the annotated original is installed a second time where nothing
# rewrites it, and the header of the live file points at it.
#
# INSTALL_DATA, not INSTALL_CONF: this copy is package metadata (refreshed by
# every upgrade so it documents the build actually installed), not user config.
$(INSTALL_DIR) $(1)/usr/share/shater
$(INSTALL_DATA) ./files/etc/config/shater $(1)/usr/share/shater/config.sample
$(INSTALL_DIR) $(1)/etc/uci-defaults
$(INSTALL_BIN) ./files/etc/uci-defaults/30_shater-core $(1)/etc/uci-defaults/30_shater-core
# sysupgrade's "keep settings" walks /lib/upgrade/keep.d/*, and without this the
# node inventory in /etc/shater/subs does NOT survive a flash: the restored box
# has its rules and its groups and no nodes for them to point at, and the only
# repair is `sub update`, which needs the internet the tunnel was going to
# provide. Package metadata, not user config, so INSTALL_DATA and not
# INSTALL_CONF. (/etc/config/shater needs no entry — it is a conffile and
# sysupgrade already keeps it that way.)
$(INSTALL_DIR) $(1)/lib/upgrade/keep.d
$(INSTALL_DATA) ./files/lib/upgrade/keep.d/shater-core $(1)/lib/upgrade/keep.d/shater-core
endef
$(eval $(call BuildPackage,shater-core))
+72 -9
View File
@@ -12,8 +12,30 @@
# from this file. There is no separate xray/dnsmasq and no generated run.json.
#
# Full schema: docs-shater/PORTING.md (PART A "uci.go — /etc/config/shater
# schema") and shater/model. This file is installed as a conffile — your edits
# survive package upgrades.
# schema") and shater/model.
#
# WHAT SURVIVES WHAT. This file is installed as a conffile, so `apk upgrade
# shater-core` will not replace it: your VALUES survive a package upgrade.
#
# These COMMENTS do not survive the first write, and that write does not need
# you to make it. The daemon and the panel persist the whole package in one go
# (`uci delete shater` + `uci import`), and a package rebuilt by `uci import`
# keeps no comments and no hand-made blank lines; the sections come back in the
# daemon's own order. Three things write here with nobody at the keyboard: a save
# in the admin panel, a subscription refresh (cron, every 6 h) and the profile
# watcher switching profiles (it looks every 25 s). So expect the annotations
# below to be gone shortly after the box is first configured.
#
# THE ANNOTATED COPY IS KEPT: /usr/share/shater/config.sample is this same file,
# installed by the package where nothing rewrites it. Read it there (`cat
# /usr/share/shater/config.sample`) and copy the fragment you need. It is
# refreshed by each package upgrade, so it always documents the build you have.
#
# BEFORE THE FIRST CHANGE, the previous file is copied to
# /etc/shater/config.pre-v<schema>.bak — once per schema version, never
# overwritten afterwards. That copy is the one taken at the transition (a schema
# migration, or the first save on a newly installed build); it is not a rolling
# backup, and it is deliberately not carried across a sysupgrade.
#
config globals 'globals'
@@ -49,17 +71,58 @@ config globals 'globals'
# Queries aimed at an EXTERNAL resolver are dropped with the rest of the LAN's
# forwarded traffic.
option dns_intercept '1'
# Carry LAN ping through the tunnel. ON by default, and the alternative is
# why: without it a ping is decided by `untunnelable` below, whose rungs are
# "drop it" (block, the default) or "let it out of the WAN interface with the
# client's real IP on it" (icmp/direct). There was no setting in which ping
# both worked and stayed inside the tunnel. With this on, the engine opens a
# TUN, LAN ICMP is routed into it, and an outbound that speaks layer 3
# (WireGuard/AmneziaWG, or a direct route) carries the echo for real. An
# outbound that does not (vless/trojan/shadowsocks) makes the ping DROP —
# honestly: no reply is forged, ping reports loss. So a ping that used to
# "work" through such a node was a ping that was leaking.
#
# It costs a permanent TUN device plus the gVisor netstack behind it, about
# 2 MB of RSS for as long as the daemon runs.
#
# Set to '0' to opt out — worth it on a 32/64 MB router, or to bisect whether
# the L3 ingress is what broke something. `shaterd apply` will tell you what
# the off state costs. Your explicit value is never overwritten: this file is
# a conffile and the daemon always writes the option back as '1'/'0'.
#
# NOTE the interaction: with this ON, `untunnelable` no longer governs ping at
# all (the L3 route decision happens before the firewall chain its verdicts
# live in). It still governs ESP/AH/GRE/IGMP/SCTP, which no tunnel of ours can
# carry. `untunnelable 'icmp'` in particular stops meaning "block plus working
# ping" and is reported as such.
option l3_tunnel '1'
option ipv6 '1'
# Reserved fwmark base and routing-table base (do not overlap fw4/other apps).
option fwmark_base '0x2000'
option table_base '0x2000'
# Seconds to auto-rollback an unconfirmed apply (0 = commit-confirm off).
# Seconds to auto-rollback an unconfirmed apply. SHIPPED AS 0, i.e.
# commit-confirm is OFF: `shaterd apply` arms nothing, and an apply that costs
# you SSH/LuCI access stays until you undo it by hand. Set a window (e.g.
# '120') to arm it, and run `shaterd confirm` inside that window to keep the
# new config. Note the option is written back only when NON-zero, so an
# explicit '0' disappears from this file on the first write by the daemon or
# the panel — absent and 0 are the same thing.
option confirm_timeout '0'
option schema_version '1'
# Master enable of the DNS blocklist/allowlist filter (D15). OFF by default;
# it needs at least one `config resolver` to have a DNS plane to filter with.
# See the "DNS filter" section at the end of this file.
option dns_filter '0'
option schema_version '2'
# LAN interception inbound. `network` is a UCI interface name; shaterd resolves
# it to its device (e.g. 'lan' -> br-lan) for the nft TPROXY plane. Enable
# globals above and adjust `network` to the interface(s) you want proxied.
#
# There is no per-inbound `sniff` option: since sing-box 1.11 sniffing is a
# leading route ACTION rule with no inbound matcher, so EVERY inbound is sniffed,
# always. Do not add one back — the hijack-dns rule matches the SNIFFED `dns`
# protocol, so a per-inbound sniff toggle would be a DNS-leak switch (D14, and
# the long argument at shater/model/model.go Inbound).
config inbound
option name 'lan'
option enabled '1'
@@ -68,7 +131,6 @@ config inbound
option tproxy_port '12345'
option tcp '1'
option udp '1'
option sniff '1'
# --- Commented examples (copy, uncomment, adjust, then enable globals) -------
#
@@ -124,8 +186,9 @@ config inbound
# record - tls_record_fragment (alternative; mutually exclusive w/ fragment)
# spoof - tls_spoof (inject a decoy ClientHello; needs root NET_RAW/NET_ADMIN)
# Point a rule's target at it for DPI-blocked-but-not-IP-blocked domains — direct
# and fragmented, no exit node, no extra binary. (The stronger external `byedpi`
# preset is Phase-2b.)
# and fragmented, no exit node, no extra binary. These three are the WHOLE set;
# the external desync egress that once stood beside them is removed (D29), and a
# `type 'byedpi'` egress left over from an older build is blocked, not routed.
#config egress
# option name 'frag'
# option type 'direct'
@@ -153,8 +216,8 @@ config inbound
#
# --- DNS filter (D15) -------------------------------------------------------
# Network-wide domain blocking, built on sing-box rule-sets + reject DNS rules.
# Turn it ON by setting `option dns_filter '1'` in `config globals` above (it is
# OFF by default). Filtering needs at least one `config resolver` (the in-engine
# Turn it ON by flipping `option dns_filter` to '1' in `config globals` above (it
# is shipped '0'). Filtering needs at least one `config resolver` (the in-engine
# DNS plane). A blocklist answers matched domains with NXDOMAIN; an allowlist
# always OVERRIDES the blocklists (allowlisted domains resolve normally).
#
+408 -15
View File
@@ -33,6 +33,30 @@
# be running. `start` raises ACTIVE_FLAG, `stop` clears it; hotplug/cron
# reconcile ONLY while the flag is up, so an admin `stop` STICKS — no
# background actor may resurrect interception behind a stopped daemon.
# * BEING REPLACED IS NOT BEING SWITCHED OFF. `restart` and `reload` (which is
# stop+start, i.e. every LuCI Save & Apply) both run through `stop`, and the
# daemon's SIGTERM teardown removes the fail-closed table unconditionally — it
# does not consult kill_switch at all. Between that teardown and the
# successor's first apply the init GUARANTEES a gap: it waits for the old
# process to exit (shater_wait_stopped), then runs `shaterd migrate`, then
# starts a daemon that still has to build an engine. So a restart is announced
# with RESTART_FLAG, which tells the outgoing daemon to leave the fail-closed
# holding plane STANDING — apply.TeardownExiting swaps it in with one nft
# transaction and then skips the delete, so the table is never absent, not even
# for the 80-90 ms the old arm-after-teardown order measured. A real `stop`
# raises no flag and therefore still means what it says.
# (A package UPGRADE does not come through here at all on apk v3: shater-core's
# script table is post-install / pre-deinstall / post-upgrade, with no
# pre-upgrade, so default_prerm — and its `stop` — runs only on REMOVAL.)
# * The FAIL-CLOSED PLANE MUST ALSO EXIST BEFORE THIS SCRIPT DOES. START=99 is
# after fw4 (19) and netifd (20), so at every boot the LAN forwards to the WAN
# in the clear for as long as it takes procd to decompress the daemon off
# flash and get an engine up. /etc/init.d/shater-armor (START=21) loads
# BOOT_ARMOR — a copy of the holding plane the daemon persists on every apply
# — to close that window. This script owns the DISARM half, and it owns it
# with a CLOSED LIST: an operator's `stop`, or a removal, and nothing else.
# Powering the box down must not — `shutdown` reaches stop_service too, and it
# is not a person switching the product off (see shater_stop_disarms).
# * The engine must never be permanently abandoned while interception stands:
# respawn retries are infinite (procd never gives up); a sustained-dead
# daemon is additionally escalated by the shater-cron watchdog.
@@ -50,11 +74,166 @@ ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Raised around a restart/reload, read by the OUTGOING `shaterd run` at SIGTERM:
# present => "you are being replaced, leave the fail-closed plane standing";
# absent => "you are being switched off, take everything down". tmpfs, so a
# power cut can never make the next boot look like a restart.
RESTART_FLAG=/var/run/shater.restarting
# The persisted fail-closed holding plane. Written by the daemon on every apply,
# loaded by /etc/init.d/shater-armor at boot. Its PRESENCE is the arm token, so
# removing it here is how a deliberate stop stops the next boot from blocking.
BOOT_ARMOR=/etc/shater/boot.nft
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# WHICH ACTION rc.common was invoked with, frozen at source time.
#
# rc.common does, in this order:
# initscript=$1; action=${2:-help}; shift 2; ...; . "$initscript"; $action "$@"
# so `action` is ALREADY assigned when this file is sourced, and every action then
# runs as a function in THAT SAME shell. MEASURED on the target (ImmortalWrt
# 25.12.1 r37978) with a throwaway probe init script, not read off documentation:
#
# /etc/init.d/X restart -> stop_service action=[restart], start_service [restart]
# /etc/init.d/X stop -> stop_service action=[stop]
# /etc/init.d/X reload -> reload_service action=[reload]
# `reboot` -> stop_service action=[SHUTDOWN] <-- see below
# the boot after it -> start_service action=[boot]
#
# A previous probe reported this variable EMPTY and the emptiness was written up as
# the defect. It was the probe: `sh -x /etc/init.d/shater restart` bypasses the
# `#!/bin/sh /etc/rc.common` shebang, so rc.common never runs, never assigns
# `action`, and the variable reads empty no matter what this file does.
#
# Frozen into our own variable because `action` is a short, generic name that other
# framework helpers also use as a local; a snapshot taken before any function runs
# cannot be shadowed later.
SHATER_RC_ACTION="$action"
# --- what an action MEANS --------------------------------------------------
#
# THE BUG THESE TWO PREDICATES REPLACE (v0.2.17, measured on the live router).
# The old stop_service was `case $action in restart|reload) keep;; *) DISARM;; esac`
# — an open default that swept up every action nobody had enumerated. `reboot` is
# one of them: procd runs the K-links with the action `shutdown`, so the shutdown
# path deleted the arm token on the way down and the next boot had nothing to load.
# The mechanism destroyed itself at exactly the moment it exists for. Instrument
# reading from the router, one minute apart across a reboot:
#
# 13:28 /etc/shater/boot.nft present
# ---- reboot (stop_service action=[shutdown] -> old `*` branch -> rm)
# 18s at_S22: NO_TABLE armor_file=NO_FILE
#
# So both lists below are POSITIVE and CLOSED. An action nobody thought about —
# `shutdown` above all, but also whatever a future procd invents — falls through
# both and changes nothing. The default now fails in the recoverable direction: at
# worst a boot arms when it need not have, which costs the second before the daemon
# applies and is still gated by shater-armor's own four state refusals. The old
# default failed in the direction of the plaintext window the feature was built to
# close.
#
# They are predicates rather than an inline `case` so the test gate can execute the
# real thing: it sources THIS FILE in /bin/sh and calls them with every action procd
# actually uses (shater/cmd/shaterd/initscript_test.go). A comment claiming
# `shutdown` is handled is what shipped last time.
# True only for the ONE action that means "the operator switched the product off".
# Deliberately not `shutdown`: powering a router down is not turning a feature off.
#
# NOT sufficient on its own — see shater_stop_disarms. `stop` is also how the
# package manager's plumbing reaches us, and a package manager is not a person.
shater_action_disarms() {
case "$1" in
stop) return 0 ;;
*) return 1 ;;
esac
}
# Is a package manager in the middle of a transaction RIGHT NOW?
#
# This is a state, read at the moment the decision is made, exactly like
# shater-armor's four refusals — not a record of an event. The same question is
# already asked (for the same reason: prerm/postinst plumbing is not a user
# action) by the detached bring-up in /etc/uci-defaults/30_shater-core.
shater_pkg_transaction() {
pidof apk >/dev/null 2>&1 && return 0
pidof opkg >/dev/null 2>&1 && return 0
return 1
}
# Is the main service still enabled at boot? Same glob, and for the same reason,
# as shater-armor's own check: `/etc/init.d/shater enabled` would source procd.sh
# and take a blocking flock, which is not something to do from inside a package
# manager's transaction.
shater_rc_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
# THE ACTUAL DISARM DECISION.
# $1 = action
# $2 = 1 when a package transaction is in flight
# $3 = 1 when the service is still enabled in rc.d
# All three are passed in rather than read inside, so the gate can drive every
# combination without a package manager or an /etc/rc.d.
#
# WHY IT IS NOT JUST THE ACTION. base-files' default_prerm runs, in this order:
#
# if [ "$PKG_UPGRADE" != "1" ]; then "$i" disable; fi
# "$i" stop
#
# so a package manager reaches stop_service wearing the operator's clothes. Two
# different intentions arrive as the same action, and the difference between them
# is readable at the moment of the decision:
#
# REMOVAL — prerm has ALREADY run `disable`, so S99shater is gone. The product
# is going away; the armor goes with it. (It is belt-and-braces even
# so: shater-armor refuses to arm without that symlink, and the whole
# init script is about to be deleted anyway.)
# REPLACED — the service is still enabled, so something intends to bring it
# back. That is not an operator switching anything off, and deleting
# the armor here would leave the next boot unprotected. "The next
# apply will rewrite it" is not an answer: the armor exists precisely
# to cover a reboot, and a reboot between an update and the first
# apply is how this product is deployed.
#
# MEASURED, because the paragraph above is about a path I got wrong once already.
# On THIS target (apk-tools 3.0.5, ImmortalWrt 25.12.1) shater-core's script table
# is post-install / pre-deinstall / post-upgrade, with NO pre-upgrade — so an apk
# UPGRADE never executes default_prerm and never calls `stop` at all. Verified with
# a real `apk fix --reinstall shater-core` while sampling the armor file: 245 625
# samples, zero disappearances, even with this guard mutated off. The upgrade half
# of this predicate is therefore defence-in-depth for a shape that is one
# `pre-upgrade` script (or a returning opkg lane) away, NOT a fix for an observed
# failure. The removal half is live today.
shater_stop_disarms() {
shater_action_disarms "$1" || return 1
# No package manager involved => a person typed it. The escape hatch must work.
[ "$2" = "1" ] || return 0
# A package transaction that has NOT disabled the service is replacing it.
[ "$3" = "1" ] && return 1
return 0
}
# True when a successor is coming, so the outgoing daemon should leave the
# fail-closed holding plane standing instead of removing it.
#
# `shutdown` is deliberately NOT a handoff either: nothing is coming, and the
# kernel that would hold the plane is going away with it. Leaving the flag down
# there also keeps the marker's meaning exact — it says "you are being replaced",
# and at shutdown nothing is.
shater_action_handoff() {
case "$1" in
restart|reload) return 0 ;;
*) return 1 ;;
esac
}
# --- helpers ---------------------------------------------------------------
# True only when the stack is explicitly enabled in UCI.
@@ -73,6 +252,161 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# A line the operator gets EVEN WITH globals.log_syslog=0, without going behind
# that setting's back.
#
# log_syslog is a statement about ONE destination: the syslog stream (see _slog
# above — the same toggle silences the daemon's own stderr->logread fan-out).
# Honouring it by staying silent everywhere turns "keep syslog quiet" into "never
# tell me the config could not be brought forward", which is not what it says and
# not what anybody means by it. So the refusal goes somewhere else instead:
#
# * THIS SCRIPT'S OWN STDERR, unconditionally. That is not the syslog stream; it
# is the reply to whoever invoked the script. Typed by hand it lands on the
# operator's terminal at the moment they are looking at it; run from
# 30_shater-core inside `apk add` / `opkg install` it lands in the package
# manager's output, which is the one screen an installing operator does read.
# At boot it goes to procd's stderr (console) — not durable, hence the file.
# * /etc/shater/migrate-failed, on flash, written on failure and REMOVED on the
# first success. That is the durable half: it survives the reboot nobody
# watched, `cat` reads it, and its absence is the honest all-clear. Written
# AFTER the stderr line on purpose — a full /overlay is one of the named
# causes of the failure it is reporting, so it must never be the only channel.
# * syslog too, but only when log_syslog allows it — that channel keeps
# obeying the operator exactly as before.
#
# One call, one text, three destinations, so the wording cannot drift between
# them. Failures only: _shout is not a status line.
SHATER_MIGRATE_BREADCRUMB=/etc/shater/migrate-failed
_shout() {
echo "shater: $*" >&2
_slog -p daemon.err "$*"
mkdir -p "$(dirname "$SHATER_MIGRATE_BREADCRUMB")" 2>/dev/null
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') $*" \
> "$SHATER_MIGRATE_BREADCRUMB" 2>/dev/null || :
}
# --- `shaterd migrate`: which of the four things happened --------------------
#
# The schema version on disk, as an integer. 0 for "absent" and 0 for anything
# non-numeric, DELIBERATELY the same two answers model.readSchemaVersion gives
# (`strconv.Atoi` of a garbage value is 0 with the error dropped) — this number
# is only ever used to name a version in a message, and a shell that disagreed
# with the binary about what v0 means would print a version the binary never saw.
shater_schema_version() {
local v
v=$(uci -q get shater.globals.schema_version) || v=""
case "$v" in
"") echo 0 ;;
*[!0-9]*) echo 0 ;;
*) echo "$v" ;;
esac
}
# shater_migrate_class <rc> <output-of-shaterd-migrate> — prints EXACTLY one of:
#
# ok the binary reported success (it may or may not have had work)
# downgrade REFUSED: the config on disk is NEWER than this build
# unreadable /etc/config/shater could not be read at all
# failed it failed for a reason this script does not recognise
#
# A CLOSED POSITIVE LIST, and the last rung is the point of it. Until v0.2.19 all
# four of these were reported with ONE sentence, and that sentence described only
# the third one: "routing rules that still carry the removed dst_domain/dst_ip
# options stay DISABLED until this succeeds. Free space on /overlay and re-run".
# On a DOWNGRADE every clause of that is false — nothing is disabled, /overlay is
# not the problem, and re-running does not help, because the fix is to put the
# newer package back. A confident wrong diagnosis costs more than no diagnosis.
#
# `downgrade` is recognised from the binary's own words. That is a CONTRACT with
# shater/model: both refusals — model.migrateWith's "config schema v%d newer than
# this build (v%d); upgrade the package" and model.ErrSchemaTooNew's "config
# schema newer than this build" — contain the substring matched below, and
# TestMigrateDowngradeSignatureIsAContract (shater/cmd/shaterd) fails if either
# stops containing it. If the wording is ever changed anyway, this degrades to
# `failed`, which names itself as unrecognised and quotes the binary verbatim —
# the recoverable side. It cannot degrade into one of the confident branches.
shater_migrate_class() {
local rc="$1" out="$2"
[ "$rc" = "0" ] && { echo ok; return 0; }
case "$out" in
*"newer than this build"*) echo downgrade; return 0 ;;
esac
# Asked LAST, so a refusal we can name is never re-labelled as an I/O problem.
# `uci export` fails both when the file is missing and when it does not parse,
# which is the same thing from here: nothing can be said about a schema that
# cannot be read.
uci -q export shater >/dev/null 2>&1 || { echo unreadable; return 0; }
echo failed
}
# Run the migration and report it. Called from start_service and mirrored by
# /etc/uci-defaults/30_shater-core; see the long note at the call site for why
# this never refuses to start.
shater_migrate() {
local before after out rc class
before=$(shater_schema_version)
out=$("$PROG" migrate 2>&1)
rc=$?
class=$(shater_migrate_class "$rc" "$out")
after=$(shater_schema_version)
case "$class" in
ok)
# The all-clear is the ABSENCE of the breadcrumb, so a fixed router stops
# claiming to be broken the moment it is fixed.
rm -f "$SHATER_MIGRATE_BREADCRUMB"
# Nothing to do is not news; obeys log_syslog like every other status line.
[ "$before" = "$after" ] && return 0
_slog -p daemon.info \
"UCI schema migrated: v$before -> v$after. The config as it was at v$before was copied to /etc/shater/config.pre-v$before.bak before the first change."
;;
downgrade)
_shout "UCI schema migration REFUSED — this is a DOWNGRADE, not a broken config. /etc/config/shater carries schema v$after, which is NEWER than this build understands, so nothing was migrated and nothing on disk was changed. Your settings are intact; they are also unchangeable, because the daemon and the panel refuse every config write for the same reason and a save from the panel will fail too. Nothing on this router fixes it: install a shater build that understands schema v$after — the one that ran here before the downgrade (docs-shater/INSTALL.md has the pinned per-version feed). Starting anyway, so the panel stays reachable. '$PROG migrate' said: ${out:-no output}"
;;
unreadable)
_shout "UCI schema migration FAILED and /etc/config/shater CANNOT BE READ ('uci export shater' fails), so this script cannot even say which schema is on disk. A config that is missing or does not parse is neither migrated nor repaired here. Starting anyway — the daemon will come up on whatever it can parse, which may be nothing, leaving it inert with the panel still reachable. Check /etc/config/shater by hand; an /etc/shater/config.pre-v*.bak copy from an earlier migration may be next to it. '$PROG migrate' said: ${out:-no output}"
;;
failed)
_shout "UCI schema migration FAILED for a reason this script does not recognise; the config on disk is still at schema v$after. Starting anyway: refusing to start would take the admin panel down with it, and the panel is the only way to fix the box. The mundane cause is a full /overlay, where 'uci commit' cannot write — check 'df /overlay' first, then re-run '$PROG migrate' or restart the service.$(
[ "$after" = "1" ] && printf ' %s' "While the config stays at v1, routing rules that still carry the removed dst_domain/dst_ip options are held DISABLED by the daemon and reported as such — those rules are not in force."
) '$PROG migrate' said: ${out:-no output}"
;;
*)
# shater_migrate_class returns a closed set and every member of it is
# handled above, so this is unreachable. It exists to say "this script
# disagrees with itself" out loud instead of picking one of the confident
# branches and being wrong quietly — which is the exact failure the closed
# list replaced.
_shout "INTERNAL: '$PROG migrate' produced a result /etc/init.d/shater cannot classify (class='$class', rc=$rc). That is a bug in this script, not a state of the router. Starting anyway. Output was: ${out:-no output}"
;;
esac
}
# Announce/withdraw "this daemon is being replaced, not switched off". Read by
# `shaterd run` when it receives SIGTERM.
shater_mark_restart() {
mkdir -p "$(dirname "$RESTART_FLAG")" 2>/dev/null
: > "$RESTART_FLAG"
}
shater_clear_restart() { rm -f "$RESTART_FLAG"; }
# Remove the persisted boot armor, so the LAN is NOT blocked at the next boot
# before the daemon starts. Called from exactly two places, both of which are a
# statement about the PRODUCT rather than about this process: an operator typing
# `stop`, and a daemon binary that is no longer on the box. In neither case is
# anything going to come along and replace the armor with a real data plane, and a
# kill switch with nothing behind it is just a brick.
#
# NOT called on the shutdown path. That is the whole fix — see
# shater_action_disarms.
shater_disarm_boot() { rm -f "$BOOT_ARMOR"; }
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
@@ -136,9 +470,20 @@ start_service() {
# Guard: never claim to run without the daemon binary. A half-removed/failed
# shaterd upgrade must degrade to "plugin off", not to a box that thinks
# interception is live with nothing behind it.
#
# "Plugin off" now has to include DISARMING. With the boot armor in play, a
# missing binary is the one case where the fail-closed plane could stand
# forever with nothing able to replace it: the armor loads at START=21, the
# daemon never starts, and every later boot repeats it. The product being gone
# is not a security event — it is an uninstall — so the plane comes down and
# the LAN returns to plain routing, loudly.
if [ ! -x "$PROG" ]; then
shater_clear_restart
shater_disarm_boot
rm -f "$ACTIVE_FLAG"
nft delete table inet shater 2>/dev/null
_slog -p daemon.err \
"shaterd binary missing/not executable at $PROG — refusing to start (LAN stays on plain routing)"
"shaterd binary missing/not executable at $PROG — refusing to start; the fail-closed plane and its boot armor have been REMOVED (LAN back to plain routing, unprotected). Reinstall shaterd."
return 0
fi
@@ -150,23 +495,27 @@ start_service() {
# running, which is the boot case.
shater_wait_stopped
# The predecessor is gone and has already consumed the flag (it reads it in its
# SIGTERM handler). Withdraw it now, so a LATER `stop` is unambiguous even if
# this start fails further down.
shater_clear_restart
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
#
# THE FAILURE IS LOGGED, NOT SWALLOWED. This is the only place the schema
# migration runs at boot (`shaterd run`, the SIGHUP reconcile and the panel's
# config write all read UCI directly), so if it fails here it does not get
# retried until the next start. And it CAN fail for a mundane reason — a full
# /overlay makes `uci commit` fail — after which the config still carries the
# schema-v1 `dst_domain`/`dst_ip` options. The daemon holds every rule that
# still has them DISABLED and reports it, so nothing is silently misrouted, but
# rules the operator wrote are then not in force and the reason has to be
# visible somewhere. Hence: log the binary's own stderr, and start anyway —
# refusing to start would take the admin panel down with it, and the panel is
# the only way to fix the box.
local migrate_out
migrate_out=$("$PROG" migrate 2>&1) || _slog -p daemon.err \
"UCI schema migration FAILED: ${migrate_out:-no output from $PROG migrate}. Starting anyway; routing rules that still carry the removed dst_domain/dst_ip options stay DISABLED until this succeeds. Free space on /overlay and re-run '$PROG migrate', or restart the service."
# THE FAILURE IS REPORTED, NOT SWALLOWED, AND IT IS NAMED. This is the only
# place the schema migration runs at boot (`shaterd run`, the SIGHUP reconcile
# and the panel's config write all read UCI directly), so if it fails here it
# does not get retried until the next start — which is also why this, and not
# the uci-defaults call, is the report that matters: it comes back at every
# boot and every restart for as long as the problem lasts.
#
# The three outcomes are three different problems with three different fixes
# (free space / put the newer package back / the file is unreadable), and
# shater_migrate says which one it was instead of asserting the middle one at
# all of them. Start regardless in every case: refusing to start would take the
# admin panel down with it, and the panel is the only way to fix the box.
shater_migrate
procd_open_instance shater
# shaterd runs in the FOREGROUND under procd (must never daemonize). `run` is
@@ -214,6 +563,45 @@ start_service() {
}
stop_service() {
# Say WHY we are stopping before procd sends the signal, because the daemon
# cannot tell from the signal alone and the answer changes what it leaves in
# the kernel. Two INDEPENDENT questions, and the old code conflated them into
# one two-armed `case` whose else-branch answered both wrongly for `shutdown`:
#
# 1. IS A SUCCESSOR COMING (this process only)? restart / reload.
# Raise RESTART_FLAG so the outgoing daemon replaces its data plane with
# the fail-closed HOLDING plane instead of removing it. The gap until the
# successor applies is not a moment: this script waits out the old
# process, runs `shaterd migrate`, then starts a daemon that must build an
# engine — all of it, before this flag existed, with `lan -> wan ACCEPT`
# and nothing else.
#
# 2. IS THE PRODUCT BEING SWITCHED OFF (across boots)? `stop` — and only
# `stop`, and only when a PERSON is behind it (shater_stop_disarms; the
# package manager reaches us through `stop` too). Then the boot armor goes
# with it, so the next boot does not quietly reinstate what the operator
# just switched off — the same rule ACTIVE_FLAG has always enforced for
# hotplug/cron.
#
# `shutdown` answers NO to both, which is the defect this replaced: a reboot is
# not a successor and it is certainly not an operator switching the product off.
# It is the boot the armor exists for. An upgrade answers NO to the second for
# the same kind of reason.
if shater_action_handoff "$SHATER_RC_ACTION"; then
shater_mark_restart
else
shater_clear_restart
fi
local in_pkg=0 rc_en=0
shater_pkg_transaction && in_pkg=1
shater_rc_enabled && rc_en=1
if shater_stop_disarms "$SHATER_RC_ACTION" "$in_pkg" "$rc_en"; then
shater_disarm_boot
elif [ "$in_pkg" = "1" ] && shater_action_disarms "$SHATER_RC_ACTION"; then
_slog -p daemon.info \
"stop came from a package transaction that left the service enabled — keeping the boot armor, so being replaced cannot leave the next boot unprotected"
fi
# Drop the live-flag FIRST so a concurrent hotplug/cron tick cannot rebuild
# what we are about to tear down. procd then sends SIGTERM to `shaterd run`,
# which runs its OWN honest teardown (engine.Close + netplane restore) — we
@@ -233,6 +621,11 @@ reload_service() {
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
#
# Marked EXPLICITLY as well as via SHATER_RC_ACTION: this is the path a routine
# Save & Apply takes, so it is the one that must not depend on reading an
# rc.common variable correctly. Belt and braces, one line.
shater_mark_restart
stop
start
}
@@ -0,0 +1,162 @@
#!/bin/sh /etc/rc.common
# /etc/init.d/shater-armor — the fail-closed plane, before the daemon exists.
#
# WHAT THIS CLOSES
#
# /etc/init.d/shater is START=99. By then fw4 (START=19) has long since loaded
# `lan -> wan ACCEPT` and netifd (START=20) has brought the LAN bridge up, so the
# router forwards LAN traffic to the WAN in the clear from the moment the link
# comes up until `shaterd run` has been decompressed off flash, has waited out any
# predecessor, has migrated UCI, has read the config and has installed its first
# table. On router-class hardware with a UPX-packed binary that is seconds — and
# they are exactly the seconds in which Wi-Fi finishes associating and every
# client on the network reconnects and starts talking. `kill_switch=closed` was
# configured the whole time and covered none of it.
#
# There was nothing in the package that could cover it either: no /etc/nftables.d
# include, no `nft -f` in uci-defaults. Protection existed only inside a Go
# process that had not started yet.
#
# HOW
#
# The daemon persists a copy of its fail-closed HOLDING plane (the same ruleset it
# installs when the engine is down: one forward chain, LAN-to-LAN and router
# traffic accepted, everything else from the diverted devices dropped) to
# $ARMOR on every apply. This script loads it early. When the daemon comes up it
# replaces the table atomically — the ruleset begins with `delete table` and adds
# its own in one netlink transaction — so there is never a moment with no table.
#
# `iifname` matches by NAME at packet time, not by ifindex at load time, so
# loading this before netifd has created br-lan is fine: the rules simply start
# matching when the device appears. That is why START can sit here rather than
# racing netifd.
#
# START=21: after fw4 (19) and netifd (20), because fw4's own start tears its
# table down and rebuilds it and we do not want to be in the middle of that, and
# because there is nothing to protect before the LAN device is being created. The
# residual exposure is the fraction of a second between netifd's `ifup` and this
# script, against seconds-to-a-minute before.
#
# THE ESCAPE HATCHES (a kill switch that cannot be switched off is a brick)
#
# These are STATE checks, evaluated here, at the moment of arming — not a record
# of something that happened on the way down. That distinction is the whole
# lesson of v0.2.17: the arm token was deleted by an EVENT on the shutdown path
# ("this looks like a stop"), and since `reboot` also runs the K-links, the
# mechanism reliably erased itself on the one transition it was built for. An
# event on the way down cannot be trusted to describe the world on the way up; a
# question asked on the way up can be.
#
# * $ARMOR only exists while the daemon's last applied config was BOTH enabled
# and fail-closed. `globals.enabled=0` and `kill_switch=open` each remove it
# at the next apply, and an operator typing `/etc/init.d/shater stop` removes
# it there and then. Powering the box off does NOT.
# * We refuse to arm when the main service is disabled in rc.d, or when the
# daemon binary is gone — in either case nothing would ever come along to
# replace the armor with a real data plane. These two are what makes a
# genuinely uninstalled/disabled product safe REGARDLESS of what the file
# says, which is why they are checked here rather than trusted to have been
# acted on earlier.
# * We refuse to arm when UCI can be read AND says the stack is disabled. A
# config that cannot be read is NOT a refusal: that case is precisely why the
# armor is a file rather than a query.
# * The chain hooks `forward` only, so SSH, LuCI and the admin panel (all input
# hook, to the router's own addresses) stay reachable. The operator can always
# get in and undo this.
#
# Note what a bare `/etc/init.d/shater stop` does NOT mean: it does not survive a
# reboot, because S99shater is still linked and procd starts the daemon again. So
# "stopped" is not a durable off-state and this script must not be designed as if
# it were — the durable ones are `disable` (no S??shater) and `globals.enabled=0`,
# and those are the two refusals above.
#
# busybox ash only — no bashisms.
START=21 # after firewall (19) and network (20), long before shater (99)
STOP=89
ARMOR=/etc/shater/boot.nft
PROG=/usr/bin/shaterd
# Syslog line that honors globals.log_syslog, like the other two inits. An
# unreadable UCI leaves the option empty => ON, which is what we want here: the
# one boot where the config cannot be read is the boot worth logging.
_slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater-armor "$@"
}
# Is the MAIN service enabled at boot? Answered by looking for its rc.d symlink
# rather than by running `/etc/init.d/shater enabled`: that is a USE_PROCD script,
# so every action of it sources procd.sh, which takes a blocking flock — and this
# runs at START=21, in the middle of boot, for a question a glob answers exactly
# as well. The START number is not hardcoded; any S<NN>shater counts.
shater_service_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
start() {
# No saved plane => the stack has never applied an enabled, fail-closed config
# (or it was explicitly switched off). Nothing to do, and nothing to say.
[ -f "$ARMOR" ] || return 0
[ -s "$ARMOR" ] || {
_slog -p daemon.err "$ARMOR is empty — NOT arming; the LAN is unprotected until shaterd starts"
return 0
}
# Never arm something nothing can disarm.
[ -x "$PROG" ] || {
_slog -p daemon.err \
"$PROG is missing — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
shater_service_enabled || {
_slog -p daemon.warn \
"the shater service is disabled in rc.d — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
# A READABLE config that says "off" wins over the saved plane (it means the
# daemon was stopped before it could disarm). An UNREADABLE config does not:
# that is the case this whole mechanism exists for.
en=$(uci -q get shater.globals.enabled 2>/dev/null)
if [ -n "$en" ] && [ "$en" != "1" ]; then
rm -f "$ARMOR"
_slog -p daemon.info "globals.enabled=$en — boot armor removed, not arming"
return 0
fi
command -v nft >/dev/null 2>&1 || {
_slog -p daemon.err "nft is not installed — cannot arm; the LAN is unprotected until shaterd starts"
return 0
}
# Validate before loading: a truncated/incompatible snapshot must not leave a
# half-built table behind on the one boot it is needed.
if ! nft -c -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.err \
"$ARMOR did not validate (nft -c) — NOT arming; the LAN is unprotected until shaterd starts"
return 0
fi
if nft -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.warn \
"fail-closed plane armed from $ARMOR: LAN->WAN forwarding is BLOCKED until shaterd applies. SSH, LuCI and the admin panel stay reachable."
else
_slog -p daemon.err \
"could not load $ARMOR — the LAN is unprotected until shaterd starts"
fi
return 0
}
stop() {
# Deliberately a NO-OP. By the time anything stops this service the daemon owns
# `inet shater`, and deleting the table here would dismantle a LIVE data plane
# on the strength of a service that only ever ran for one second at boot. The
# disarm paths that matter live where the decision is actually made:
# /etc/init.d/shater stop (operator switched it off) and the daemon itself
# (globals.enabled=0 / kill_switch=open).
return 0
}
+295 -28
View File
@@ -3,11 +3,12 @@
# plus the data-plane watchdog (v0.2).
#
# A tiny procd-supervised loop that, once per tick, checks every enabled
# subscription and url-ruleset against its per-item `update_interval` and runs
# subscription against its per-item `update_interval` and runs
# shaterd sub update <name> (subscriptions)
# shaterd ruleset update <name> (url rulesets)
# when the item is due, then a single `shaterd reconcile` if anything changed
# (the daemon's config-hash gate rebuilds the engine only on a real change).
# `config ruleset` items are NOT touched here — see shater_run_due for who owns
# their refresh and where the remaining gap is.
#
# RELIABILITY CONTRACT (same "железно" posture as /etc/init.d/shater):
# * The loop body is fully INERT unless globals.enabled=1 AND the main shater
@@ -27,6 +28,15 @@
# the main service is STOPPED (tears interception down — fail-open, LAN
# returns to plain routing); with kill_switch=closed the rules stay
# (blocked-by-design) and we log loudly.
# * CRASH-LOOP WATCHDOG: the dead-daemon counter above cannot see the failure
# it matters most for. /etc/init.d/shater sets `respawn 3600 5 0`, so a
# daemon that dies a few seconds into startup is back within 5s and a single
# `pidof` per 60s tick nearly always finds a process — the counter resets and
# never reaches WATCHDOG_TICKS, while the fail-closed plane keeps the LAN shut
# and the panel (served BY the daemon) never comes up. So the tick's sleep is
# spent SAMPLING the daemon's identity instead of sleeping blind, and a tick in
# which several different daemons lived is counted as churn. See
# shater_churn_scan / shater_churn_verdict / shater_churn_action.
# * The loop never self-exits (procd would respawn-churn an exiting body);
# it idles on its guards instead. busybox ash only — no bashisms.
@@ -42,14 +52,34 @@ INIT_SCRIPT=/etc/init.d/shater-cron
SHATER_INIT=/etc/init.d/shater
SHATERD=/usr/bin/shaterd
ACTIVE_FLAG=/var/run/shater.active
# Raised by /etc/init.d/shater around a restart/reload and cleared by the
# successor's start_service. Read here ONLY as a "a person/package asked for this
# bounce" veto on the crash-loop verdict — never as a liveness signal.
RESTART_FLAG=/var/run/shater.restarting
# Written by `shaterd run` itself (main.go writePidfile) before it builds anything,
# and removed by that same process on a clean exit. It is the only handle that
# names THE daemon: `pidof shaterd` also matches the short-lived CLI verbs this
# very loop runs (`sub update`, `reconcile`, `schedule due`).
PIDFILE=/var/run/shaterd.pid
STAMP_DIR=/var/run/shater/cron
TICK=60 # seconds between due-checks
RETRY_SECS=300 # backoff before retrying a FAILED fetch
WATCHDOG_TICKS=5 # consecutive dead-daemon ticks before escalating
DEFAULT_SUB_INTERVAL=6h
DEFAULT_RS_INTERVAL=24h
DEFAULT_BL_INTERVAL=24h # url blocklist refresh interval (D16)
# --- crash-loop watchdog tuning --------------------------------------------
#
# Every number here is chosen against ONE question: what can a legitimate restart
# produce? A legitimate bounce (`restart`, LuCI Save & Apply -> reload, a package
# transaction) replaces the daemon EXACTLY ONCE, and it is announced twice over —
# /etc/init.d/shater raises RESTART_FLAG in stop_service and clears ACTIVE_FLAG for
# the duration. A crash loop is announced by nothing and repeats without bound.
LOOP_POLL=5 # seconds between identity samples inside one tick
LOOP_MIN_GENS=3 # distinct daemons in ONE tick that count as churn
LOOP_WINDOWS=2 # consecutive churn ticks before we call it a loop
LOOP_REPORT_TICKS=30 # do not repeat the report more often than this
# --- helpers ---------------------------------------------------------------
shater_enabled() {
@@ -132,8 +162,33 @@ shater_stamp_retry() {
# Walk anonymous `config subscription` / `config ruleset` sections by index and
# run any that are due. Echoes non-empty on stdout if at least one item updated.
#
# WHERE A SUBSCRIPTION'S BYTES TRAVEL, and what this loop sees when they cannot.
#
# `shaterd sub update <name>` honours the subscription's own `fetch_via`:
# * fetch_via != proxy — fetched DIRECT by the short-lived CLI process, exactly
# as it always was. Needs no daemon; works at cold start and first boot.
# * fetch_via == proxy — DELEGATED to the running daemon over its control
# socket, because only the daemon owns the engine the fetch has to travel
# through. Until 2026-07-27 the CLI printed one `daemon.warn` line and
# fetched DIRECT instead, so every tick of THIS loop and every fetch-at-boot
# put the feed URL and the router's real address on the plain WAN — while the
# panel's Refresh button (the same operation, through the daemon) worked, so
# it looked configured. It now FAILS instead, exit 1.
#
# CONSEQUENCE HERE, stated rather than discovered later: with a proxy-fetched
# subscription and no live daemon, the `if` below takes the else arm every time —
# one syslog line and shater_stamp_retry, i.e. a fresh attempt every RETRY_SECS
# (300s) for as long as the daemon is down. That is ~288 lines a day per such
# subscription. It is NOT the `ruleset update` situation this file used to have:
# there the work had no owner and the retry could never succeed, whereas here the
# retry succeeds within RETRY_SECS of the daemon coming back, and the loop only
# runs at all while `shater_enabled && shater_active` — a dead daemon is already
# being escalated by shater_watchdog, and with kill_switch=open it stops the stack
# (clearing ACTIVE_FLAG), which makes this loop inert. Deliberately left noisy:
# a subscription that is silently going stale is worse than a repeated line.
shater_run_due() {
local i name en ivl secs stamp src changed=""
local i name en ivl secs stamp changed=""
# Subscriptions.
i=0
@@ -159,28 +214,36 @@ shater_run_due() {
i=$(( i + 1 ))
done
# Rulesets (only url sources auto-update; others have nothing to fetch).
i=0
while uci -q get "shater.@ruleset[$i]" >/dev/null 2>&1; do
name=$(uci -q get "shater.@ruleset[$i].name")
src=$(uci -q get "shater.@ruleset[$i].source")
if [ -n "$name" ] && [ "$src" = "url" ]; then
ivl=$(uci -q get "shater.@ruleset[$i].update_interval")
secs=$(shater_ivl_secs "$ivl" "$DEFAULT_RS_INTERVAL")
stamp="$STAMP_DIR/rs.$(shater_safe_name "$name")"
if shater_due "$stamp" "$secs"; then
if "$SHATERD" ruleset update "$name" >/dev/null 2>&1; then
shater_stamp "$stamp"
changed=1
else
_slog -p daemon.warn \
"ruleset update '$name' failed; retrying in ${RETRY_SECS}s"
shater_stamp_retry "$stamp" "$secs"
fi
fi
fi
i=$(( i + 1 ))
done
# RULE-SETS ARE NOT UPDATED FROM HERE, AND NEVER WERE.
#
# There used to be a second loop that ran `shaterd ruleset update <name>` for
# every `config ruleset` with source=url. That verb has never existed: it
# printed a note and exited 0, so this loop stamped the item as freshly updated
# and raised `changed`, which cost a reconcile per item per interval and told
# the operator the list was current when not one byte had been fetched. The verb
# now exits non-zero (shater/cmd/shaterd/main.go, notImpl), which turns the same
# loop into one failed attempt and one syslog line every RETRY_SECS — ~288 lines
# a day, per rule-set, about work that has no owner here. Noise in the log hides
# real problems as effectively as a lie about success does.
#
# WHO REFRESHES A url RULE-SET NOW, so the next reader does not think this was
# forgotten. `source=url` splits into two shapes in shater/generate/ruleset.go:
#
# * the URL serves an engine-native .srs/.json -> it stays a REMOTE rule-set
# and sing-box owns fetch/cache/refresh through RemoteRuleSet.UpdateInterval
# on the running box. This cron loop never had anything to contribute.
# * the URL serves a plain-text list -> it is compiled locally into
# /etc/shater/lists/<tag>.srs, and that artifact is refreshed by the
# GENERATOR, "when missing or older than update_interval" — i.e. only when
# something else already caused a generate. Nothing schedules one, so this
# shape has NO periodic refresh at all today. That is a real gap, and it is
# stated here rather than papered over with a call to a verb that does
# nothing: closing it needs a daemon-side timer (or a real `ruleset update`),
# not a shell loop, because only the daemon can force a rebuild past the
# config-hash gate.
#
# `config blocklist` url items are a different mechanism and DO refresh — see
# shater_run_due_blocklists below.
[ -n "$changed" ] && echo 1
}
@@ -256,6 +319,171 @@ shater_watchdog() {
echo "$dead"
}
# --- crash-loop watchdog ----------------------------------------------------
#
# THE HOLE. shater_watchdog above answers "is a daemon there?" once every TICK
# seconds. /etc/init.d/shater sets `respawn 3600 5 0`, so a daemon that dies a few
# seconds into startup is back 5s later and that one sample nearly always finds a
# process: the dead-counter resets, never reaches WATCHDOG_TICKS, and the single
# failure the watchdog exists for — a new binary or a bad config that cannot get
# an engine up while the fail-closed plane holds the LAN shut — is the one it can
# never see. The daemon also serves the panel, so in that state the operator has
# neither internet nor a way to look at the box.
#
# THE SIGNAL. Not "is it there" but "is it the SAME one". The tick's sleep is
# spent taking an identity sample every LOOP_POLL seconds instead of sleeping
# blind, and a tick in which LOOP_MIN_GENS different daemons lived is a churn
# tick. LOOP_WINDOWS consecutive churn ticks is the verdict.
#
# WHY A LEGITIMATE RESTART CANNOT REACH IT. Four independent reasons, in order of
# how much they are relied on:
#
# 1. A bounce replaces the daemon ONCE. One restart scores 2 generations in the
# tick it happens in and 1 in every tick after, so it cannot even produce a
# single churn tick at LOOP_MIN_GENS=3, let alone LOOP_WINDOWS of them in a
# row. Reaching the verdict takes >= 4 replacements inside 2 consecutive
# minutes, >= 2 in each.
# 2. Bounces are ANNOUNCED. /etc/init.d/shater raises RESTART_FLAG in
# stop_service and clears ACTIVE_FLAG for the whole stop->start, and either
# one seen in any sample of a tick discards that tick outright.
# 3. The panel's Apply does not restart anything: it writes UCI and applies over
# the daemon's control socket, in process. Only `restart`, a LuCI Save &
# Apply (config.change -> reload) and a package transaction bounce the
# daemon, and a human cannot produce those at four a minute.
# 4. The sample names THE daemon via its pidfile, not `pidof shaterd` — the
# short-lived CLI verbs this very loop runs share that process name.
#
# WHAT IT DOES NOT COVER, stated rather than implied: a daemon that dies
# INSTANTLY (well under a second) is almost never caught alive by a 5s sample, so
# it scores few generations and this detector stays quiet. That case is exactly
# the one the existing dead-tick counter does see — its `pidof` misses too, tick
# after tick — so the two cover opposite ends and are deliberately left as two
# independent instruments rather than merged into one clever number.
# One identity sample: echoes the pid of the live `shaterd run`, or "-" for none.
#
# Through the PIDFILE, which `shaterd run` writes before it builds anything and
# removes on a clean exit, because that is the only handle that names THE daemon:
# `pidof shaterd` also matches `shaterd sub update` / `reconcile` / `schedule due`.
# /proc/<pid>/comm is checked so a stale pidfile whose pid has been reused by an
# unrelated process cannot read as a live daemon. No forks: `read` is a builtin.
shater_sample_pid() {
local pid="" comm=""
[ -r "$PIDFILE" ] && read -r pid 2>/dev/null < "$PIDFILE"
case "$pid" in
''|*[!0-9]*) echo -; return ;;
esac
[ -r "/proc/$pid/comm" ] && read -r comm 2>/dev/null < "/proc/$pid/comm"
[ "$comm" = "shaterd" ] || { echo -; return; }
echo "$pid"
}
# shater_churn_scan <sample>... -> "<generations> <absent-samples>"
#
# A GENERATION is one distinct daemon lifetime observed during the tick: a live
# pid that differs from the last live pid seen. A daemon that simply keeps running
# therefore scores exactly 1 generation and 0 absent samples for as long as it
# runs — the signal is flat unless something is actually being replaced.
#
# A GAP (samples with no daemon at all, e.g. procd's 5s respawn hole) is counted
# but does NOT by itself open a new generation: only a different pid does. An
# earlier draft reset the comparison across a gap so that "same pid seen again
# after a gap" would score two. That case cannot occur — a respawn always gets a
# fresh pid — so it was unfalsifiable code, and resetting also meant a momentarily
# unreadable pidfile could inflate the count. Not resetting is both simpler and
# the safer direction.
#
# Pure: no I/O, no globals, every input on the command line. That is what lets the
# gate drive it with synthetic sample streams instead of a live router.
shater_churn_scan() {
local gens=0 absent=0 last="" s
for s in "$@"; do
if [ "$s" = "-" ]; then
absent=$(( absent + 1 ))
continue
fi
[ "$s" = "$last" ] || gens=$(( gens + 1 ))
last="$s"
done
echo "$gens $absent"
}
# shater_churn_verdict <gens> <samples> <announced> <churn-so-far>
# -> the new consecutive-churn-tick count
#
# Also pure. `announced`=1 means a sample during the tick saw RESTART_FLAG up or
# ACTIVE_FLAG down, i.e. /etc/init.d/shater said out loud that it was bouncing the
# daemon: that tick proves nothing and resets the run. A tick with no samples at
# all (the first pass through the loop) likewise scores 0 rather than guessing.
shater_churn_verdict() {
local gens="$1" n="$2" announced="$3" churn="$4"
[ "$announced" = "1" ] && { echo 0; return; }
[ "$n" -gt 0 ] || { echo 0; return; }
if [ "$gens" -ge "$LOOP_MIN_GENS" ]; then
echo $(( churn + 1 ))
return
fi
echo 0
}
# shater_churn_action <churn-ticks> <kill_switch> -> none | log | stop
#
# WHAT TO DO, and why it is not our call to make twice. A crash loop leaves the
# box in the same state a dead daemon does — no engine, fail-closed plane standing
# — so the answer is the one the operator already gave with kill_switch, not a new
# policy invented here:
#
# open The operator asked for connectivity over interception. Stop the stack,
# exactly as shater_watchdog does for a sustained-dead daemon: the plane
# comes down and the LAN returns to plain routing. It also disarms the
# boot armor, so the NEXT boot is clean too instead of repeating the loop
# behind a closed LAN. Nothing else can end the loop: procd's retries are
# infinite by design.
# closed The operator asked for blocked-over-leaking. Blocked is what they get,
# and opening their LAN from a background loop would be the opposite of
# what the knob says. Report it loudly and let the person decide; the
# message names the one command that opens it.
#
# The list is POSITIVE and CLOSED, and the fall-through goes to `log`: an absent
# or unrecognised kill_switch is the model's documented default ("closed", see
# shater/model/model.go DefaultGlobals), and `log` is the recoverable side — it
# changes nothing and can be acted on, where a wrong `stop` silently drops a
# household onto the unproxied WAN.
#
# NOTE (not changed here, deliberately): shater_watchdog above answers the same
# question with `if closed ... else stop`, so for an ABSENT kill_switch it fails
# open — the opposite of the documented default. It is left alone because that
# behaviour predates this file's crash-loop work; it is reported upward instead.
shater_churn_action() {
local churn="$1" ks="$2"
[ "$churn" -ge "$LOOP_WINDOWS" ] || { echo none; return; }
case "$ks" in
open) echo stop ;;
closed) echo log ;;
*) echo log ;;
esac
}
# Sleep out one tick in LOOP_POLL slices, sampling the daemon's identity as we go.
# Publishes CHURN_SAMPLES / CHURN_N / CHURN_ANNOUNCED for the next pass of loop().
# Deliberately NOT a subshell (globals must survive), and it always returns 0 so a
# false `[ -f ]` at the end cannot look like a failure.
shater_tick_sample() {
local slept=0
CHURN_SAMPLES=""
CHURN_N=0
CHURN_ANNOUNCED=0
while [ "$slept" -lt "$TICK" ]; do
sleep "$LOOP_POLL"
slept=$(( slept + LOOP_POLL ))
CHURN_SAMPLES="$CHURN_SAMPLES $(shater_sample_pid)"
CHURN_N=$(( CHURN_N + 1 ))
[ -f "$RESTART_FLAG" ] && CHURN_ANNOUNCED=1
[ -f "$ACTIVE_FLAG" ] || CHURN_ANNOUNCED=1
done
return 0
}
# loop: the foreground body supervised by procd. Never exits on its own — it
# idles while disabled/inactive so procd is not respawn-churned by a
# self-exiting body when the stack is off.
@@ -274,7 +502,12 @@ loop() {
# the flock immediately and keeps children (sleep/shaterd) from inheriting
# it. A no-op where fd 1000 is not open (older procd.sh without procd_lock).
exec 1000>&-
local changed dead=0
local changed dead=0 churn=0 quiet=0 scan gens absent ks act
# No tick has been sampled yet on the first pass; shater_churn_verdict scores
# an empty tick as 0 rather than guessing.
CHURN_SAMPLES=""
CHURN_N=0
CHURN_ANNOUNCED=0
mkdir -p "$STAMP_DIR"
while :; do
if shater_enabled && shater_active; then
@@ -299,10 +532,44 @@ loop() {
"$SHATERD" schedule due >/dev/null 2>&1
fi
dead=$(shater_watchdog "$dead")
# Crash-loop verdict on the tick that has just elapsed. Unquoted on
# purpose: CHURN_SAMPLES is a whitespace-separated token list and word
# splitting is how it becomes arguments.
scan=$(shater_churn_scan $CHURN_SAMPLES)
gens=${scan%% *}
absent=${scan##* }
churn=$(shater_churn_verdict "$gens" "$CHURN_N" "$CHURN_ANNOUNCED" "$churn")
ks=$(uci -q get shater.globals.kill_switch)
act=$(shater_churn_action "$churn" "$ks")
case "$act" in
stop)
_slog -p daemon.crit \
"shaterd is CRASH-LOOPING: $gens distinct daemons in the last ${TICK}s (absent in $absent of $CHURN_N samples), $churn such windows in a row — it is being respawned faster than it can bring an engine up. kill_switch=open, so shater is being STOPPED: interception comes down and the LAN returns to plain, UNPROXIED routing. Find the reason with 'logread -e shaterd', then '/etc/init.d/shater start'."
"$SHATER_INIT" stop
churn=0
quiet="$LOOP_REPORT_TICKS"
;;
log)
# Rate-limited: a standing condition, not an event. Never
# silent for good, though — an operator who looks at the log an
# hour later must still find it being said.
if [ "$quiet" -le 0 ]; then
_slog -p daemon.crit \
"shaterd is CRASH-LOOPING: $gens distinct daemons in the last ${TICK}s (absent in $absent of $CHURN_N samples), $churn such windows in a row — it is being respawned faster than it can bring an engine up. kill_switch=${ks:-closed} keeps the fail-closed plane standing, so the LAN stays blocked and the admin panel is down with the daemon that serves it. Nothing is decided for you: find the reason with 'logread -e shaterd', or open the LAN with '/etc/init.d/shater stop'."
quiet="$LOOP_REPORT_TICKS"
fi
churn=0
;;
esac
[ "$quiet" -gt 0 ] && quiet=$(( quiet - 1 ))
else
dead=0
churn=0
quiet=0
fi
sleep "$TICK"
# Sleeps out the tick, sampling the daemon's identity while it does.
shater_tick_sample
done
}
@@ -36,31 +36,344 @@ mkdir -p /etc/shater
# transaction, and any /etc/init.d/* invocation in that window risks blocking
# the transaction on rc.common's per-service flock (procd_lock).
# Seed the built-in preset packs (disabled) so the LuCI Rules page renders their
# toggles. Idempotent: only creates a section that does not yet exist.
seed_preset() {
local sid="$1" name="$2" s n
uci -q get "shater.$sid" >/dev/null 2>&1 && return 0
# A pack section may already exist under a DIFFERENT section id (created by
# the LuCI seeding or an older release) — match by pack name, not just id,
# or we would duplicate the toggle.
for s in $(uci -q show shater 2>/dev/null | sed -n "s/^shater\.\([^.=]*\)=preset$/\1/p"); do
n=$(uci -q get "shater.$s.name")
[ "$n" = "$name" ] && return 0
# NO preset packs are seeded here, and the ones older releases seeded are removed.
#
# Until now this script created three `config preset` sections (block_ads,
# ru_bypass, private; all `enabled=0`) "so the LuCI Rules page renders their
# toggles". Both halves of that stopped being true in v0.2, and what was left was
# a knob wired to nothing:
#
# * `preset` IS NOT A SECTION TYPE. The type switch in model.ParseUCIExport
# (shater/model/uci.go) has no `preset` branch, and an unknown section type is
# dropped on the floor rather than rejected — TestUnknownSectionAndOptionIgnored
# pins that a config carrying one still parses, because the daemon has to come
# up on whatever it finds. So `uci set shater.block_ads.enabled=1; uci commit`
# edited the file and changed NOTHING about the running router, and there was
# no error anywhere to say so. Nothing in shater/, panel/ or luci-app-shater
# reads the type either.
# * v0.2's luci-app-shater is a launcher for the daemon's own admin panel. There
# is no LuCI Rules page for the toggles to appear on.
# * The sections did not even persist. model.writeUCIWith replaces the WHOLE
# package (`uci delete shater` + `uci import`), so the first save from the
# panel deleted all three. A placeholder that erases itself reads as "this
# broke", not as "this was never here" — which is worse than its absence.
#
# So the packs are not "missing": nothing in v0.2 lost a feature when the sections
# stopped being written, because nothing ever read them. And re-adding the seed is
# not how presets would come back. v0.1's packs were xray `geosite:`/`geoip:`
# matcher lists materialised into synthetic rules (`xrayctl/preset.go` on the v0.1
# branch); in v0.2 a rule's destination IS a `config ruleset` (schema v2), so the
# same pack is an ordinary ruleset + rule — which the panel's Routing page already
# builds, geosite/geoip sources included. Anything richer needs a section type the
# parser knows about, which has to land in shater/model FIRST; seeding UCI ahead
# of the parser only produces silence.
#
# The purge is narrow and safe by construction: it matches on the section TYPE
# being exactly `preset`, and that type is read by no consumer, so there is no
# setting to lose. Bounded and re-querying each round because a `config preset`
# may also be ANONYMOUS (`shater.@preset[0]`), where deleting from a list captured
# up front would shift the remaining indices out from under it.
purge_presets() {
local s n=0 changed=""
[ -f /etc/config/shater ] || return 0
while [ "$n" -lt 32 ]; do
s=$(uci -q show shater 2>/dev/null |
sed -n 's/^shater\.\([^.=]*\)=preset$/\1/p' | head -n 1)
[ -n "$s" ] || break
uci -q delete "shater.$s" || break
changed=1
n=$((n + 1))
done
uci set "shater.$sid=preset"
uci set "shater.$sid.name=$name"
uci set "shater.$sid.enabled=0"
[ -n "$changed" ] && uci -q commit shater
return 0
}
if uci -q get shater.globals >/dev/null 2>&1 || [ -f /etc/config/shater ]; then
seed_preset block_ads block-ads
seed_preset ru_bypass ru-bypass
seed_preset private private
uci -q commit shater
fi
purge_presets
# Introduce the daemon-created `shater-l3*` TUN to fw4 (L3 ingress, D-L3). The
# daemon policy-routes LAN ICMP into that device from OUR nft table
# `inet shater`, but nftables runs EVERY table on every packet and a drop in
# any one of them wins — an accept in `inet shater` cannot override fw4. And
# fw4 WILL drop this forward: netifd knows nothing about a device the daemon
# creates at runtime, so it belongs to no zone and falls into fw4's zone-less
# defaults (REJECT). The device has to be declared to fw4 itself; it cannot be
# fixed from our own table.
#
# Seeded UNCONDITIONALLY (not gated on globals.l3_tunnel): uci-defaults run
# once, so gating on the option would require re-running this script when the
# option is flipped later — which never happens. An idle zone is harmless: its
# device match is a plain iifname/oifname STRING compare that simply never hits
# while the TUN does not exist.
#
# Idempotency: `config zone`/`config forwarding` are normally ANONYMOUS
# sections, and a naive `uci add firewall zone` would append a duplicate on
# every re-run (uci-defaults re-run on package upgrade/reinstall). The zone is
# NAMED instead, guarded by an existence check — a re-run re-finds the section
# and touches nothing. The forwardings are named too where the name is free, but
# their guard is a scan of the actual src/dest pairs, which is stronger; see
# seed_l3_forwarding below.
seed_l3_zone() {
# No fw4 on this image (bare nftables build) => nothing drops the forward
# on fw4's behalf and there is nothing to punch through.
[ -f /etc/config/firewall ] || return 0
if ! uci -q get firewall.shater_l3 >/dev/null; then
uci set firewall.shater_l3=zone
uci set firewall.shater_l3.name='shater_l3'
uci set firewall.shater_l3.input='REJECT'
uci set firewall.shater_l3.output='ACCEPT'
uci set firewall.shater_l3.forward='REJECT'
uci set firewall.shater_l3.masq='0'
# INERT TODAY, kept for the day it is not. mtu_fix clamps forwarded TCP
# MSS to the route MTU — but the L3 TUN is 65535 (deliberately: at any
# smaller value the kernel fragments into the device, and the flow
# dispatcher refuses to judge a fragment and lets the stack forge the
# echo reply — see l3MTU in shater/generate/inbound.go), so the clamp has
# nothing to clamp to. And only ICMP is ever marked into this device, so
# no TCP rides here to be clamped in the first place. It earns its keep
# the moment either of those changes; removing it would make that day
# silent.
uci set firewall.shater_l3.mtu_fix='1'
# `list device`, deliberately NOT the usual `list network`: fw4
# resolves a zone's networks through netifd, and netifd never learns
# about a device the daemon creates at runtime — a stub interface
# (proto none) would need to be brought UP to contribute an l3_device,
# and nothing ever brings it up, so `list network` resolves to an
# EMPTY device set and fw4 keeps dropping the forward. `list device`
# instead compiles to an iifname/oifname STRING match, valid before
# the TUN exists and matching from the moment shaterd creates it —
# no netifd involvement and no firewall reload at enable time. Do not
# "normalize" this to `list network` in a refactor; it breaks silently.
#
# The WILDCARD is load-bearing too. The daemon no longer opens one fixed
# device: it alternates between `shater-l3a` and `shater-l3b` so that a
# new engine generation never has to reopen the name the previous one is
# still holding (that collision — TUNSETIFF: device or resource busy —
# took the whole LAN down on the production router, because the recovery
# path rebuilt the same config and hit the same busy name). fw4 compiles
# `shater-l3*` to `iifname "shater-l3*"` / `oifname "shater-l3*"`,
# verified on ImmortalWrt 25.12.1 with nftables 1.1.6, so ONE zone covers
# every slot and no firewall reload is needed when the slot changes.
uci add_list firewall.shater_l3.device='shater-l3*'
fi
seed_l3_forwardings
uci -q commit firewall
}
# EVERY zone gets a forwarding into shater_l3, not just `lan`.
#
# The bug this closes is silent by construction. The daemon's divert set is built
# from every enabled `config inbound`'s network PLUS every device a rule names
# through an `iface:`/`zone:` source (shater/netplane/nft.go, nftDivertRefs) — so
# on a router with several LAN zones, ICMP from ALL of them is marked and routed
# into the TUN by our table. Our table then accepts it and fw4 drops it anyway,
# because the forward is judged in `forward_<source zone>` and only `lan` had a
# jump to `accept_to_shater_l3`. Result: ping through the tunnel works from one
# subnet and not from the next, with nothing in any log to say why — fw4's drop
# is the zone's policy verdict, not a rule with a name. The owner's production
# router has a single LAN zone, which is exactly why this went unnoticed; his
# second router has three.
#
# Every zone, including an uplink zone, and that is deliberate rather than lazy:
#
# - The alternative is guessing which zones hold clients, and every available
# signal is wrong somewhere. `masq='1'` marks the WAN on a stock config and
# also marks a double-NAT LAN. The name `wan*` is a convention, not a rule.
# A guess that is wrong reintroduces exactly the silent breakage above, while
# a superfluous entry costs a line of ruleset.
# - A forwarding into shater_l3 permits nothing on its own. It authorises the
# forward of packets ROUTED INTO the TUN, and the only thing that routes a
# packet there is our own fwmark rule, which matches solely on the divert
# device set. A packet arriving on the WAN is not marked and never reaches
# this decision; if an operator ever puts a WAN device in the divert set,
# they meant to and this is the entry that makes it work.
# - The reverse direction is NOT opened: no `src shater_l3` forwarding exists,
# so nothing comes out of the TUN into a zone by way of these sections. The
# engine's own replies return on the conntrack `established,related accept`
# at the top of fw4's forward chain.
#
# LIMIT, stated because it is not obvious: this is a SNAPSHOT. uci-defaults run
# at first boot and on package install/upgrade, so a zone created AFTER the last
# shater-core install has no forwarding until the next one. Re-running this
# script (or reinstalling the package) re-seeds. The durable fix belongs in the
# daemon, which recomputes the divert set on every apply and already knows which
# zones are in it; it is deliberately not attempted from here.
seed_l3_forwardings() {
uci -q show firewall 2>/dev/null |
sed -n "s/^firewall\.\([^.=]*\)=zone\$/\1/p" |
while read -r sid; do
zone=$(uci -q get "firewall.$sid.name")
# Unnamed zone: fw4 cannot reference it from a forwarding either.
[ -n "$zone" ] || continue
# Our own zone: a forwarding from shater_l3 to itself is meaningless.
[ "$zone" = "shater_l3" ] && continue
seed_l3_forwarding "$zone"
done
}
# One `config forwarding` <zone> -> shater_l3, created only if no such forwarding
# exists yet.
#
# The guard scans the ACTUAL src/dest pairs rather than trusting a section id,
# which covers all three ways one can already be there: the legacy named section
# `shater_l3_fwd` seeded by earlier releases (src=lan), the per-zone names this
# function writes, and an anonymous one an operator added by hand. Without that,
# a re-run — uci-defaults re-run on every package upgrade — would append a
# duplicate for `lan` on every upgrade.
seed_l3_forwarding() {
local zone="$1" sid found
found=$(uci -q show firewall 2>/dev/null |
sed -n "s/^firewall\.\([^.=]*\)=forwarding\$/\1/p" |
while read -r f; do
[ "$(uci -q get "firewall.$f.dest")" = "shater_l3" ] || continue
[ "$(uci -q get "firewall.$f.src")" = "$zone" ] || continue
echo yes
break
done)
[ -n "$found" ] && return 0
# Section ids are [a-zA-Z0-9_] only, while a zone name may legally carry a
# hyphen — sanitise, and keep the legacy id for `lan` so an existing install
# is recognised as already seeded rather than gaining a second section.
if [ "$zone" = "lan" ]; then
sid="shater_l3_fwd"
else
sid="shater_l3_fwd_$(printf '%s' "$zone" | sed 's/[^a-zA-Z0-9_]/_/g')"
fi
# The id may still be taken — by a section for a DIFFERENT zone whose name
# sanitises to the same thing, or by something else entirely. Fall back to an
# anonymous section rather than overwrite: the src/dest scan above is what
# makes this idempotent, the name is only there to be readable.
if uci -q get "firewall.$sid" >/dev/null; then
sid=$(uci add firewall forwarding) || return 0
else
uci set "firewall.$sid=forwarding"
fi
uci set "firewall.$sid.src=$zone"
uci set "firewall.$sid.dest=shater_l3"
}
seed_l3_zone
# Upgrade path for routers seeded by a pre-slot build.
#
# The block above only runs when the zone does NOT exist, which is exactly right
# for idempotency and exactly wrong here: an already-installed router has the
# zone with the OLD exact device `shater-l3`, that name matches no slot, and fw4
# would go back to dropping the forward — i.e. LAN ping through the tunnel dies
# silently on upgrade while everything reports healthy. Rewrite it in place.
#
# Narrow on purpose: only the literal legacy entry is replaced, and only when the
# wildcard is not already listed, so an operator who added devices of their own
# keeps them and a re-run changes nothing (uci-defaults re-run on every package
# upgrade). No `fw4 reload` here — uci-defaults run before the firewall starts on
# boot, and on a package upgrade the daemon's next apply is what needs the zone,
# not this script.
migrate_l3_zone_wildcard() {
[ -f /etc/config/firewall ] || return 0
uci -q get firewall.shater_l3 >/dev/null || return 0
devs=$(uci -q get firewall.shater_l3.device) || return 0
case " $devs " in
*" shater-l3* "*) return 0 ;; # already migrated
*" shater-l3 "*) ;; # legacy exact name present
*) return 0 ;;
esac
uci -q del_list firewall.shater_l3.device='shater-l3'
uci add_list firewall.shater_l3.device='shater-l3*'
uci -q commit firewall
}
migrate_l3_zone_wildcard
# Bring the UCI schema forward on upgrade (idempotent; refuses a newer schema).
[ -x /usr/bin/shaterd ] && /usr/bin/shaterd migrate >/dev/null 2>&1
#
# THE RESULT IS NO LONGER THROWN AWAY. `>/dev/null 2>&1` discarded stdout, stderr
# AND the exit status, so a refusal here was indistinguishable from success — on
# the one screen the operator who caused it is actually looking at.
#
# WHY THIS STILL `exit 0`s (the script ends with one, and this block does not
# change that). A uci-defaults script that exits non-zero is NOT deleted and runs
# again at every boot. That is the wrong trade here, three times over:
#
# * This file does far more than migrate — it seeds rt_tables, the shater_l3
# fw4 zone and its per-zone forwardings, applies sysctl, and launches a
# DETACHED BRING-UP that enables and RESTARTS shater/shater-cron and reloads
# the firewall. Re-running all of that at every boot to carry one bit of "the
# migration failed" would bounce the tunnel on every boot, after S99 had
# already started it. One recoverable failure would become permanent churn.
# * The exit status is not a reporting channel: nothing reads a uci-defaults
# script's status, and neither procd nor the package manager surfaces it. It
# buys no diagnosis, only the re-run.
# * The retry it would buy already exists, and is better. /etc/init.d/shater
# runs the same migration on EVERY start with the full classified report, so
# a failure is retried at every boot and every restart regardless. And
# re-running cannot fix either real cause anyway: a downgrade is fixed by
# installing the right package, a full /overlay by freeing space.
#
# So: capture the status, name the cause, record it durably, and exit 0. The
# script has done everything it can do, and the failure is not lost.
#
# The classification is the same closed positive list /etc/init.d/shater uses
# (shater_migrate_class there, with the long argument for why the list is closed
# and why `downgrade` is matched on the binary's own words). It is repeated here
# rather than shared because the package installs no shell library the two could
# both source; TestMigrateClassifiersAgree (shater/cmd/shaterd) runs both over
# the same inputs and fails if they ever disagree.
SHATER_MIGRATE_BREADCRUMB=/etc/shater/migrate-failed
shater_migrate_class() {
local rc="$1" out="$2"
[ "$rc" = "0" ] && { echo ok; return 0; }
case "$out" in
*"newer than this build"*) echo downgrade; return 0 ;;
esac
uci -q export shater >/dev/null 2>&1 || { echo unreadable; return 0; }
echo failed
}
# stderr, unconditionally: inside `apk add` / `opkg install` that is the package
# manager's own output, i.e. the installing operator's screen. Plus the durable
# breadcrumb on flash, which /etc/init.d/shater removes on the first successful
# migration. Deliberately NOT syslog — globals.log_syslog may be off by the
# operator's choice, and this path honours it by using other channels instead.
shater_migrate_shout() {
echo "shater: $*" >&2
mkdir -p "$(dirname "$SHATER_MIGRATE_BREADCRUMB")" 2>/dev/null
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') $*" \
> "$SHATER_MIGRATE_BREADCRUMB" 2>/dev/null || :
}
run_migrate() {
local out rc class schema
[ -x /usr/bin/shaterd ] || return 0
out=$(/usr/bin/shaterd migrate 2>&1)
rc=$?
class=$(shater_migrate_class "$rc" "$out")
schema=$(uci -q get shater.globals.schema_version)
case "$schema" in ""|*[!0-9]*) schema=0 ;; esac
case "$class" in
ok)
rm -f "$SHATER_MIGRATE_BREADCRUMB"
;;
downgrade)
shater_migrate_shout "install: UCI schema migration REFUSED — /etc/config/shater is schema v$schema and the shater build being installed understands an older one, so this is a DOWNGRADE. Nothing was migrated and nothing on disk was changed: your settings are intact, and also unchangeable, because the daemon and the panel refuse every config write for the same reason. Install a build that understands schema v$schema again (docs-shater/INSTALL.md has the pinned per-version feed). 'shaterd migrate' said: ${out:-no output}"
;;
unreadable)
shater_migrate_shout "install: UCI schema migration FAILED and /etc/config/shater CANNOT BE READ ('uci export shater' fails), so the schema on disk cannot even be named. Check the file by hand before configuring anything. 'shaterd migrate' said: ${out:-no output}"
;;
failed)
shater_migrate_shout "install: UCI schema migration FAILED for a reason this script does not recognise; the config is still at schema v$schema. The mundane cause is a full /overlay ('uci commit' cannot write) — check 'df /overlay'. /etc/init.d/shater retries this on every start and reports it there too. 'shaterd migrate' said: ${out:-no output}"
;;
*)
# Unreachable: shater_migrate_class returns a closed set, all of it
# handled above. Named rather than swept up, so a script that disagrees
# with itself says so instead of picking a confident branch and being
# wrong quietly.
shater_migrate_shout "install: INTERNAL — 'shaterd migrate' produced a result this script cannot classify (class='$class', rc=$rc). That is a bug in /etc/uci-defaults/30_shater-core. Output was: ${out:-no output}"
;;
esac
return 0
}
run_migrate
# Apply our sysctl knobs NOW (boot applies them via procd's sysctl service, but
# on a live opkg/apk install nothing else re-reads sysctl.d — without this, an
@@ -110,8 +423,27 @@ SHATER_BRINGUP='
done
[ -x /etc/init.d/shater ] && /etc/init.d/shater enable
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron enable
# The boot-time fail-closed armor. `enable` only — it is a one-shot that loads
# the persisted holding plane at START=21, and running it NOW would install a
# block on a live box moments before the daemon replaces it anyway. It has to
# be enabled here regardless of whether the stack is on: the file it loads only
# exists while the daemon wants it to, so an enabled-but-unarmed service is a
# no-op, and enabling it later would mean the first boot after an upgrade is
# the one boot still exposed.
[ -x /etc/init.d/shater-armor ] && /etc/init.d/shater-armor enable
[ -x /etc/init.d/shater ] && /etc/init.d/shater restart
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron restart
# Fold the seeded shater_l3 zone into the LIVE ruleset — matters on a live
# opkg/apk install only, where firewall started long before our commit and
# nothing else would re-read it until the next reboot. Gated on the fw4
# table actually being loaded: at FIRST boot this job can run before the
# S19 firewall start, and an early reload would install a ruleset built
# from a half-initialized netifd AND make the later start a no-op (fw4
# start skips when its table already exists). No table => the pending S19
# start reads the committed config by itself, no reload needed.
if nft list tables 2>/dev/null | grep -q "inet fw4"; then
[ -x /etc/init.d/firewall ] && /etc/init.d/firewall reload
fi
exit 0
'
SHATER_TMO=""
@@ -0,0 +1,80 @@
# /lib/upgrade/keep.d/shater-core — what sysupgrade and LuCI "Backup" must carry
# out of /etc/shater.
#
# HOW THIS FILE IS READ. /sbin/sysupgrade (base-files, list_static_conffiles):
#
# find $(sed -ne '/^[[:space:]]*$/d; /^#/d; p' \
# /etc/sysupgrade.conf /lib/upgrade/keep.d/* 2>/dev/null) \
# \( -type f -o -type l \) $filter 2>/dev/null
#
# so blank lines and lines starting with '#' are stripped, and every other line is
# a path handed to `find`: a directory is recursed, a path that does not exist is
# silently skipped (hence a trailing '/' for the two directories, and no need to
# guard for a fresh install that has neither). The result is tarred and, on a real
# sysupgrade, HELD IN RAM across the flash — which is why this is a per-file
# decision and not simply "/etc/shater/".
#
# WHY IT EXISTS. Everything the product knows besides /etc/config/shater lives in
# /etc/shater, and nothing shipped a keep.d entry for it. A "keep settings"
# sysupgrade, or a LuCI backup restored onto a new router, therefore produced a
# box whose config looked complete and whose node inventory was EMPTY — silently.
#
# /etc/config/shater is NOT listed here: it is declared in
# Package/shater-core/conffiles, and sysupgrade backs CHANGED conffiles up on its
# own (list_changed_conffiles). Listing it again would work, but it would claim
# ownership of a mechanism that already covers it.
# THE NODE INVENTORY. Subscription-fetched nodes deliberately live OUTSIDE UCI
# (shater/model/subcache.go) — one JSON file per subscription. Without them the
# restored box has groups and rules that reference nodes which do not exist, so no
# tunnel comes up, and the only repair is `sub update`, which needs the internet
# the tunnel was supposed to be providing. Indented JSON: a few hundred KiB even
# for a several-hundred-node subscription.
/etc/shater/subs/
# THE BOOT-ARMOR ARM TOKEN. Its PRESENCE is what lets /etc/init.d/shater-armor
# (START=21) load the fail-closed plane before fw4's `lan -> wan ACCEPT` is the
# only rule on the box. Without it the first boot after a restore forwards LAN to
# WAN in the clear until the daemon has built an engine. One small nft script.
/etc/shater/boot.nft
# COMPILED LIST ARTIFACTS (.srs). Losing these fails SILENTLY in the worst
# direction: a missing LOCAL rule-set is left out of the generated config and the
# engine starts perfectly happily with the filtering simply gone
# (shater/generate/ruleset.go, compiledListRuleSet). "It will re-download itself"
# is NOT true for them either — a compiled url list is rebuilt only by the next
# generate, and nothing schedules one (see the note in /etc/init.d/shater-cron
# about `ruleset update`). Cheap to keep: compiled .srs is 3-6% of the source
# text (~80 KiB for a 150k-domain list), under a 4 MiB soft cap.
/etc/shater/lists/
# ALERT DE-DUPLICATION STATE. A few hundred bytes mapping subscription -> when its
# expiry warning last fired. Without it every subscription already announced
# announces itself again on the restored box — the exact re-alert storm the file
# was created to prevent (shater/alert/expiry.go).
/etc/shater/alert-state.json
# DELIBERATELY NOT KEPT. Each of these is history or cache, and the backup is
# built in RAM:
#
# /etc/shater/stats.db Traffic/query HISTORY, not configuration. Bounded
# only by globals.stats_disk_limit_mb, whose default is
# 64 MB and whose 0 means UNLIMITED — one file able to
# outweigh everything else here by two orders of
# magnitude, and the only entry whose loss costs the
# operator nothing but a chart.
# /etc/shater/cache.db sing-box's own cache (8 MiB cap, deleted above it).
# Rebuilt on demand by design, and a stale rule-set
# cache carried onto a different box is worse than no
# cache at all.
# /etc/shater/shaterd.log A log (capped by globals.log_max_kb). A restored box
# wants its own log, and this one carries the DNS query
# history of the box it came from — which is not
# something to move into an archive a person then puts
# somewhere else.
#
# ON SECRECY, since this archive routinely ends up in cloud storage: subs/*.json
# carries every node's credentials (UUID/password/keys). That is not a NEW
# exposure — /etc/config/shater already carries the subscription URLs and every
# manual node's credentials, and it is already in the backup as a conffile — but a
# shater backup is a secret-bearing file and should be treated as one.
+226
View File
@@ -262,6 +262,65 @@
}
}
/* ---- fixture band (dev builds only; see App.tsx MockBanner) ----
Deliberately outside the crit/amber vocabulary: nothing is wrong with the
router, there is no router. The hazard hatch is the service-sticker language a
piece of network hardware already uses for "this unit is not in service". */
.mock-band {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
margin-top: calc(var(--u, 8px) * 2);
padding: 10px 14px;
border: 1px dashed var(--faint);
border-radius: 9px;
background: repeating-linear-gradient(
-45deg,
var(--sink),
var(--sink) 9px,
var(--panel) 9px,
var(--panel) 18px
);
}
.mock-band-tag {
flex-shrink: 0;
align-self: flex-start;
padding: 3px 7px;
border: 1px solid var(--faint);
border-radius: 4px;
background: var(--raised);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 700;
letter-spacing: 0.14em;
color: var(--dim);
}
.mock-band-copy {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 2px;
}
.mock-band-headline {
font-family: var(--font-mono);
font-size: 12.5px;
font-weight: 700;
letter-spacing: 0.02em;
color: var(--ink);
}
.mock-band-detail {
font-size: 12.5px;
line-height: 1.5;
color: var(--dim);
max-width: 76ch;
}
.mock-band-detail code {
font-family: var(--font-mono);
font-size: 11.5px;
color: var(--ink);
}
/* ---- commit-confirm band (every page except Apply, which has the full panel) ----
Same plate as the protection banner so the two read as one family; the seconds
are the loud element because they are the only thing that is running out. */
@@ -373,6 +432,18 @@
letter-spacing: 0.06em;
color: var(--faint);
}
/* Findings come from the last SUCCESSFUL apply. When the config on disk was
refused, that is a different configuration from the one the reader just saved —
and under a heading reading "Last apply" a clean list means "your edit is
fine". A sentence about the whole list, so it gets its own line above it rather
than a third cell in the header, where 390 px left it four words a column. */
.findings-stale {
margin: calc(var(--u, 8px) * 1.25) 0 0;
max-width: 76ch;
font-size: 12px;
line-height: 1.5;
color: var(--crit);
}
.findings-list {
margin: calc(var(--u, 8px) * 1.5) 0 0;
padding: 0;
@@ -397,6 +468,13 @@
.finding--warning {
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
}
/* The daemon's "the list is capped" disclosure. Dashed, because the row is about
what ISN'T here — it must not read as one more finding to work through. */
.finding--truncated {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
background: var(--panel);
}
.finding-copy {
flex: 1;
min-width: 0;
@@ -560,3 +638,151 @@
.inline-rename-input:disabled {
opacity: 0.55;
}
/* ---- pre-apply hazard band (Apply.tsx ApplyRiskBand ← planeState.applyRisk) ----
Shown on Apply and Overview when pressing the button below would leave the
network with no way out.
It borrows the CRIT vocabulary because the outcome really is crit-severity, but
a forecast must not be mistaken for an observed fault — the panel's other crit
bands all report something that has already happened. Two things keep them
apart: the band leads with an eyebrow naming the tense, and it ends in the
ACCENT rather than in more red. Crit says how bad this is; the accent marks the
controls that prevent it. Deliberately taller and quieter-edged than
.plane-band, because unlike a banner this one is meant to be read, not
glanced at. */
.risk-band {
display: flex;
flex-direction: column;
gap: calc(var(--u, 8px) * 0.75);
margin-top: calc(var(--u, 8px) * 2);
padding: 14px 16px 15px;
border: 1px solid color-mix(in srgb, var(--crit) 55%, var(--groove));
border-left: 3px solid var(--crit);
border-radius: 9px;
background: linear-gradient(180deg, color-mix(in srgb, var(--crit) 12%, var(--raised)), var(--raised));
box-shadow: 0 1px 0 var(--edge) inset;
}
.risk-band-hd {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.25);
}
.risk-eyebrow {
font-family: var(--font-mono);
font-size: 10.5px;
font-weight: 700;
letter-spacing: var(--track-label-wide, 0.24em);
text-transform: uppercase;
color: var(--crit);
}
.risk-headline {
margin: 0;
font-family: var(--font-mono);
font-size: 13.5px;
font-weight: 700;
letter-spacing: 0.01em;
line-height: 1.35;
color: var(--ink);
}
.risk-detail,
.risk-undo {
margin: 0;
max-width: 76ch;
font-size: 12.5px;
line-height: 1.55;
color: var(--dim);
}
/* The missing auto-rollback is the part that decides whether a mistake here costs
two minutes or an SSH session, so it is the one line drawn at full ink. */
.risk-undo.hot {
color: var(--ink);
font-weight: 600;
}
.risk-steps {
margin: calc(var(--u, 8px) * 0.5) 0 0;
padding: 0;
list-style: none;
display: flex;
flex-direction: column;
gap: calc(var(--u, 8px) * 0.75);
max-width: 76ch;
}
.risk-steps li {
position: relative;
padding-left: 18px;
font-size: 12.5px;
line-height: 1.55;
color: var(--ink);
}
/* A square tick in the accent — the panel's "this is a control you touch" colour,
and the only accent in a band that is otherwise entirely red. */
.risk-steps li::before {
content: '';
position: absolute;
left: 0;
top: 0.55em;
width: 7px;
height: 7px;
border-radius: 1px;
background: var(--accent);
}
@media (max-width: 560px) {
.risk-band {
padding: 12px 13px 13px;
}
.risk-detail,
.risk-undo,
.risk-steps li {
font-size: 12px;
}
}
/* ---- the config on disk is not the one running (Overview.tsx NotAppliedBand) ----
Same crit vocabulary and the same internals as .risk-band, and deliberately so:
the severity is identical. What separates them is the TENSE, which the eyebrow
states — the hazard band forecasts what a button would do, this one reports a
state the box is already in. So this band ends in crit rather than in the
accent: there is nothing to prevent any more, and the control that clears it is
an edit to the configuration, not a button on this page.
It sits ABOVE the findings because it says what the findings are about. */
.stale-band {
display: flex;
flex-direction: column;
gap: calc(var(--u, 8px) * 0.75);
margin-top: calc(var(--u, 8px) * 2);
padding: 14px 16px 15px;
border: 1px solid color-mix(in srgb, var(--crit) 55%, var(--groove));
border-left: 3px solid var(--crit);
border-radius: 9px;
background: linear-gradient(180deg, color-mix(in srgb, var(--crit) 12%, var(--raised)), var(--raised));
box-shadow: 0 1px 0 var(--edge) inset;
}
/* The step that refused — the one word an operator acts on, so it is the one
thing here drawn at full ink inside a line of body copy. */
.stale-stage {
color: var(--ink);
font-weight: 700;
}
/* The daemon's reason, verbatim. Monospaced because it is machine text quoted
into prose, and boxed so a long generator error cannot be mistaken for our own
sentence about it. */
.stale-cause {
padding: 8px 10px;
border: 1px solid var(--groove);
border-radius: 6px;
background: color-mix(in srgb, var(--sink) 75%, transparent);
font-size: 11.5px;
line-height: 1.5;
color: var(--ink);
overflow-wrap: anywhere;
}
@media (max-width: 560px) {
.stale-band {
padding: 12px 13px 13px;
}
.stale-cause {
font-size: 11px;
}
}
+60 -5
View File
@@ -7,7 +7,8 @@ import type { Status } from './api'
import { usePendingConfirm } from './pendingConfirm'
import { bootstrapSession } from './session'
import { ROUTES, navigate, useRoute } from './router'
import { engineState, protectionState } from './planeState'
import { engineState, protectionState, serviceIntent } from './planeState'
import { truncationNote } from './findings'
import type { Route } from './router'
import { Overview, Placeholder, Nodes, Routing, Apply, DNS, Devices, Targets, Settings, Profiles, Insights, Networks } from './pages'
@@ -100,6 +101,7 @@ export function App() {
footer={<StatusBar status={status} />}
>
<Nav route={route} />
<MockBanner />
<PlaneBanner status={status} route={route} />
<ConfirmBand route={route} onChanged={() => void refreshStatus()} />
<Page route={route} status={status} onStatusChange={() => void refreshStatus()} />
@@ -170,6 +172,31 @@ function ConfirmBand({ route, onChanged }: { route: Route; onChanged: () => void
)
}
/**
* Says, on every page, that nothing on screen came from a router.
*
* Only a DEV build can ever render this — the fixtures are not in a production
* bundle (api.ts initMockBackend), so an operator cannot reach this state at all.
* It is here for the person who CAN: a footer line reading "DEMO DATA" is easy to
* work past for an afternoon and then screenshot into a bug report, and every
* number above it is invented.
*/
function MockBanner() {
if (!MOCK) return null
return (
<div className="mock-band" role="status">
<span className="mock-band-tag">FIXTURES</span>
<div className="mock-band-copy">
<span className="mock-band-headline">No router is being read</span>
<span className="mock-band-detail">
Every reading on this page is invented by <code>src/mock.ts</code> for offline
development. Drop <code>?mock</code> from the address to talk to a daemon.
</span>
</div>
</div>
)
}
/**
* The protection state, pinned under the nav on every page EXCEPT Overview
* (which shows the same state as its own headline readout — see planeState.ts).
@@ -195,12 +222,15 @@ function PlaneBanner({ status, route }: { status: Status | null; route: Route })
const criticals = (status.warnings ?? []).filter((w) => w.severity === 'critical').length
const state = protectionState(status)
// The published list is capped at 50, so with a note attached the count is a
// floor. Say "at least" rather than quoting a total the daemon didn't send.
const atLeast = truncationNote(status.warnings) ? 'At least ' : ''
// Wording comes from the shared source of truth so the banner and Overview can
// never describe the same router differently.
const headline = state.alarm
? state.headline
: `${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
: `${atLeast}${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
const detail = state.alarm
? state.detail
: 'Something you configured isn’t in effect. Review the findings before relying on it.'
@@ -240,7 +270,9 @@ function Page({
if (route === 'dns') return <DNS />
if (route === 'targets') return <Targets />
if (route === 'devices') return <Devices />
if (route === 'insights') return <Insights />
// Insights takes `status` for ONE reason: so it can say that its ten empty
// sections are empty because the service is off. It polls its own stats.
if (route === 'insights') return <Insights status={status} />
if (route === 'settings') return <Settings />
if (route === 'profiles') return <Profiles />
if (route === 'apply') return <Apply />
@@ -298,13 +330,29 @@ function StatusBar({ status }: { status: Status | null }) {
* engine that had failed to start. It now asks {@link engineState}, whose whole
* job is to be able to answer "down", and refuses to guess when nothing has been
* reported: an unlit socket, not a green light.
*
* AND IT ASKS WHETHER THE ENGINE WAS SUPPOSED TO BE RUNNING. `engineState` alone
* cannot tell a crashed engine from one nobody started: `Status.running` is the
* DAEMON's liveness and `engine_running` is the ENGINE's, and neither is the
* operator's switch. So the lamp above every page on a correctly installed,
* not-yet-configured router was crit "Engine down" — the product's most visible
* lamp reporting a failure that had not happened. {@link serviceIntent} is the
* missing question, and only a positive `off` takes this branch: with the
* configuration unreadable the intent is `unknown` and the crit stands, because
* that is the case where the LAN really has been cut off.
*/
function masterIndicator(
phase: Phase,
status: Status | null,
): { label: string; variant: LedVariant; pulse?: boolean } {
if (phase === 'loading' || !status) return { label: 'Linking', variant: 'off' }
switch (engineState(status)) {
const engine = engineState(status)
if (engine === 'down' && serviceIntent(status) === 'off') {
// Named for the act, not the symptom: "Switched off" says a person did this
// and can undo it, where "Engine down" says the appliance broke.
return { label: 'Switched off', variant: 'off' }
}
switch (engine) {
case 'down':
return { label: 'Engine down', variant: 'crit' }
case 'up':
@@ -321,8 +369,15 @@ function UnauthPlate() {
<Faceplate ariaLabel="shater — not authenticated" header={<FaceplateHeader wordmark="SHATER" subline="v0.2 · openwrt appliance" />}>
<div className="plate-msg">
<Module name="Session" value="LOCKED" led={{ variant: 'amber' }}>
{/* THE ONLY RECOVERY INSTRUCTION THE PRODUCT GIVES, so it has to point at
the real menu entry. It said "System → shater"; the page is registered
at `admin/services/shater` (luci-app-shater/root/usr/share/luci/menu.d/
luci-app-shater.json, title "Shater"), which LuCI renders under
SERVICES. Anyone reading this line has just lost access to the panel
and is looking for the one door back — sending them to the wrong menu
costs far more than its size. */}
<p className="placeholder-note">
No active session. Open the panel from the LuCI menu (System → shater →{' '}
No active session. Open the panel from the LuCI menu (Services → Shater →{' '}
<strong>Open panel</strong>) to hand off a fresh access token.
</p>
</Module>
+181
View File
@@ -0,0 +1,181 @@
// Editing an alert channel — and, above all, editing a secret the panel refuses
// to show.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// What these protect:
//
// 1. AN EMPTY TOKEN BOX KEEPS THE STORED TOKEN. The row masks the bot token
// deliberately, so it cannot be prefilled; an empty box that meant "clear it"
// would destroy a secret on every save of an unrelated field, and the only
// way to get it back is BotFather. This is the project's rule — a save may
// clear only a field the editor was able to SHOW — applied to the one field
// that can never be shown.
// 2. A TYPED TOKEN STILL REPLACES. Otherwise the fix for a typo'd token is no
// fix at all.
// 3. THE OTHER TYPE'S SETTINGS SURVIVE A TYPE SWITCH, for the same reason: the
// form in Telegram mode renders no URL box, so it may not clear a URL.
// alert/notifier.go's buildPayload reads only the selected type's fields, so
// carrying them costs nothing and switching back costs nothing either.
// 4. VALIDATION KNOWS THE DIFFERENCE between "no token was given" and "no token
// exists". The add form's rule ("Telegram needs a bot token") is one an edit
// can never satisfy, and that is exactly why they now share a validator.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { Alert } from './api'
import {
buildAlert,
draftFromAlert,
hasStoredToken,
validateAlertDraft,
TOKEN_KEEP_HINT,
} from './alertEdit.ts'
import type { AlertDraft } from './alertEdit.ts'
const tg: Alert = {
Name: 'tg',
Enabled: true,
Type: 'telegram',
Token: '123456:REAL-SECRET',
ChatID: '-1001',
Events: ['killswitch'],
}
const hook: Alert = {
Name: 'hook',
Enabled: false,
Type: 'webhook',
URL: 'https://hooks.example.com/abc?key=xyz',
Events: ['apply_fail'],
}
const draft = (over: Partial<AlertDraft> = {}): AlertDraft => ({
name: 'tg',
type: 'telegram',
token: '',
chatId: '-1001',
url: '',
events: ['killswitch'],
via: 'direct',
fallback: false,
...over,
})
// --- 1 + 2. the secret -------------------------------------------------------
test('an edit form starts with the token box EMPTY, never prefilled', () => {
const d = draftFromAlert(tg, 'direct')
assert.equal(d.token, '', 'the token must not be handed back to the browser')
// Everything the form CAN show is prefilled, so an edit is an edit and not a
// re-entry exercise.
assert.equal(d.chatId, '-1001')
assert.deepEqual(d.events, ['killswitch'])
})
test('saving with an empty token box KEEPS the stored token', () => {
// The whole point: fix the chat ID without going back to BotFather.
const next = buildAlert(tg, draft({ chatId: '-1002' }))
assert.equal(next.Token, '123456:REAL-SECRET', 'an unshown secret must never be cleared by a save')
assert.equal(next.ChatID, '-1002')
})
test('a typed token replaces the stored one, trimmed', () => {
const next = buildAlert(tg, draft({ token: ' 999:NEW-SECRET ' }))
assert.equal(next.Token, '999:NEW-SECRET')
})
test('the interface SAYS that empty means keep — it is not left to be inferred', () => {
assert.equal(hasStoredToken(tg), true)
assert.equal(hasStoredToken(hook), false)
assert.equal(hasStoredToken(null), false)
assert.match(TOKEN_KEEP_HINT, /leave this empty to keep/i)
})
// --- 3. a type switch is not a delete ----------------------------------------
test('switching telegram → webhook keeps the token and chat ID stored', () => {
const next = buildAlert(
tg,
draft({ type: 'webhook', url: 'https://hooks.example.com/x', token: '' }),
)
assert.equal(next.Type, 'webhook')
assert.equal(next.URL, 'https://hooks.example.com/x')
assert.equal(next.Token, '123456:REAL-SECRET', 'the form showed no token box — it may not clear one')
assert.equal(next.ChatID, '-1001')
})
test('saving a telegram channel does not clear a URL the form never showed', () => {
const both: Alert = { ...tg, URL: 'https://hooks.example.com/keep' }
const next = buildAlert(both, draft({ chatId: '-1003' }))
assert.equal(next.URL, 'https://hooks.example.com/keep')
})
test('a field the form DID show can be emptied — that is the difference', () => {
// In webhook mode the URL box is on screen, so clearing it is a decision the
// operator made and can see. (Validation refuses to save it; the builder's job
// is only to not invent a value.)
const next = buildAlert(hook, draft({ name: 'hook', type: 'webhook', url: '' }))
assert.equal(next.URL, undefined)
})
test('the row switch owns Enabled — an edit never flips it', () => {
assert.equal(buildAlert(hook, draft({ name: 'hook', type: 'webhook', url: hook.URL! })).Enabled, false)
assert.equal(buildAlert(tg, draft()).Enabled, true)
// A NEW channel arrives on.
assert.equal(buildAlert(null, draft({ token: 't' })).Enabled, true)
})
test('delivery: a direct channel carries no Via and no Fallback key at all', () => {
const direct = buildAlert(tg, draft({ via: 'direct', fallback: true }))
assert.equal('Via' in direct, false)
assert.equal('Fallback' in direct, false)
const routed = buildAlert(tg, draft({ via: 'group:eu', fallback: true }))
assert.equal(routed.Via, 'group:eu')
assert.equal(routed.Fallback, true)
// Turning the detour back off drops both — a stale Fallback on a direct channel
// would describe a retry path that does not exist.
const back = buildAlert(routed, draft({ via: 'direct', fallback: false }))
assert.equal('Via' in back, false)
assert.equal('Fallback' in back, false)
})
// --- 4. validation ------------------------------------------------------------
test('an edit with an empty token box passes, because one is already stored', () => {
assert.equal(validateAlertDraft(draft(), new Set(['tg']), tg), null)
})
test('a NEW telegram channel with no token is refused', () => {
const problem = validateAlertDraft(draft({ name: 'fresh', token: '' }), new Set(), null)
assert.notEqual(problem, null)
assert.match(problem!, /bot token/i)
})
test('converting a webhook to telegram demands a token — there is none to keep', () => {
const problem = validateAlertDraft(
draft({ name: 'hook', type: 'telegram', token: '', chatId: '-1' }),
new Set(['hook']),
hook,
)
assert.notEqual(problem, null)
assert.match(problem!, /none is stored/i)
})
test('renaming to its own name is allowed; colliding with another is not', () => {
const taken = new Set(['tg', 'other'])
assert.equal(validateAlertDraft(draft({ name: 'tg' }), taken, tg), null)
const problem = validateAlertDraft(draft({ name: 'other' }), taken, tg)
assert.match(problem!, /already exists/i)
})
test('the remaining shape rules still bite', () => {
assert.match(validateAlertDraft(draft({ name: ' ' }), new Set(), tg)!, /name/i)
assert.match(validateAlertDraft(draft({ chatId: '' }), new Set(), tg)!, /chat ID/i)
assert.match(validateAlertDraft(draft({ events: [] }), new Set(), tg)!, /event/i)
assert.match(
validateAlertDraft(draft({ type: 'webhook', url: 'hooks.example.com' }), new Set(), tg)!,
/http/i,
)
})
+143
View File
@@ -0,0 +1,143 @@
import type { Alert } from './api'
/**
* Editing an alert channel that already exists — and, in particular, editing a
* secret the panel refuses to show.
*
* WHY THIS MODULE EXISTS. Alerts could be created and toggled, and their delivery
* path (Via/Fallback) changed, but Type / Token / ChatID / URL / Events were
* write-once: the only way to fix a typo in a chat ID was to delete the channel
* and build it again. For the bot token that is worse than tedious — the row
* masks it deliberately (`secret hidden`), so "re-enter it" means going back to
* BotFather for a token you already own.
*
* THE RULE THIS FILE ENCODES. A save may clear only a field the editor was able
* to SHOW. The token is never shown, so an empty token box cannot mean "erase the
* stored token" — it means "keep it". The same reasoning covers the fields of the
* OTHER channel type: while the form is in Telegram mode it shows no URL box, so
* a save in Telegram mode leaves a stored URL alone (and vice versa). That is
* also why switching type is non-destructive — the settings of the other kind sit
* there unused until you switch back. Nothing reads them meanwhile:
* alert/notifier.go `buildPayload` switches on Type and touches only that type's
* fields.
*
* It lives outside `pages/Alerts.tsx` because it is the part that must be TESTED,
* and the panel's runner is `node --test src/*.test.ts`: plain modules only, no
* JSX, no DOM (same reason as ruleset.ts and egressEdit.ts).
*/
export type AlertType = 'telegram' | 'webhook'
/** The live fields of either alert form (add or edit). */
export interface AlertDraft {
name: string
type: AlertType
/** Telegram bot token. EMPTY MEANS "keep whatever is stored" — never "clear it". */
token: string
chatId: string
url: string
events: string[]
/** Canonical delivery value from the picker; 'direct' (or '') ⇒ no detour. */
via: string
fallback: boolean
}
const HTTP_RE = /^https?:\/\//i
/** Does this channel already hold a bot token? Drives which hint the form shows. */
export function hasStoredToken(a: Alert | null): boolean {
return !!(a?.Token ?? '').trim()
}
/** Prefill an edit form from a stored channel. The token box always starts EMPTY. */
export function draftFromAlert(a: Alert, via: string): AlertDraft {
return {
name: a.Name,
type: a.Type === 'webhook' ? 'webhook' : 'telegram',
token: '',
chatId: a.ChatID ?? '',
url: a.URL ?? '',
events: [...(a.Events ?? [])],
via,
fallback: a.Fallback ?? false,
}
}
/**
* What the form is allowed to submit, or the message to show instead.
*
* `stored` is the channel being edited (null when adding). It is what makes the
* token rule work in both directions: an edit passes with an empty token box
* because one is already stored, and a NEW Telegram channel — or one being
* converted from a webhook, which has no token — still has to be given one.
*/
export function validateAlertDraft(
d: AlertDraft,
taken: ReadonlySet<string>,
stored: Alert | null,
): string | null {
const nm = d.name.trim()
if (!nm) return 'Give the alert a name.'
if (nm !== (stored?.Name ?? '') && taken.has(nm)) return `An alert named “${nm}” already exists.`
if (d.type === 'telegram') {
if (!d.token.trim() && !hasStoredToken(stored)) {
return 'Telegram needs a bot token — none is stored for this alert yet.'
}
if (!d.chatId.trim()) return 'Telegram needs a chat ID.'
} else if (!HTTP_RE.test(d.url.trim())) {
return 'Enter an http(s):// webhook URL.'
}
if (d.events.length === 0) return 'Pick at least one event to notify on.'
return null
}
/**
* Build the channel a save writes.
*
* The stored channel is spread in first, so anything this form does not model
* survives untouched — the same idiom the rule editor uses for Order/Enabled/Kill.
* Enabled is deliberately taken from storage too: the row's own switch owns it,
* and an edit form that has no switch must not decide it.
*/
export function buildAlert(stored: Alert | null, d: AlertDraft): Alert {
const typed = d.token.trim()
// Telegram fields: the token box, when filled, replaces; when empty it keeps.
// In webhook mode the box is not rendered at all, so it can never speak here.
const token = d.type === 'telegram' && typed ? typed : (stored?.Token ?? '')
const chatId = d.type === 'telegram' ? d.chatId.trim() : (stored?.ChatID ?? '')
const url = d.type === 'webhook' ? d.url.trim() : (stored?.URL ?? '')
const routed = d.via !== '' && d.via !== 'direct'
const out: Alert = {
...(stored ?? ({} as Alert)),
Name: d.name.trim(),
Enabled: stored ? stored.Enabled : true,
Type: d.type,
Events: [...d.events],
}
// Written when non-empty, dropped when the operator emptied a field the form
// actually showed (a webhook URL in webhook mode, a chat ID in Telegram mode).
// `delete` rather than `''` so the PUT body stays the shape the add form sends.
if (token) out.Token = token
else delete out.Token
if (chatId) out.ChatID = chatId
else delete out.ChatID
if (url) out.URL = url
else delete out.URL
if (routed) out.Via = d.via
else delete out.Via
if (routed && d.fallback) out.Fallback = true
else delete out.Fallback
return out
}
/** Said under the token box when one is already stored. */
export const TOKEN_KEEP_HINT =
'A bot token is stored. Leave this empty to keep it — type a new one only to replace it. It is never shown back.'
/** Said under the token box when there is nothing to keep. */
export const TOKEN_NEW_HINT = 'Stored secretly and never shown again — you can replace it later.'
/** Said when an edit switches an existing channel to the other type. */
export const TYPE_SWITCH_NOTE =
'The other type’s settings stay stored and unused, so you can switch back without entering them again.'
+754 -70
View File
File diff suppressed because it is too large Load Diff
+210
View File
@@ -0,0 +1,210 @@
// The configuration on disk was refused, and everything else on the page is
// about a different one.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// THE MEASUREMENT THIS EXISTS FOR (stand, shater/apply/apply.go): with a config
// the engine could not accept, reconciliation retried it every 60 seconds — each
// retry a full engine swap that failed and rolled back — while the status said
// `engine_running: true` and carried the OLD config's hash and the OLD config's
// warnings, and named the cause nowhere.
//
// WHAT THESE PROTECT:
//
// 1. THE REFUSED STATE AND THE APPLIED STATE MUST NOT RENDER ALIKE, and the
// refused one must NAME THE STAGE.
// 2. THE CONTROL, BOTH WAYS. On a healthy box nothing new appears at all —
// otherwise "it warns when refused" is satisfiable by a helper that warns
// always. And a daemon that does not publish the field must not be alarmed
// either, because there is no evidence about it.
// 3. IT MUST NOT CONTRADICT `engine_running`. True is TRUE in this state. The
// reading has to say WHICH configuration is running, not deny that one is.
// 4. THE REST OF THE PAGE IS DISOWNED BY NAME. "Refused" alone leaves the hash,
// the traffic verdict and the findings reading as if they were about the
// edit that was just saved.
// 5. NO BROWSER-COMPUTED ELAPSED TIME. The router has no RTC.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { Status } from './api.ts'
import { appliedState, rejectedReading, shortHash } from './appliedConfig.ts'
const status = (over: Partial<Status>): Status => ({
running: true,
enabled: true,
active: true,
table: true,
hash: 'a1b2c3d4e5f60718293a4b5c6d7e8f90',
version: '0.2.19',
engine_running: true,
config_readable: true,
plane: 'full',
...over,
})
/** A box whose saved edit was refused at validation while the tunnel kept working. */
const refused = status({
config_applied: false,
apply_error: 'rule "kids": target "de-hysteria" resolves to nothing',
apply_error_stage: 'the schema gate',
apply_attempts: 7,
apply_failed_since_unix: 1_700_000_000,
})
const healthy = status({ config_applied: true, apply_error: '', apply_error_stage: '' })
// ---- the three states -------------------------------------------------------
test('the three states are three, and false is not folded into absent', () => {
assert.equal(appliedState(healthy), 'applied')
assert.equal(appliedState(refused), 'rejected')
assert.equal(appliedState(status({})), 'unknown') // a daemon without the field
assert.equal(appliedState(null), 'unknown')
})
test('a refused configuration produces a statement, and it names the stage', () => {
const r = rejectedReading(refused, '09:14:00')
assert.ok(r, 'the refusal produced nothing to show')
assert.equal(r.stage, 'schema')
assert.match(r.stageText, /the schema gate/)
assert.match(r.stageHint, /did not pass validation/)
// The daemon's reason, verbatim — not summarised into a mood.
assert.match(r.cause, /target "de-hysteria" resolves to nothing/)
})
// ---- the controls, both ways ------------------------------------------------
test('CONTROL: a healthy box gets nothing new at all', () => {
// Without this, every assertion above is satisfied by a helper that reports a
// refusal unconditionally.
assert.equal(rejectedReading(healthy, '09:14:00'), null)
})
test('CONTROL: a daemon that does not publish the field is not alarmed either', () => {
// Absent is not false. There is no evidence this box's hash is stale, and
// stamping "your edit is not in effect" across the page on no evidence is the
// same class of lie as hiding it when it is true.
assert.equal(rejectedReading(status({}), ''), null)
assert.equal(rejectedReading(null, ''), null)
})
test('CONTROL: refused and applied cannot be drawn the same way', () => {
// The one assertion the whole defect reduces to. It fails if the two states
// ever produce the same screen.
const a = rejectedReading(healthy, '09:14:00')
const b = rejectedReading(refused, '09:14:00')
assert.notEqual(a === null, b === null, 'refused and applied produced the same rendering')
assert.equal(a, null)
assert.ok(b)
})
// ---- what it says, and what it must not say ---------------------------------
test('it does NOT contradict engine_running — it says which config is running', () => {
const r = rejectedReading(refused, '09:14:00')
assert.ok(r)
assert.match(r.scope, /The engine IS running/)
assert.match(r.scope, /PREVIOUS configuration/)
// The running config named by its own short hash, so the two can be told apart.
assert.match(r.scope, /a1b2c3d4e5f6/)
})
test('the hash, the traffic verdict and the findings are disowned BY NAME', () => {
// "It was refused" is not enough. Every other reading on the page stays green
// and keeps describing the configuration that is running.
const r = rejectedReading(refused, '09:14:00')
assert.ok(r)
for (const named of [/config hash/, /where traffic goes/, /every finding/]) {
assert.match(r.scope, named)
}
assert.match(r.scope, /Your edit is not in effect/)
})
test('a stopped engine is not told the tunnel still works', () => {
const r = rejectedReading(status({ ...refused, engine_running: false }), '09:14:00')
assert.ok(r)
assert.doesNotMatch(r.scope, /tunnel still works/)
assert.match(r.scope, /Nothing of this configuration is running/)
// …and the rest of the page is still disowned, because the readings are still
// there and still about something else.
assert.match(r.scope, /every finding/)
})
test('persistence says whether anything is still trying, with the router’s clock', () => {
const r = rejectedReading(refused, '09:14:00')
assert.ok(r)
assert.match(r.persistence, /Tried 7 times/)
assert.match(r.persistence, /first refused at 09:14:00 on the router’s clock/)
assert.match(r.persistence, /retried on a widening interval/)
assert.match(r.persistence, /saving any change to the configuration cancels the wait/)
})
test('no elapsed time is computed in the browser — the router has no RTC', () => {
// Passing no formatted clock must not make the panel invent one from
// `apply_failed_since_unix` against the browser's clock.
const r = rejectedReading(refused, '')
assert.ok(r)
assert.doesNotMatch(r.persistence, /minute\(s\)|ago|for \d+ /)
assert.match(r.persistence, /Tried 7 times\./)
})
// ---- the closed stage vocabulary --------------------------------------------
test('the three stages are recognised and each says something different', () => {
const seen = new Set<string>()
for (const [phrase, kind] of [
['the schema gate', 'schema'],
['building the engine configuration', 'generate'],
['starting the engine', 'engine'],
] as const) {
const r = rejectedReading(status({ ...refused, apply_error_stage: phrase }), '')
assert.ok(r)
assert.equal(r.stage, kind)
assert.match(r.stageText, new RegExp(phrase))
assert.ok(r.stageHint.length > 0, `${kind} has no hint`)
seen.add(r.stageHint)
}
assert.equal(seen.size, 3, 'two stages tell the operator the same thing')
})
test('an unrecognised stage is shown verbatim and labelled, never guessed', () => {
const r = rejectedReading(status({ ...refused, apply_error_stage: 'installing the data plane' }), '')
assert.ok(r)
assert.equal(r.stage, 'unrecognised')
assert.match(r.stageText, /installing the data plane/)
assert.match(r.stageText, /does not recognise/)
assert.equal(r.stageHint, '', 'a stage we cannot name must not carry advice about another one')
})
test('a stage the daemon never named is its own state, not one of the three', () => {
const r = rejectedReading(status({ ...refused, apply_error_stage: '' }), '')
assert.ok(r)
assert.equal(r.stage, 'unnamed')
assert.match(r.stageText, /did not name/)
})
test('a refusal with no recorded reason still reads as a refusal', () => {
const r = rejectedReading(status({ ...refused, apply_error: '' }), '')
assert.ok(r)
assert.match(r.cause, /recorded no reason/)
assert.notEqual(r.cause, '', 'an empty cause line reads as “nothing wrong”')
})
test('shortHash matches the daemon’s own short form, prefix or not', () => {
assert.equal(shortHash('a1b2c3d4e5f60718293a4b5c6d7e8f90'), 'a1b2c3d4e5f6')
assert.equal(shortHash('abc'), 'abc')
assert.equal(shortHash(''), '')
// A prefixed hash must shorten to the SAME twelve characters, or the band and
// the Engine module's hash row print two different strings for one config on
// one screen — in the one state where telling configs apart is the whole job.
assert.equal(shortHash('sha256:a1b2c3d4e5f60718293a4b5c6d7e8f90'), 'a1b2c3d4e5f6')
})
test('the band names the running config exactly as the hash row shortens it', () => {
const h = 'sha256:9f7c0abc12345678'
const r = rejectedReading(status({ ...refused, hash: h }), '')
assert.ok(r)
assert.match(r.scope, new RegExp(`\\(${shortHash(h)}\\)`))
assert.doesNotMatch(r.scope, /sha256:/)
})
+208
View File
@@ -0,0 +1,208 @@
// Is the configuration on disk the one that is RUNNING — and if not, what is
// everything else on the screen actually describing?
//
// Backend contract: shater/apply/apply.go — Status.ConfigApplied / ApplyError /
// ApplyErrorStage / ApplyAttempts / ApplyFailedSinceUnix, and Applier.rejected,
// the record that makes them possible.
//
// THE MEASUREMENT THIS EXISTS FOR. With a configuration on disk the engine
// cannot accept, the daemon's reconciliation retried it every 60 seconds — each
// retry a full engine swap that failed and rolled back — while the status body
// said `engine_running: true` and carried the OLD configuration's hash and the
// OLD configuration's warnings. The reason for the refusal appeared nowhere in
// it. The owner watched a healthy engine while the tunnel restarted once a
// minute and their edit did nothing.
//
// `engine_running: true` is TRUE in that state and this module never contradicts
// it. The whole job here is to say WHICH configuration is running, and to mark
// every value on the page that describes that older one instead of the one the
// operator just saved.
import type { Status } from './api'
/**
* A CLOSED set of three, and `unknown` is a real member.
*
* applied — the configuration on disk IS the one running.
* rejected — it was read and REFUSED. Something else is running.
* unknown — the key is ABSENT: nobody measured it, so there is nothing to
* assert in either direction.
*
* WHY `unknown` IS NOT FOLDED INTO `rejected`. `false` is a MEASURED refusal and
* earns the alarm; absence is not a measurement at all, and the daemon spends a
* pointer to keep the two apart (apply.Status.ConfigApplied is a *bool with
* `omitempty`). It has to: `false` publishes "THE CONFIGURATION ON DISK HAS NOT
* BEEN APPLIED … your edit is not in effect", so with a plain bool that alarm
* would be the ZERO VALUE OF THE TYPE — raised by any object that simply never
* mentioned the field, such as the offline stub `shaterd status` prints with no
* daemon answering, over a data plane that may have been running for weeks.
*
* A live status always carries one of the two verdicts, so absence here means
* either that stub or a daemon older than the field — and neither is evidence
* that this router's hash is stale. Stamping "your edit is not in effect" across
* a page on no evidence is the same class of lie as hiding it when it is true.
* `unknown` therefore claims nothing: it neither raises the alarm nor certifies
* the page. This is the same treatment `plane` and `traffic` already get.
*/
export type AppliedState = 'applied' | 'rejected' | 'unknown'
export function appliedState(status: Status | null | undefined): AppliedState {
const v = status?.config_applied
if (v === true) return 'applied'
if (v === false) return 'rejected'
return 'unknown'
}
/** The stage vocabulary, normalised. POSITIVE and CLOSED — `unrecognised` is how
* a future fourth stage arrives, and it is shown verbatim rather than guessed
* at. `unnamed` is the daemon saying nothing, which is not the same thing. */
export type ApplyStage = 'schema' | 'generate' | 'engine' | 'unnamed' | 'unrecognised'
/** The three phrases apply.go publishes (stageSchemaGate / stageGenerate /
* stageEngine), verbatim. Matched exactly: the daemon writes these constants, so
* a value that is merely close is a value this build does not know. */
const STAGES: Record<string, ApplyStage> = {
'the schema gate': 'schema',
'building the engine configuration': 'generate',
'starting the engine': 'engine',
}
export interface RejectedReading {
stage: ApplyStage
/** The step, named. Always non-empty — silence here is itself reported. */
stageText: string
/** What the operator does about this stage. '' when the stage is not known. */
stageHint: string
/** The daemon's reason, verbatim, or a statement that it gave none. */
cause: string
/** Whether anything is still trying, and how long it has been like this. */
persistence: string
/**
* WHAT THE REST OF THE PAGE IS ABOUT. This is the sentence the defect was
* missing: it is not enough to say the config was refused, because every other
* reading on screen stays green and keeps describing the configuration that is
* running.
*/
scope: string
}
/**
* The full statement of a standing refusal — or `null` when there is nothing to
* state, which is every status that is not `rejected`.
*
* `hash` is the short form of the RUNNING configuration's hash, and it is in the
* text on purpose: the single most misleading thing about this state is that
* `hash` in the same status body looks entirely normal. It names the config the
* operator is actually running so the two can be told apart.
*
* NO ELAPSED TIME IS COMPUTED HERE. `apply_failed_since_unix` is the router's
* clock, the router has no RTC, and a browser-side "failing for 12 minutes" would
* be fiction whenever the two clocks differ. The attempt COUNT carries
* persistence; the instant is handed back to the caller to format against the
* router's clock, like every other router timestamp in this panel.
*/
export function rejectedReading(
status: Status | null | undefined,
/** Already-formatted router-clock time of `apply_failed_since_unix`; '' when
* there is none. The caller owns formatting — this module has no locale. */
sinceClock = '',
): RejectedReading | null {
if (appliedState(status) !== 'rejected') return null
const raw = (status?.apply_error_stage ?? '').trim()
const stage: ApplyStage = raw === '' ? 'unnamed' : (STAGES[raw] ?? 'unrecognised')
const err = (status?.apply_error ?? '').trim()
const attempts = status?.apply_attempts ?? 0
return {
stage,
stageText: stageText(stage, raw),
stageHint: stageHint(stage),
// A refusal whose reason was not recorded is still a refusal, and saying so is
// the point — an empty line here would read as "no reason, so probably fine".
cause: err || 'The daemon recorded no reason for the refusal.',
persistence: persistenceText(attempts, sinceClock),
scope: scopeText(status),
}
}
function stageText(stage: ApplyStage, raw: string): string {
switch (stage) {
case 'schema':
return 'the schema gate'
case 'generate':
return 'building the engine configuration'
case 'engine':
return 'starting the engine'
case 'unrecognised':
// Verbatim, and labelled. A step this build does not know is still the step
// that refused, and guessing which of the three it resembles would send the
// operator to the wrong page.
return `a step this panel does not recognise — the daemon calls it “${raw}”`
case 'unnamed':
return 'a step the daemon did not name'
}
}
function stageHint(stage: ApplyStage): string {
switch (stage) {
case 'schema':
return 'The configuration did not pass validation, so nothing was built from it. The cause names the setting.'
case 'generate':
return 'The configuration is valid and could not be turned into an engine configuration — a reference that resolves to nothing, or a combination the generator refuses.'
case 'engine':
return 'The configuration built, and the engine would not come up on it — a port already taken, or a transport that fails at start.'
case 'unrecognised':
case 'unnamed':
return ''
}
}
/**
* Is anything still trying, and since when.
*
* The retry is on a widening interval and ANY edit cancels the wait, so "it will
* be retried" is true and worth saying: an operator who reads "refused" alone
* cannot tell whether the box has given up.
*/
function persistenceText(attempts: number, sinceClock: string): string {
const tries =
attempts > 0
? `Tried ${attempts} time${attempts === 1 ? '' : 's'}`
: 'The daemon did not say how many times it has been tried'
const since = sinceClock ? `, first refused at ${sinceClock} on the router’s clock` : ''
return `${tries}${since}. It is retried on a widening interval, and saving any change to the configuration cancels the wait and tries it again at once.`
}
/**
* The sentence that names what everything else on screen describes.
*
* It is built from what the status actually carries, so it never promises a
* reading that is not on the page: with no `hash` there is no hash to disown.
*/
function scopeText(status: Status | null | undefined): string {
const running = shortHash(status?.hash ?? '')
const engineUp = status?.engine_running === true
const head = engineUp
? `The engine IS running — on the PREVIOUS configuration${running ? ` (${running})` : ''}, not this one. That is why the tunnel still works.`
: `Nothing of this configuration is running${running ? `; the hash shown is ${running}` : ''}.`
return `${head} Everything else on this page describes that older configuration: the config hash, where traffic goes, and every finding below. Your edit is not in effect.`
}
/**
* The leading 12 characters of a config hash — the same short form the daemon's
* own warning uses (apply.shortHash), so the two texts name the config
* identically.
*
* The `sha256:` prefix is stripped first. The daemon's hash is bare hex
* (engine.hashOptions), but a prefixed form exists in fixtures and older data, and
* without the strip the twelve characters would be spent on the word "sha256:" —
* so this band and the Engine module's hash row would print two different strings
* for one configuration, on one screen, in the one state where telling two
* configurations apart is the whole job. Overview's hash row calls THIS function
* for that reason; there is deliberately only one normalisation.
*/
export function shortHash(h: string): string {
const bare = h.replace(/^sha256:/, '')
return bare.length <= 12 ? bare : bare.slice(0, 12)
}
+287
View File
@@ -0,0 +1,287 @@
// applyRisk — the warning shown BEFORE the apply that takes the network down.
//
// Traced through the daemon: nothing rejects an empty config (model.Validate is
// advisory, generate errors only on a nil model, the engine starts because the
// outbound list always holds direct+block); with the kill-switch closed and no
// catch-all rule generate/route.go sets `route.final = "block"`; and the tproxy
// divert for the shipped `lan` inbound is installed, so every TCP connection and
// UDP flow from the LAN is handed to the engine and dropped.
//
// A predictive warning is only worth having if it is quiet on configs that are
// fine, so every "it fires" case below is paired with the single-field change
// that must silence it. Four conditions, four ways for the danger to be absent.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { applyRisk, isCatchAllRule, isInterceptingInbound } from './planeState.ts'
import type { ApplyRisk, ApplyRiskInput, CatchAllRule } from './planeState.ts'
/** The shipped config with the service switched on and nothing else changed:
* fail-closed, one enabled tproxy inbound, no rules, no resolvers, no window. */
function shipped(over: Partial<ApplyRiskInput> = {}): ApplyRiskInput {
return {
Globals: { Enabled: true, KillSwitch: 'closed', ConfirmTimeout: 0 },
Rules: [],
Inbounds: [{ Enabled: true, Type: 'tproxy' }],
Resolvers: [],
...over,
}
}
/** A rule with no matcher of any kind — the effective catch-all (route.Final). */
function catchAll(over: Partial<CatchAllRule> = {}): CatchAllRule {
return { Enabled: true, ...over }
}
// --- it fires on the dangerous config ---------------------------------------
test('warns before an apply that blocks every device', () => {
const r = applyRisk(shipped())
assert.ok(r, 'the shipped config, switched on, blocks everything')
// Says what the ACTION does, in traffic terms, not what is wrong with a field.
assert.match(r.headline, /cuts every device off the internet/)
assert.match(r.detail, /drops it/)
// Only certain claims: TCP and UDP are what tproxy diverts. ICMP depends on
// the Untunnelable policy, so it is not promised here either way.
assert.match(r.detail, /TCP connection and UDP flow/)
// What still works, so nobody power-cycles a router they can still reach.
assert.match(r.detail, /reach each other and this panel/)
// And the fix, naming the page and the shape of the rule.
assert.match(r.steps[0], /Routing page/)
assert.match(r.steps[0], /no conditions/)
})
test('names the missing auto-rollback, and the value to set', () => {
const r = applyRisk(shipped())
assert.ok(r)
assert.equal(r.noAutoRollback, true)
assert.match(r.undo, /no auto-rollback/)
assert.match(r.undo, /SSH/)
// The README's documented first-apply value, which the panel never mentioned.
assert.ok(
r.steps.some((s) => /120 seconds/.test(s) && /Settings/.test(s)),
'the confirm-window remedy must be offered, with the documented value',
)
})
test('an armed confirm window changes the undo line, not the warning', () => {
const r = applyRisk(shipped({ Globals: { Enabled: true, KillSwitch: 'closed', ConfirmTimeout: 120 } }))
assert.ok(r, 'a confirm window does not make blocking the network unremarkable')
assert.equal(r.noAutoRollback, false)
assert.match(r.undo, /120s/)
assert.match(r.undo, /reverts to the last-good one/)
assert.doesNotMatch(r.undo, /SSH/)
// ...and it stops offering a remedy the operator has already applied.
assert.ok(!r.steps.some((s) => /120 seconds/.test(s)))
})
test('zero resolvers is named, because DNS is the one thing that still leaves', () => {
// CONTROL FIRST: one configured resolver and the step is gone. A DNS warning
// that appears whatever the DNS config says is not a reading of the DNS config.
const withResolver = applyRisk(shipped({ Resolvers: [{}] }))
assert.ok(withResolver)
assert.ok(!withResolver.steps.some((s) => /DNS page/.test(s)))
const without = applyRisk(shipped())
assert.ok(without)
const dns = without.steps.find((s) => /DNS page/.test(s))
assert.ok(dns, 'with no resolver, lookups go to the provider unprotected — say so')
assert.match(dns, /in the clear/)
// IT IS NOT ONLY THE QUERIES ADDRESSED TO THE ROUTER. The shipped
// `dns_intercept '1'` pulls a query aimed at a resolver the device picked for
// ITSELF into the engine too, and it leaves as the same clear UDP/53 — measured
// both ways on the stand, each producing its own plaintext packet on the WAN.
// The earlier wording covered only the router-addressed half, which reads as a
// promise that the other half is contained. It is not.
assert.match(dns, /to the router or to a resolver they picked themselves/)
// ...and encrypted DNS is not the way out of it either: :853 out of the LAN
// measured connects=0, because the plan rejects it.
assert.match(dns, /:853/)
// The claim the detail below is not allowed to contradict.
assert.match(dns, /only thing that still leaves this network/)
})
// --- the band may not contradict itself -------------------------------------
//
// MEASURED on the stand: in this exact state 0 packets left the WAN across the
// whole run (27 in the control that differs only by an added catch-all) — except
// for exactly 2, both plaintext UDP/53. So the DNS step is right and the detail's
// "nothing reaches the internet" was wrong by those 2 packets.
//
// The wording fix alone does not survive the next editor, because the two halves
// live 15 lines apart and each reads fine on its own. What follows is therefore
// not a check of the words but of the INVARIANT between them: no clause of the
// band may claim that nothing leaves while another clause names something that
// does. Edit either half back without the other and this fails.
/** The band is one utterance to one operator, so a claim in the steps and a claim
* in the detail are claims in the same breath. Flatten it to clauses. */
function bandClauses(r: ApplyRisk): string[] {
return [r.headline, r.detail, r.undo, ...r.steps]
.join(' ')
.split(/[.;:](?=\s|$)/)
.map((c) => c.trim())
.filter(Boolean)
}
/** "…<verb> … <the outside>" — a clause that talks about traffic leaving. */
const EXIT = /\b(?:reach(?:es)?|leaves?|leaving|gets? out|goes? out|going out)\b.*?\b(?:internet|this network|your provider)\b/i
/** Universal negation, and ONLY universal negation: "No resolver is configured"
* must NOT count, or the instrument goes blind on the very clause it exists to
* see. */
const NONE = /\bnothing\b|\bno traffic\b|\bnone of it\b/i
/** What turns "nothing leaves" into a survivable claim by admitting an exception. */
const QUALIFIED = /\belse\b|\bexcept\b|\bapart from\b|\bother than\b|\baside from\b/i
const saysNothingLeaves = (c: string) => EXIT.test(c) && NONE.test(c) && !QUALIFIED.test(c)
const saysSomethingLeaves = (c: string) => EXIT.test(c) && !NONE.test(c)
test('the band never says both "nothing leaves" and "DNS leaves"', () => {
// The instrument, proved on a fabricated band before it is trusted on a real
// one: the absolute claim IS detectable, and it IS detected next to the DNS
// clause. Without this, a regex that matches nothing would pass silently.
const control = bandClauses({
headline: 'x',
detail: 'Devices can still reach each other and this panel; nothing reaches the internet.',
undo: 'x',
steps: ['It is the only thing that still leaves this network.'],
noAutoRollback: true,
})
assert.ok(control.some(saysNothingLeaves), 'the absolute-claim detector must be able to fire')
assert.ok(control.some(saysSomethingLeaves), 'the something-leaves detector must be able to fire')
// Now the real bands, in every shape that renders one.
const bands = [
shipped(),
shipped({ Resolvers: [{}] }),
shipped({ Globals: { Enabled: true, KillSwitch: 'closed', ConfirmTimeout: 120 } }),
shipped({ Rules: [catchAll({ Enabled: false })] }),
shipped({ Rules: [catchAll({ DstRuleset: ['ru'] })], Resolvers: [{}, {}] }),
].map((cfg) => {
const r = applyRisk(cfg)
assert.ok(r, 'this config must still produce a band, or the check below is vacuous')
return r
})
for (const r of bands) {
const clauses = bandClauses(r)
const denies = clauses.filter(saysNothingLeaves)
const admits = clauses.filter(saysSomethingLeaves)
assert.ok(
!(denies.length > 0 && admits.length > 0),
`the band contradicts itself:\n denies: ${JSON.stringify(denies)}\n admits: ${JSON.stringify(admits)}`,
)
}
// AND THE INSTRUMENT IS LOOKING AT THE LIVE TEXT. Without a resolver the band
// must contain a clause admitting that something leaves — if that ever stops
// being true, the loop above passes for the wrong reason.
const dangerous = applyRisk(shipped())
assert.ok(dangerous)
assert.ok(
bandClauses(dangerous).some(saysSomethingLeaves),
'the no-resolver band must still name the traffic that gets out',
)
// The detail's half of the invariant, pinned by hand so the one-word edit that
// breaks it is named in the failure and not just inferred.
assert.match(dangerous.detail, /nothing else reaches the internet/)
assert.doesNotMatch(dangerous.detail, /nothing reaches the internet/)
// CONTROL: with a resolver configured nothing is claimed to leave at all, so
// the invariant is satisfied by the other side of the same test.
const safe = applyRisk(shipped({ Resolvers: [{}] }))
assert.ok(safe)
assert.equal(bandClauses(safe).filter(saysSomethingLeaves).length, 0)
})
// --- the four ways it must stay silent (the controls) -----------------------
test('silent when the service will be off — no plane is built at all', () => {
assert.equal(
applyRisk(shipped({ Globals: { Enabled: false, KillSwitch: 'closed', ConfirmTimeout: 0 } })),
null,
)
})
test('silent when the kill-switch is open — unmatched traffic leaves, it is not dropped', () => {
assert.equal(
applyRisk(shipped({ Globals: { Enabled: true, KillSwitch: 'open', ConfirmTimeout: 0 } })),
null,
)
// ...and the daemon's own normalisation is what decides "closed", so every
// spelling that blocks on the router must still warn here.
for (const spelling of ['closed', 'Closed', ' closed ', '', 'CLOSED', 'whatever']) {
assert.ok(
applyRisk(shipped({ Globals: { Enabled: true, KillSwitch: spelling, ConfirmTimeout: 0 } })),
`kill_switch '${spelling}' blocks on the router, so it must warn here`,
)
}
// Only a case-insensitive, trimmed "open" is fail-open.
assert.equal(
applyRisk(shipped({ Globals: { Enabled: true, KillSwitch: ' Open ', ConfirmTimeout: 0 } })),
null,
)
})
test('silent when nothing intercepts — no enabled tproxy inbound diverts anything', () => {
assert.equal(applyRisk(shipped({ Inbounds: [] })), null)
assert.equal(applyRisk(shipped({ Inbounds: null })), null)
assert.equal(applyRisk(shipped({ Inbounds: [{ Enabled: false, Type: 'tproxy' }] })), null)
// socks/http/dokodemo are local listeners; they intercept nothing on their own.
assert.equal(applyRisk(shipped({ Inbounds: [{ Enabled: true, Type: 'socks' }] })), null)
assert.equal(applyRisk(shipped({ Inbounds: [{ Enabled: true, Type: 'http' }] })), null)
assert.equal(applyRisk(shipped({ Inbounds: [{ Enabled: true, Type: 'dokodemo' }] })), null)
// An absent/empty Type IS tproxy (model.Inbound.EffectiveType), so it warns.
assert.ok(applyRisk(shipped({ Inbounds: [{ Enabled: true, Type: '' }] })))
assert.ok(applyRisk(shipped({ Inbounds: [{ Enabled: true }] })))
})
test('silent when a default rule exists — that is what stops final being block', () => {
assert.equal(applyRisk(shipped({ Rules: [catchAll()] })), null)
// A DISABLED catch-all is not one: it is not emitted, so final stays block.
assert.ok(applyRisk(shipped({ Rules: [catchAll({ Enabled: false })] })))
// Neither is a rule that only matches SOME traffic.
assert.ok(applyRisk(shipped({ Rules: [catchAll({ DstRuleset: ['ru'] })] })))
assert.ok(applyRisk(shipped({ Rules: [catchAll({ Src: ['192.168.1.0/24'] })] })))
assert.ok(applyRisk(shipped({ Rules: [catchAll({ DstPort: '443' })] })))
assert.ok(applyRisk(shipped({ Rules: [catchAll({ Proto: 'tcp' })] })))
// One catch-all among specific rules is still a catch-all.
assert.equal(
applyRisk(shipped({ Rules: [catchAll({ DstPort: '443' }), catchAll()] })),
null,
)
})
test('an UNMIGRATED rule is never the default — the case that would hide the warning', () => {
// Its destination is still in schema-v1 options the parser no longer reads, so
// "no matchers" means "unreadable destination", not "matches everything" — and
// the daemon holds it disabled. Counting it would silence this warning on
// exactly the config that needs it.
assert.equal(isCatchAllRule(catchAll({ LegacyDst: ['example.com'] })), false)
assert.ok(
applyRisk(shipped({ Rules: [catchAll({ LegacyDst: ['example.com'] })] })),
'an unmigrated rule must not be mistaken for a default route',
)
// CONTROL: the same rule once migrated does silence it.
assert.equal(applyRisk(shipped({ Rules: [catchAll({ LegacyDst: [] })] })), null)
})
// --- the small predicates, directly -----------------------------------------
test('isInterceptingInbound is a closed positive list', () => {
assert.equal(isInterceptingInbound({ Enabled: true, Type: 'TPROXY' }), true)
assert.equal(isInterceptingInbound({ Enabled: true, Type: ' tproxy ' }), true)
assert.equal(isInterceptingInbound({ Enabled: true }), true)
// Anything the panel does not know about must NOT be assumed to intercept —
// an open `!== 'socks'` test would warn about a router that diverts nothing.
assert.equal(isInterceptingInbound({ Enabled: true, Type: 'something-new' }), false)
assert.equal(isInterceptingInbound({ Enabled: false }), false)
})
test('null and missing input produce no warning rather than a guess', () => {
assert.equal(applyRisk(null), null)
assert.equal(applyRisk(undefined), null)
assert.equal(applyRisk(shipped({ Rules: null, Resolvers: null })) !== null, true)
})
+137
View File
@@ -0,0 +1,137 @@
// A BLOCKED CHAIN IS A FIELD NOW, NOT A SENTENCE.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// WHAT THIS PROTECTS. The prober walks a chain in order and stops at the first
// hop that does not answer, so the chain's exit is never dialled. There is no
// end-to-end measurement, and the daemon correctly files the row `source:''` —
// which a client reading `source` strictly puts in the "nobody looked" bucket.
// That is the wrong colour: a probe DID run, at the hop, and it failed.
//
// The panel used to keep such a row red by matching a fragment of the daemon's
// error sentence. It was the last place prose decided anything here, and a
// reworded message would have silently turned a red row grey. `blocked_by` is
// the same fact as a NUMBER.
//
// 1. A NON-ZERO blocked_by KEEPS THE ROW RED, though nothing measured it.
// 2. THE CONTROL: zero does NOT. Written so that a helper which reds
// everything, or one which reds nothing, fails.
// 3. NO PROSE. The daemon's sentence can be reworded to anything at all and
// the colour does not move.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { GroupTestResult } from './api.ts'
import { blockedHop, readingTone, testReading } from './testResult.ts'
const chainRow = (over: Partial<GroupTestResult>): GroupTestResult => ({
group: 'ewan-wg-subs',
selected: '',
kind: 'chain',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: '',
tested_unix: 1_700_000_000,
source: '',
blocked_by: 0,
...over,
})
// --- 1 + 2. red on a hop, not red on zero -------------------------------------
test('a chain blocked at a hop stays RED, though nothing measured its exit', () => {
const blocked = chainRow({
blocked_by: 3,
error:
'hop 3 of this chain was probed and did not answer, so nothing reaches the exit through it — fix that hop first',
})
assert.equal(blockedHop(blocked), 3, 'the hop is 1-based and comes off the field')
assert.equal(testReading(blocked), 'not-measured', 'there is still no end-to-end measurement')
assert.equal(readingTone(blocked), 'crit', 'but a probe did run, at the hop, and failed')
})
test('CONTROL: blocked_by = 0 is NOT red — the escalation is exactly one case', () => {
// Every non-chain result carries 0. If this went red too, the field would have
// replaced a narrow prose match with a blanket one, which is worse than what
// it replaced.
const quiet = chainRow({
blocked_by: 0,
error:
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
})
assert.equal(blockedHop(quiet), 0)
assert.equal(
readingTone(quiet),
'unknown',
'an unlit lamp: painting this red reports a fault nobody has found',
)
assert.notEqual(readingTone(quiet), 'crit')
})
test('CONTROL: the helper can still say a row is good, and a measured failure bad', () => {
// Without these, "0 is not crit" would also be satisfied by a classifier that
// never says anything at all.
const alive = chainRow({ ok: true, delay_ms: 42, source: 'observatory' })
const dead = chainRow({
source: 'observatory',
error: 'the observatory’s probe through this path failed',
})
assert.equal(readingTone(alive), 'good')
assert.equal(readingTone(dead), 'crit')
assert.equal(blockedHop(alive), 0, 'a working chain names no blocking hop')
})
// --- 3. the prose no longer decides anything -----------------------------------
test('the daemon may reword its sentence freely — the colour comes off the field', () => {
// This is the regression the change exists for. Under the old string match,
// every one of these rows would have gone quiet grey.
for (const error of [
'hop 3 did not respond',
'chain stopped at hop 3',
'',
'a sentence this build has never seen',
]) {
const r = chainRow({ blocked_by: 3, error })
assert.equal(readingTone(r), 'crit', JSON.stringify(error))
}
})
test('…and the old sentence WITHOUT the field no longer reds a row on its own', () => {
// The mirror: prose alone is not evidence any more. A daemon that sends the
// sentence and no field is one this panel does not ship with, and guessing
// from its wording is exactly the coupling that was removed.
const proseOnly = chainRow({
blocked_by: undefined,
error:
'hop 3 of this chain was probed and did not answer, so nothing reaches the exit through it — fix that hop first',
})
assert.equal(blockedHop(proseOnly), 0)
assert.equal(readingTone(proseOnly), 'unknown')
})
test('a nonsense hop number claims nothing — the escalation is not entered by accident', () => {
for (const raw of [0, -1, 0.5, NaN, Infinity, '3' as unknown as number]) {
assert.equal(blockedHop(chainRow({ blocked_by: raw })), 0, String(raw))
}
assert.equal(blockedHop(chainRow({ blocked_by: 1 })), 1, 'one IS a valid hop — 1-based')
assert.equal(blockedHop(chainRow({ blocked_by: 4.9 })), 4, 'a fractional hop floors to a real one')
})
test('a measured failure is never downgraded by a missing hop, nor upgraded by one', () => {
const measuredFail = chainRow({
source: 'observatory',
blocked_by: 0,
error: 'the observatory’s probe through this path failed',
})
assert.equal(readingTone(measuredFail), 'crit')
const measuredOK = chainRow({ ok: true, source: 'observatory', blocked_by: 3 })
assert.equal(
readingTone(measuredOK),
'good',
'blocked_by only ever escalates a not-measured row; it cannot overturn a measurement',
)
})
+132
View File
@@ -0,0 +1,132 @@
// TWO KINDS OF ROW ON ONE BOARD, and they may not look alike.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// WHAT THIS PROTECTS. Starting a test run used to replace the results outright,
// so pressing Test on one NODE blanked every group and chain card on the Targets
// screen — the daemon threw true measurements away and nothing on screen
// explained it. It no longer does: the rows a run will not itself re-measure are
// carried forward (engine.startTestRun), capped at 64, oldest evicted first.
//
// That fixes one lie and opens the door to another. A card can now show a
// reading taken twenty minutes ago beside a card showing one taken a second ago,
// and if both are drawn as a bare timestamp the whole board reads "as of now".
//
// Attribution is NOT by comparing timestamps: the router has no RTC, so its
// clock can sit far from the browser's and any computed "n minutes ago" would be
// fiction. The daemon publishes `scope`, the set of names this run covers.
//
// 1. THIS RUN'S ROW AND A CARRIED ROW ARE DIFFERENT ON SCREEN. Written so
// that drawing them identically fails.
// 2. THE CONTROL. The same helper must produce the plain, unqualified stamp
// too, or "they differ" is satisfied by a function that marks everything.
// 3. AN EMPTY SCOPE IS "CANNOT ATTRIBUTE", NOT "EVERYTHING IS CARRIED". A run
// always covers at least one target, so an empty set only ever means the
// daemon published none — and the panel's normalizer turns an absent field
// into exactly that.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { GroupTestResult } from './api.ts'
import { originStamp, rowOrigin } from './testResult.ts'
const row = (over: Partial<GroupTestResult>): GroupTestResult => ({
group: 'auto',
selected: 'nl-reality-2',
kind: 'group',
delay_ms: 42,
exit_ip: '185.12.34.56',
exit_country: 'NL',
ok: true,
error: '',
tested_unix: 1_700_000_000,
source: 'observatory',
...over,
})
const CLOCK = '14:02:11'
// --- 1 + 2. the distinction, and the control ---------------------------------
test('a row this run measured and a row it carried are attributed differently', () => {
const scope = ['stealth']
assert.equal(rowOrigin(row({ group: 'stealth' }), scope), 'this-run')
assert.equal(
rowOrigin(row({ group: 'auto' }), scope),
'carried',
'this run never touched `auto`; its reading is from an earlier press',
)
})
test('and they are DRAWN differently — the same fact, at the pixel level', () => {
const mine = originStamp('this-run', CLOCK)
const carried = originStamp('carried', CLOCK)
assert.notEqual(
mine.text,
carried.text,
'a screen that stamps both with the same bare time says the board is current when half of it is not',
)
assert.match(carried.text, /earlier run/, 'the words have to name the fact, not hint at it')
assert.equal(mine.text, CLOCK, 'this run needs no qualifier — it IS the reading just taken')
assert.notEqual(mine.hint, carried.hint)
})
test('CONTROL: the same helper does produce an unqualified stamp', () => {
// Without this, "they differ" would also be satisfied by a function that
// stamped EVERY row "earlier run" — which would be a different lie, and the
// one that makes an operator distrust a number they just asked for.
for (const origin of ['this-run', 'unattributed'] as const) {
assert.equal(
originStamp(origin, CLOCK).text,
CLOCK,
`${origin} must not be dressed as carried`,
)
assert.doesNotMatch(originStamp(origin, CLOCK).text, /earlier/)
}
})
test('a carried row says so even when it carries no timestamp at all', () => {
// The worst case for silence: nothing to print, and the row is still not this
// run's. `tested_unix: 0` reaches here as an empty clock string.
const carried = originStamp('carried', '')
assert.equal(carried.text, 'earlier run')
assert.notEqual(carried.text, originStamp('this-run', '').text)
assert.equal(originStamp('this-run', '').text, '', 'and an unstamped fresh row prints nothing')
})
// --- 3. an empty scope claims nothing ------------------------------------------
test('an EMPTY scope is "cannot attribute", not "everything is carried"', () => {
// The panel's normalizer turns an absent `scope` into `[]` before this sees
// it, so this branch is the pre-scope daemon. Reading it as a real scope would
// stamp "earlier run" on every row of a run that had just measured them all.
for (const scope of [[], undefined, null]) {
assert.equal(rowOrigin(row({}), scope), 'unattributed', JSON.stringify(scope))
}
assert.equal(originStamp('unattributed', CLOCK).text, CLOCK)
assert.doesNotMatch(
originStamp('unattributed', CLOCK).hint,
/earlier run/,
'saying nothing is the honest answer here; guessing "carried" is not',
)
})
test('a run over every target leaves nothing carried', () => {
const all = ['auto', 'stealth', 'via-tunnel']
for (const group of all) {
assert.equal(rowOrigin(row({ group }), all), 'this-run', group)
}
})
test('no row is not a row from an earlier run', () => {
assert.equal(rowOrigin(undefined, ['auto']), 'unattributed')
assert.equal(rowOrigin(null, ['auto']), 'unattributed')
})
test('the three origins produce three distinct hints', () => {
const hints = new Set(
(['this-run', 'carried', 'unattributed'] as const).map((o) => originStamp(o, CLOCK).hint),
)
assert.equal(hints.size, 3, 'three different situations, three different explanations')
})
+168
View File
@@ -0,0 +1,168 @@
/* ListPicker — the additions on top of SrcPicker.css.
*
* Everything structural (the machined slot, the flush-docked popover, the shelf
* strips, the sunken rows, the chip shell and its ×) is SrcPicker's and is reused
* verbatim; only what this picker says differently lives here:
*
* .lstp-lane / .lstp-step the two ranked lanes inside one slot, carrying the
* engine step number that ties the control back to the
* order rail on the device card;
* .lstp-dot / .lstp-load the load reading — the one colour on a list chip. It
* is driven by data-load, never by "is it attached":
* on = matching, warn = loaded but empty, crit = not
* loaded / no such list, unknown = the engine has not
* said, which is DIM and never green.
*/
/* --- ranked lanes --------------------------------------------------------- */
.lstp-lane {
display: inline-flex;
flex-wrap: wrap;
align-items: center;
gap: 5px;
min-width: 0;
padding: 1px 4px 1px 1px;
border-radius: 6px;
background: color-mix(in srgb, var(--panel) 55%, transparent);
}
/* The step number. A machined stamp, not a bullet: it is the same number printed
* on the card's order rail, so the two read as one instrument. */
.lstp-step {
flex: 0 0 auto;
display: inline-flex;
align-items: center;
justify-content: center;
min-width: 15px;
height: 15px;
padding: 0 3px;
border: 1px solid var(--groove);
border-radius: 3px;
background: var(--sink);
box-shadow: 0 1px 1px var(--shadow) inset;
color: var(--dim);
font-family: var(--font-mono, ui-monospace, monospace);
font-size: 9.5px;
font-weight: 700;
line-height: 1;
}
/* --- chips ---------------------------------------------------------------- */
/* A typed entry is hand-made: amber dot, same as SrcPicker's custom value. It
* carries NO kind tag — the marker it was typed with is already the first word of
* the label, and "full:discord.com FULL" says the same thing twice. */
/* An attached list carries the engine's verdict, and nothing else. */
.lstp-dot {
background: var(--faint);
}
[data-load='on'] > .lstp-dot,
.lstp-dot[data-load='on'] {
background: var(--led-on);
box-shadow: 0 0 4px var(--led-on);
}
[data-load='warn'] > .lstp-dot,
.lstp-dot[data-load='warn'] {
background: var(--amber);
box-shadow: 0 0 4px var(--amber);
}
[data-load='crit'] > .lstp-dot,
.lstp-dot[data-load='crit'] {
background: var(--crit);
box-shadow: 0 0 4px var(--crit);
}
[data-load='unknown'] > .lstp-dot,
.lstp-dot[data-load='unknown'] {
background: var(--faint);
box-shadow: none;
}
/* The verdict, in full. It is never truncated: "not loaded" clipped to
* "NOT LOAD…" is the one word on this card a parent has to be able to read. */
.lstp-load {
flex: 0 0 auto;
white-space: nowrap;
color: var(--faint);
font-family: var(--font-mono, ui-monospace, monospace);
font-size: 9.5px;
letter-spacing: 0.04em;
text-transform: uppercase;
}
[data-load='on'] > .lstp-load,
.lstp-load[data-load='on'] {
color: var(--led-on);
}
[data-load='warn'] > .lstp-load,
.lstp-load[data-load='warn'] {
color: var(--amber);
}
[data-load='crit'] > .lstp-load,
.lstp-load[data-load='crit'] {
color: var(--crit);
}
.lstp-chip-list[data-load='crit'] {
border-color: color-mix(in srgb, var(--crit) 45%, var(--groove));
}
.lstp-chip-list[data-load='warn'] {
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
}
/* --- popover rows --------------------------------------------------------- */
.lstp-list {
max-height: 168px;
}
.lstp-row {
gap: 6px;
}
.lstp-row-detail {
flex: 1 1 auto;
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
color: var(--faint);
font-size: 10px;
}
.lstp-row-tags {
flex: 0 0 auto;
display: inline-flex;
align-items: center;
gap: 5px;
}
.lstp-tag {
padding: 0 4px;
border: 1px solid var(--groove);
border-radius: 3px;
color: var(--dim);
font-family: var(--font-mono, ui-monospace, monospace);
font-size: 9px;
letter-spacing: var(--track-label, 0.08em);
text-transform: uppercase;
}
/* --- notes ---------------------------------------------------------------- */
.lstp-hint,
.lstp-foot {
margin: 0;
padding: 6px 9px;
border-top: 1px solid var(--groove);
background: var(--panel);
font-family: var(--font-sans, system-ui, sans-serif);
font-size: 11px;
line-height: 1.5;
color: var(--dim);
}
.lstp-hint .mono {
color: var(--ink);
font-size: 10.5px;
}
.lstp-foot {
color: var(--faint);
}
@media (max-width: 560px) {
.lstp-row-detail {
display: none;
}
}
+394
View File
@@ -0,0 +1,394 @@
import { useEffect, useMemo, useRef, useState } from 'react'
import { describeDomainEntry, parseDomainEntry } from '../deviceLists'
import type { ListLoad } from '../deviceLists'
import './SrcPicker.css'
import './ListPicker.css'
/**
* ListPicker — the Faceplate chooser for ONE direction of a device's domain
* policy: the named lists attached to it plus the domains typed on it, in a
* single slot.
*
* It borrows SrcPicker's LANGUAGE wholesale (a machined slot of chips, a popover
* docked flush under it with hairline shelves, a Custom row at the bottom, Esc and
* click-outside, a dot encoding the chip's kind) and its stylesheet — but not its
* code. SrcPicker is welded to `useSrcOptions()`, to IP/CIDR validation, and to an
* empty state that reads "everyone · all LAN clients"; on a BLOCK list that
* sentence would mean the exact opposite of the truth. Generalising it would have
* produced a component with two of everything and a props list nobody reads.
*
* ## What the slot has to make legible
*
* The two chip kinds are not equal, and they are not adjacent in the engine's
* order either. A device is evaluated:
*
* 1 allow typed · 2 block typed · 3 allow attached · 4 block attached · 5 network
*
* so the typed chips and the attached chips inside ONE control sit two steps
* apart. The slot therefore renders them as two labelled lanes carrying their own
* step number, which is what ties this control back to the order rail above it on
* the card. Sorting them the other way round would draw a precedence that does not
* exist.
*
* ## What a chip is not allowed to say
*
* A list chip reports what the ENGINE says about that list (via
* /api/ruleset/status, through `listLoad`), never the fact that someone attached
* it. A list whose URL is unreachable, whose file is missing, or whose geosite
* category resolved to nothing is a parental control that does not work; it is
* drawn `not loaded`, in crit, and a list the engine has not mentioned at all is
* drawn `load unknown` rather than green.
*/
export type ListPickerKind = 'block' | 'allow'
/** One attachable named list, as the page knows it. */
export interface ListOption {
/** The `config blocklist` / `config allowlist` section name — the stored value. */
name: string
/** Where its content comes from: "url · big.oisd.nl", "3 domains", "geosite · telegram". */
detail: string
/** Blocklist/Allowlist.Enabled — whether it is in the NETWORK-wide filter.
* It does NOT gate this device: attaching the list is the switch here. */
networkEnabled: boolean
/** Blocklist.Response — the reply a blocked name gets. Block lists only. */
response?: string
/** What the running engine reports about it. */
load: ListLoad
}
export interface ListPickerProps {
kind: ListPickerKind
/** Attached list names — Device.Blocklists / Device.Allowlists. */
lists: string[]
/** Hand-typed entries — Device.Block / Device.Allow. */
domains: string[]
/** Every list that could be attached, in config order. */
options: ListOption[]
/** The engine step the TYPED lane is (1 for allow, 2 for block). */
typedStep: number
/** The engine step the ATTACHED lane is (3 for allow, 4 for block). */
listStep: number
disabled?: boolean
ariaLabel: string
onListsChange: (v: string[]) => void
onDomainsChange: (v: string[]) => void
}
const WORD: Record<ListPickerKind, { noun: string; verb: string; shelf: string }> = {
block: { noun: 'block', verb: 'Blocked', shelf: 'Blocklists' },
allow: { noun: 'allow', verb: 'Allowed', shelf: 'Allowlists' },
}
const EMPTY_TEXT: Record<ListPickerKind, string> = {
block: 'nothing blocked for this device',
allow: 'nothing forced through for this device',
}
const CUSTOM_PLACEHOLDER: Record<ListPickerKind, string> = {
block: 'example.com or keyword:tiktok',
allow: 'school.example.edu',
}
/** Said at the moment of choosing, because both facts change what the operator is
* about to do — and neither is visible from the chip afterwards. */
const FOOTNOTE: Record<ListPickerKind, string> = {
block:
'Attaching a list runs it for this device even when the list is off for the network. The reply a blocked name gets comes from the list, not from the device.',
allow:
'An attached allow list is terminal: everything it covers is also lifted out of the network blocklists for this device. A big list here removes a lot of filtering.',
}
/** The reading for a name the config no longer has. Built once — it never varies. */
const MISSING_LOAD: ListLoad = {
tone: 'crit',
tag: 'no such list',
detail: 'This device points at a list that is not in the config — it filters nothing.',
ruleCount: 0,
}
export function ListPicker({
kind,
lists,
domains,
options,
typedStep,
listStep,
disabled,
ariaLabel,
onListsChange,
onDomainsChange,
}: ListPickerProps) {
const [open, setOpen] = useState(false)
const [custom, setCustom] = useState('')
const [customErr, setCustomErr] = useState<string | null>(null)
const rootRef = useRef<HTMLDivElement>(null)
const fieldRef = useRef<HTMLDivElement>(null)
const customRef = useRef<HTMLInputElement>(null)
const byName = useMemo(() => new Map(options.map((o) => [o.name, o])), [options])
const attached = useMemo(() => new Set(lists), [lists])
const free = useMemo(() => options.filter((o) => !attached.has(o.name)), [options, attached])
// A stored name with no matching option is a device pointing at a deleted list.
// It keeps its chip and says so, rather than vanishing from the card.
const listChips = useMemo(
() =>
lists.map((name) => {
const opt = byName.get(name)
return { name, opt, load: opt ? opt.load : MISSING_LOAD }
}),
[lists, byName],
)
const domainChips = useMemo(
() => domains.map((d) => ({ value: d, ...describeDomainEntry(d) })),
[domains],
)
useEffect(() => {
if (!open) return
const onDown = (e: MouseEvent) => {
if (!rootRef.current?.contains(e.target as Node)) setOpen(false)
}
const onKey = (e: KeyboardEvent) => {
if (e.key === 'Escape') {
setOpen(false)
fieldRef.current?.focus()
}
}
document.addEventListener('mousedown', onDown)
document.addEventListener('keydown', onKey)
return () => {
document.removeEventListener('mousedown', onDown)
document.removeEventListener('keydown', onKey)
}
}, [open])
const w = WORD[kind]
const attachList = (name: string) => {
if (attached.has(name)) return
onListsChange([...lists, name])
}
const detachList = (name: string) => onListsChange(lists.filter((x) => x !== name))
const removeDomain = (v: string) => onDomainsChange(domains.filter((x) => x !== v))
const addCustom = () => {
const parsed = parseDomainEntry(custom)
if (!parsed.ok) {
setCustomErr(parsed.reason)
return
}
if (domains.some((d) => d.toLowerCase() === parsed.value)) {
setCustomErr(`${parsed.value} is already on this list.`)
return
}
onDomainsChange([...domains, parsed.value])
setCustom('')
setCustomErr(null)
customRef.current?.focus()
}
const toggleOpen = () => {
if (disabled) return
setOpen((o) => !o)
}
const onFieldKeyDown = (e: React.KeyboardEvent) => {
if (disabled) return
if (e.key === 'Enter' || e.key === ' ' || e.key === 'ArrowDown') {
e.preventDefault()
setOpen(true)
}
}
const empty = listChips.length === 0 && domainChips.length === 0
return (
<div className="srcp lstp" ref={rootRef}>
<div
ref={fieldRef}
className={disabled ? 'srcp-field disabled' : 'srcp-field'}
role="button"
tabIndex={disabled ? -1 : 0}
aria-haspopup="dialog"
aria-expanded={open}
aria-label={ariaLabel}
aria-disabled={disabled || undefined}
onClick={toggleOpen}
onKeyDown={onFieldKeyDown}
>
{empty ? (
<span className="srcp-empty">{EMPTY_TEXT[kind]}</span>
) : (
<>
{domainChips.length > 0 && (
<span
className="lstp-lane"
role="group"
aria-label={`Step ${typedStep} — domains typed here`}
>
<span className="lstp-step" aria-hidden="true">
{typedStep}
</span>
{domainChips.map((c) => (
<span key={c.value} className="srcp-chip lstp-chip-typed" title={c.detail}>
<span className="srcp-dot srcp-dot-custom" aria-hidden="true" />
<span className="srcp-chip-name mono">{c.label}</span>
<button
type="button"
className="srcp-chip-x"
aria-label={`Remove ${c.value} from the ${w.noun} list`}
onClick={(e) => {
e.stopPropagation()
removeDomain(c.value)
}}
disabled={disabled}
>
×
</button>
</span>
))}
</span>
)}
{listChips.length > 0 && (
<span
className="lstp-lane"
role="group"
aria-label={`Step ${listStep} — attached ${w.shelf.toLowerCase()}`}
>
<span className="lstp-step" aria-hidden="true">
{listStep}
</span>
{listChips.map((c) => (
<span
key={c.name}
className="srcp-chip lstp-chip-list"
data-load={c.load.tone}
title={`${c.name} — ${c.load.detail}`}
>
<span className="srcp-dot lstp-dot" aria-hidden="true" />
<span className="srcp-chip-name mono">{c.name}</span>
<span className="lstp-load">{c.load.tag}</span>
<button
type="button"
className="srcp-chip-x"
aria-label={`Detach list ${c.name} — ${c.load.detail}`}
onClick={(e) => {
e.stopPropagation()
detachList(c.name)
}}
disabled={disabled}
>
×
</button>
</span>
))}
</span>
)}
</>
)}
<span className="srcp-caret" aria-hidden="true">
▾
</span>
</div>
{open && !disabled && (
<div className="srcp-pop" role="dialog" aria-label={`${ariaLabel} — choose lists and domains`}>
{/* Named lists ---------------------------------------------------- */}
<div className="srcp-shelf" aria-hidden="true">
<span>
{w.shelf} · step {listStep}
</span>
<span className="srcp-shelf-n">{free.length || '—'}</span>
</div>
{free.length > 0 ? (
<ul className="srcp-list lstp-list">
{free.map((o) => (
<li key={o.name}>
<button
type="button"
className="srcp-row lstp-row"
onClick={() => attachList(o.name)}
title={o.load.detail}
>
<span className="srcp-dot lstp-dot" data-load={o.load.tone} aria-hidden="true" />
<span className="srcp-row-name">{o.name}</span>
<span className="lstp-row-detail mono">{o.detail}</span>
<span className="lstp-row-tags">
{o.response === 'zero' && <span className="lstp-tag">0.0.0.0</span>}
{!o.networkEnabled && (
<span className="lstp-tag" title="Off for the network — attaching it still runs it here">
network off
</span>
)}
<span className="lstp-load" data-load={o.load.tone}>
{o.load.tag}
</span>
</span>
</button>
</li>
))}
</ul>
) : (
<p className="srcp-none">
{options.length
? `every ${w.shelf.toLowerCase().replace(/s$/, '')} is already attached`
: `no ${w.shelf.toLowerCase()} configured — add one on the DNS page`}
</p>
)}
{/* Typed domains -------------------------------------------------- */}
<div className="srcp-shelf" aria-hidden="true">
<span>
Type a domain · step {typedStep}
</span>
</div>
<div className="srcp-custom">
<input
ref={customRef}
className="srcp-custom-input mono"
type="text"
value={custom}
onChange={(e) => {
setCustom(e.target.value)
if (customErr) setCustomErr(null)
}}
onKeyDown={(e) => {
if (e.key === 'Enter') {
e.preventDefault()
addCustom()
}
}}
placeholder={CUSTOM_PLACEHOLDER[kind]}
aria-label={`Domain to ${w.noun} for this device`}
autoComplete="off"
spellCheck={false}
/>
<button
type="button"
className="srcp-custom-add"
onClick={addCustom}
disabled={!custom.trim()}
>
{kind === 'block' ? 'Block' : 'Allow'}
</button>
</div>
{customErr && (
<p className="srcp-note" role="alert">
{customErr}
</p>
)}
<p className="lstp-hint">
A bare entry covers the domain and its subdomains.{' '}
<span className="mono">full:</span> one exact name,{' '}
<span className="mono">keyword:</span> any host containing it.
</p>
<p className="lstp-foot">{FOOTNOTE[kind]}</p>
</div>
)}
</div>
)
}
+19
View File
@@ -2,6 +2,25 @@
* former page-local .set-select in Settings.css so promoting it to a shared
* component changed nothing visually. */
.fp-select {
/* A <select> shrink-wraps to its WIDEST OPTION and, as a flex/grid item,
* refuses to shrink below that min-content width. Nothing capped it here, so
* one long label — "Auto — country codes from SagerNet, the rest from
* Loyalsoldier" on Settings → Geo data provider — measured 501px inside a
* 375px viewport and gave the whole PAGE a horizontal scrollbar (measured:
* documentElement.scrollWidth 559 vs clientWidth 375 at a 390px window).
*
* The labels are load-bearing — they are where the coverage and cost
* difference between providers is stated — so the control is capped rather
* than the sentence shortened; truncation is the browser's job once there is a
* definite width to truncate against. Both declarations are needed: max-width
* bounds it against the containing block, min-width lets it actually shrink
* there instead of insisting on min-content.
*
* This lives on the component because every page that puts a long label in a
* dropdown inherits the same bug; Settings.css and Networks.css each carried
* their own narrow copy of this fix before it. */
max-width: 100%;
min-width: 0;
padding: 8px 10px;
border: 1px solid var(--groove);
border-radius: 7px;
+2
View File
@@ -22,5 +22,7 @@ export type { ConfirmDialogProps, ConfirmOptions, ConfirmTone } from './ConfirmD
export { Clock } from './Clock'
export { CatSuggest } from './CatSuggest'
export { SrcPicker } from './SrcPicker'
export { ListPicker } from './ListPicker'
export type { ListOption, ListPickerKind, ListPickerProps } from './ListPicker'
export { ThemeSwitch } from './ThemeSwitch'
export { usePrefersReducedMotion } from './usePrefersReducedMotion'
+121
View File
@@ -0,0 +1,121 @@
// A connection the router KILLED must not read as traffic it carried.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// THE DEFECT THIS EXISTS FOR. `outbound:"block"` is a legitimate rule target and
// the value of route.Final on a fail-closed box (shater/generate/route.go), so it
// is what a kill-switch drop looks like in the connection log. The row was drawn
// as plain mono text, indistinguishable from `outbound:"nl-reality-1"` — on the
// same page where a DNS row about the same host gets a crit rail and a BLOCK
// mark. The connection log is the first place a "this site does not open" report
// is read, and the row that IS the answer looked like carried traffic.
//
// WHAT THESE PROTECT:
//
// 1. KILLED AND CARRIED MUST NOT RENDER ALIKE. Asserted on the state, the label
// AND the tooltip, because a single shared field would let two of the three
// collapse unnoticed.
// 2. THE CONTROL. The same helper must produce the ordinary reading too, or "the
// two differ" is satisfiable by a function that calls everything killed.
// 3. NOT RECORDED IS ITS OWN STATE. An absent outbound is neither a kill nor a
// carry, and it may not be drawn as either.
// 4. THE MATCH IS EXACT. `block` is reserved by an exact string comparison in
// the generator (generate/outbound.go, generate/group.go skip a node or group
// named exactly that), so a node someone named `Block` IS a real destination
// and calling it a kill would be a lie about where the traffic went.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { ConnLogEntry } from './api.ts'
import { connFate, connRule } from './logRoute.ts'
const conn = (over: Partial<ConnLogEntry>): ConnLogEntry => ({
unix: 1_700_000_000,
src_ip: '192.168.1.42',
src_name: 'ipad-kids',
dest: 'blocked-ad.example',
dest_ip: '203.0.113.77',
port: 80,
network: 'tcp',
proto: 'http',
outbound: 'block',
seq: 12,
rule_kind: 'matched',
rule: 'rule_set=rs-ads',
chain: ['block'],
...over,
})
const killed = connFate(conn({ outbound: 'block' }))
const carried = connFate(conn({ outbound: 'nl-reality-1', chain: ['nl-reality-1'] }))
test('a killed connection and a carried one do not render alike', () => {
assert.notEqual(killed.fate, carried.fate, 'both connections got the same state')
assert.notEqual(killed.label, carried.label, 'both connections got the same label')
assert.notEqual(killed.title, carried.title, 'both connections got the same tooltip')
})
test('the killed row SAYS the router did not carry it', () => {
assert.equal(killed.fate, 'killed')
assert.match(killed.label, /killed/i)
assert.match(killed.title, /did NOT carry/)
// It must not name the decision — that is connRule's job beside it, and a second
// copy could only disagree with the first.
assert.doesNotMatch(killed.title, /kill-switch|rule_set/)
})
test('CONTROL: the same helper produces the ordinary carried reading', () => {
// Without this, "killed differs from carried" is satisfied by a function that
// marks every row killed.
assert.equal(carried.fate, 'carried')
assert.match(carried.title, /nl-reality-1/, 'the carried reading does not name the exit it took')
})
test('carried claims a ROUTE, never a result — the log records neither success nor health', () => {
// The engine logs the routing decision. A row that says "carried" must not be
// readable as "the flow worked", which is the failure mode the DNS log already
// paid for once (a proxied lookup that timed out drawn as the healthiest row).
assert.match(carried.title, /which WAY it went/)
// The only mention of success is a DENIAL of one. Asserted as the negated
// phrase rather than as a banned word, because banning the word would also ban
// the sentence that does the work.
assert.match(carried.title, /not whether the flow then succeeded/)
})
test('no outbound recorded is a third state, and claims nothing either way', () => {
const blank = connFate(conn({ outbound: '' }))
assert.equal(blank.fate, 'unrecorded')
for (const other of [killed, carried]) {
assert.notEqual(blank.fate, other.fate)
assert.notEqual(blank.label, other.label)
assert.notEqual(blank.title, other.title)
}
assert.match(blank.title, /NOT a claim that it was carried/)
assert.match(blank.title, /NOT a claim that it was blocked/)
})
test('the reserved tag is matched EXACTLY — a node named “Block” is a destination', () => {
// generate/outbound.go and generate/group.go reserve the tag with `==`, so a
// node named `Block` is emitted under that tag and really is where the traffic
// went. Folding the comparison here would report a kill that never happened.
assert.equal(connFate(conn({ outbound: 'Block' })).fate, 'carried')
assert.equal(connFate(conn({ outbound: 'block-list-egress' })).fate, 'carried')
assert.equal(connFate(conn({ outbound: 'noblock' })).fate, 'carried')
// …and surrounding whitespace is not a way to hide a kill.
assert.equal(connFate(conn({ outbound: ' block ' })).fate, 'killed')
})
test('the fate axis is independent of the rule axis — both readings survive together', () => {
// A kill-switch drop and a rule-target drop are both `killed`, and the DIFFERENCE
// between them is carried by connRule alone. If the fate ever started answering
// that too, the row would state it twice and could state it twice differently.
const byRule = conn({ outbound: 'block', rule_kind: 'matched', rule: 'rule_set=rs-ads' })
const byDefault = conn({ outbound: 'block', rule_kind: 'default', rule: '', chain: ['block'] })
assert.equal(connFate(byRule).fate, 'killed')
assert.equal(connFate(byDefault).fate, 'killed')
assert.equal(connFate(byRule).title, connFate(byDefault).title)
// …while the rule axis still tells them apart.
assert.notEqual(connRule(byRule).kind, connRule(byDefault).kind)
})
+81
View File
@@ -0,0 +1,81 @@
// The default route, named.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// What these protect: the Routing page's lead and empty state now NAME the route
// unmatched traffic takes, instead of calling it "the default route" and leaving
// it at that. The case that matters is the fresh install — no rules, kill-switch
// closed — where the answer is `block`, i.e. the LAN has no internet. That is the
// state the kill-switch alarm sends people to this page in, so getting it wrong
// means telling someone whose network is down that everything is following the
// default route, which is true and useless.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { effectiveTarget, liveDefaultRoute } from './defaultRoute.ts'
import type { TargetedRule } from './defaultRoute.ts'
/** Stand-in for Routing's isLiveDefault: conditionless, on, not shadowed. */
const live = (names: string[]) => (r: TargetedRule) => names.includes(r.Name)
test('no rules + fail-closed ⇒ the default route is block, and no rule owns it', () => {
const d = liveDefaultRoute([], live([]), 'closed')
assert.equal(d.target, 'block')
assert.equal(d.rule, null)
})
test('no rules + fail-open ⇒ direct', () => {
const d = liveDefaultRoute([], live([]), 'open')
assert.equal(d.target, 'direct')
assert.equal(d.rule, null)
})
test('an unreadable kill-switch value falls to the blocking side, like the daemon', () => {
// Callers pass planeState.killSwitchClosed's verdict, and its rule is
// "everything that is not literally open is closed". Anything but 'open' here
// must therefore not become 'direct'.
for (const ks of ['', 'closed', 'Closed', 'nonsense']) {
assert.equal(liveDefaultRoute([], live([]), ks).target, 'block', ks)
}
})
test('a live catch-all owns the default and is named', () => {
const rules: TargetedRule[] = [
{ Name: 'stream', Target: 'direct' },
{ Name: 'fallback', Target: 'group:eu' },
]
const d = liveDefaultRoute(rules, live(['fallback']), 'closed')
assert.equal(d.target, 'group:eu')
assert.equal(d.rule, 'fallback')
})
test('among several catch-alls the LAST one wins — the generator overwrites Final in order', () => {
const rules: TargetedRule[] = [
{ Name: 'first', Target: 'block' },
{ Name: 'last', Target: 'direct' },
]
const d = liveDefaultRoute(rules, live(['first', 'last']), 'closed')
assert.equal(d.rule, 'last')
assert.equal(d.target, 'direct')
})
test('a catch-all that is not in force does not own the default', () => {
const rules: TargetedRule[] = [{ Name: 'off', Target: 'direct' }]
// Switched off, shadowed, or overridden by the active profile — the predicate
// is the caller's, and when it says no, the kill-switch decides.
const d = liveDefaultRoute(rules, live([]), 'closed')
assert.equal(d.target, 'block')
assert.equal(d.rule, null)
})
test('a rule with only a bare Egress routes there — it is a target too', () => {
assert.equal(effectiveTarget({ Name: 'x', Egress: 'wan2' }), 'egress:wan2')
assert.equal(effectiveTarget({ Name: 'x', Target: 'group:eu', Egress: 'wan2' }), 'group:eu')
assert.equal(effectiveTarget({ Name: 'x' }), 'direct')
assert.equal(effectiveTarget({ Name: 'x', Target: ' ' }), 'direct')
// …and the default route reports it, rather than reporting `direct` at a row
// the panel draws as egress:wan2.
const d = liveDefaultRoute([{ Name: 'wan2only', Egress: 'wan2' }], live(['wan2only']), 'closed')
assert.equal(d.target, 'egress:wan2')
})
+80
View File
@@ -0,0 +1,80 @@
/**
* Where traffic that matched no rule actually goes — the one fact the routing
* page never stated.
*
* The page said "traffic that reaches the bottom follows the default route", and
* the empty state said "No rules — all traffic follows the default route". Both
* are true and neither is an answer. On a fresh install the default route is
* `block`: with no catch-all rule the generator sets `final := tagBlock`, and only
* an open kill-switch swaps that for `direct` (generate/route.go). So "no rules"
* means "the LAN has no internet", and the page described it as a routine
* fallback.
*
* The omission is expensive in one specific place: the kill-switch alarm SENDS
* PEOPLE HERE — its banner says "Add a default rule on the Routing page". Someone
* whose network has just gone down arrives, reads that everything follows the
* default route, and cannot learn from this screen that the default route is what
* took it down.
*
* It lives outside `pages/Routing.tsx` so it can be tested: the panel's runner is
* `node --test src/*.test.ts`, plain modules only (same reason as ruleset.ts).
*/
/** The rule shape this module needs. Structurally a subset of Routing's RRule. */
export interface TargetedRule {
Name: string
Target?: string
Egress?: string
}
/**
* A rule's effective routing target: `Target` wins, and a bare `Egress` is a
* target too.
*
* The last clause is the daemon's (model.EffectiveRuleTarget, mirrored by
* generate's effectiveRuleTarget). A rule carrying only `option egress wan2` used
* to fall through to resolveTarget("") == "direct" and leave over the DEFAULT WAN
* while the panel showed `egress:wan2` — a silent mis-route on exactly the
* multi-WAN setups the field exists for. Reading it the same way here is what
* keeps the panel's answer and the router's answer the same answer.
*/
export function effectiveTarget(r: TargetedRule): string {
if (r.Target && r.Target.trim()) return r.Target.trim()
if (r.Egress && r.Egress.trim()) return `egress:${r.Egress.trim()}`
return 'direct'
}
/** What the router does with unmatched traffic, and which rule decided it. */
export interface DefaultRoute {
/** The outbound tag: a rule's target, or the kill-switch's `block` / `direct`. */
target: string
/** The rule that owns route.Final, or null when the kill-switch decides. */
rule: string | null
}
/**
* The default route in force right now.
*
* `isLive` is the caller's own "is this rule the engine's route.Final" predicate
* (Routing's isLiveDefault: conditionless, not shadowed, in force after the
* active profile has had its say). It is a parameter and not a reimplementation
* because the badge on the row, the delete confirmation and this sentence must
* name the SAME rule — three copies of that predicate is how they came to name
* three different ones.
*
* The scan runs backwards because the generator overwrites `Final` as it walks
* the list in order: among several conditionless rules the LAST one wins.
*/
export function liveDefaultRoute<T extends TargetedRule>(
rules: readonly T[],
isLive: (r: T) => boolean,
killSwitch: string,
): DefaultRoute {
for (let i = rules.length - 1; i >= 0; i--) {
if (isLive(rules[i])) return { target: effectiveTarget(rules[i]), rule: rules[i].Name }
}
// No rule claims it, so generate/route.go's own default stands: `block`, unless
// the kill-switch is open. Normalised the daemon's way (planeState.killSwitchClosed
// is the canonical comparison; callers pass its verdict as 'open' / 'closed').
return { target: killSwitch === 'open' ? 'direct' : 'block', rule: null }
}
+323
View File
@@ -0,0 +1,323 @@
// Per-device domain policy — the three things this panel is not allowed to get
// wrong about a parental control.
//
// Run with `npm test` (node's built-in runner + native type stripping).
// deviceLists.ts has no runtime imports, so this runs against the real module.
//
// THE CASES THESE WERE WRITTEN FOR, each one a control that looked applied and
// was not:
//
// 1. The form demanded /^[a-z0-9.-]+$/, so a colon could not be typed and the
// engine's own `full:` / `suffix:` / `keyword:` vocabulary was unreachable
// from the panel. Opening the charset instead of listing the prefixes would
// have been worse: a lone `keyword:` is strings.Contains(host, "") — true
// for EVERY host — and a bare `keyword:` in a child's Block list takes that
// device off the internet with no error anywhere.
//
// 2. A chip said "attached", which is not evidence of anything. A list whose
// fetch never succeeded blocks nothing, and has to look like it.
//
// 3. "Blocklists below are configured but inactive" stopped being true the day
// a device could attach one — attaching IS the switch for that device.
//
// Each test therefore has a POSITIVE and a NEGATIVE half where one exists: a
// reading that only ever says "bad" proves as little as one that only says "good".
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
attachedDeviceNames,
compactCount,
describeDomainEntry,
devicesWithAttachedLists,
dnsFilterOffNote,
listLoad,
listRowState,
nameList,
parseDomainEntry,
} from './deviceLists.ts'
import type { Device, RulesetStatus } from './api.ts'
// ---------------------------------------------------------------------------
// 1. What the form accepts
// ---------------------------------------------------------------------------
test('the three engine prefixes are typeable — the panel is no longer poorer than the engine', () => {
const full = parseDomainEntry('full:Discord.com')
assert.deepEqual(full, { ok: true, value: 'full:discord.com', kind: 'full' })
const kw = parseDomainEntry('keyword:TikTok')
assert.deepEqual(kw, { ok: true, value: 'keyword:tiktok', kind: 'keyword' })
// `suffix:` and a bare entry build the SAME matcher in generate/dnsfilter.go,
// so they must canonicalise to one stored value — otherwise the two spellings
// read as two different rules and the duplicate check misses.
const suf = parseDomainEntry('suffix:Example.COM')
assert.deepEqual(suf, { ok: true, value: 'example.com', kind: 'suffix' })
assert.equal(parseDomainEntry('example.com').ok && parseDomainEntry('example.com').value, 'example.com')
assert.equal(parseDomainEntry('.example.com').ok && parseDomainEntry('.example.com').value, 'example.com')
assert.equal(parseDomainEntry('*.example.com').ok && parseDomainEntry('*.example.com').value, 'example.com')
})
test('a bare marker is refused, and the keyword one is refused for the reason that matters', () => {
const kw = parseDomainEntry('keyword:')
assert.equal(kw.ok, false)
// Not a generic "invalid entry": an empty keyword matches every host, so the
// message has to say what accepting it would do.
assert.match(kw.ok === false ? kw.reason : '', /every host/i)
assert.match(kw.ok === false ? kw.reason : '', /off the internet/i)
for (const bad of ['full:', 'suffix:', 'full: ', 'keyword: ']) {
assert.equal(parseDomainEntry(bad).ok, false, `${bad} must be refused`)
}
// A lone dot is not an entry either — bareDomain strips it to nothing.
assert.equal(parseDomainEntry('.').ok, false)
assert.equal(parseDomainEntry('').ok, false)
})
test('an unrecognised prefix is refused and named, because a domain cannot contain a colon', () => {
const r = parseDomainEntry('regexp:.*ads.*')
assert.equal(r.ok, false)
assert.match(r.ok === false ? r.reason : '', /regexp:/)
assert.match(r.ok === false ? r.reason : '', /full:, suffix: or keyword:/)
const g = parseDomainEntry('geosite:youtube')
assert.equal(g.ok, false)
assert.match(g.ok === false ? g.reason : '', /geosite:/)
})
test('an address is refused as an address, not mistaken for a prefixed entry', () => {
for (const addr of ['192.168.1.10', '10.0.0.0/8', '2001:db8::1', 'fe80::1']) {
const r = parseDomainEntry(addr)
assert.equal(r.ok, false, `${addr} must be refused`)
assert.match(r.ok === false ? r.reason : '', /IP address/i, addr)
}
})
test('a keyword may only be hostname material, and a stored entry reads back as what it does', () => {
assert.equal(parseDomainEntry('keyword:has space').ok, false)
assert.equal(parseDomainEntry('keyword:a:b').ok, false)
assert.equal(describeDomainEntry('example.com').detail, 'example.com and its subdomains')
assert.equal(describeDomainEntry('full:example.com').kind, 'full')
assert.match(describeDomainEntry('full:example.com').detail, /subdomains are not matched/)
assert.equal(describeDomainEntry('keyword:tiktok').kind, 'keyword')
assert.match(describeDomainEntry('keyword:tiktok').detail, /containing/)
// An unknown prefix already in the config is still shown, verbatim — the chip
// is not the place to make a stored value disappear.
assert.equal(describeDomainEntry('weird:thing').label, 'weird:thing')
})
// ---------------------------------------------------------------------------
// 2. Did the attached list load?
// ---------------------------------------------------------------------------
const st = (o: Partial<RulesetStatus>): RulesetStatus => ({
tag: 'bl-x',
name: 'x',
category: '',
kind: 'blocklist',
remote: true,
last_updated: '',
interval_seconds: 86_400,
rule_count: 0,
...o,
})
test('the load reading gives a POSITIVE result — a fetched list with rules is green and counted', () => {
const r = listLoad([st({ last_updated: '2026-07-26T10:00:00Z', rule_count: 218_431 })], true)
assert.equal(r.tone, 'on')
assert.equal(r.tag, '218k')
assert.equal(r.ruleCount, 218_431)
// CONTROL that the green branch is reachable at all, and on a real payload:
// only a remote set carries a count, so this is the shape the daemon really
// sends for a list that loaded and holds entries.
const small = listLoad([st({ last_updated: '2026-07-26T10:00:00Z', rule_count: 3 })], true)
assert.equal(small.tone, 'on')
assert.equal(small.tag, '3')
})
test('an INLINE list has no entry count, and "not published" is not "empty"', () => {
// THE DEFECT, AND THE BROKEN INSTRUMENT THAT HID IT. engine.go fills
// RuleSetStat.RuleCount only for *rule.RemoteRuleSet; `LocalRuleSet.RuleCount`
// does not exist anywhere in the tree. So an inline list ALWAYS arrives with
// rule_count 0, and the panel called a perfectly working list
// "empty — nothing matches".
//
// The old test for this asked `st({remote:false, rule_count:3})` — a record the
// daemon cannot produce. A fixture of an impossible state is not a control; it
// is a green light wired to nothing.
const local = listLoad([st({ remote: false, last_updated: '', rule_count: 0 })], true)
assert.equal(local.tone, 'unknown', 'unlit: not a fault, and not a confirmation either')
assert.equal(local.tag, 'size unknown')
assert.notEqual(local.tag, 'empty', 'the list may hold five hundred entries — nobody said')
assert.doesNotMatch(local.detail, /nothing matches/)
assert.notEqual(local.tone, 'on', 'and an unknown is never drawn as healthy')
})
test('CONTROL: a REMOTE list that really is empty still says so', () => {
// Without this, "a zero count is unknown" would also be satisfied by deleting
// the empty state outright — and a fetched list that genuinely came back with
// no entries is a real finding the operator needs.
const empty = listLoad([st({ remote: true, last_updated: '2026-07-26T10:00:00Z', rule_count: 0 })], true)
assert.equal(empty.tone, 'warn')
assert.equal(empty.tag, 'empty')
assert.notEqual(empty.tone, listLoad([st({ remote: false, rule_count: 0 })], true).tone)
})
test('a MIXED group is counted as a floor, not as a whole', () => {
// One name resolving to a remote set and an inline one: the sum covers only the
// part that reports. Printing it bare would understate the list while looking
// exact.
const mixed = listLoad(
[
st({ category: 'ads', last_updated: '2026-07-26T10:00:00Z', rule_count: 1_284 }),
st({ category: '', remote: false, last_updated: '', rule_count: 0 }),
],
true,
)
assert.equal(mixed.tone, 'on')
assert.equal(mixed.tag, '1,284+', 'the plus is the whole point')
assert.match(mixed.detail, /At least/)
})
test('and a NEGATIVE one — never fetched, empty, unreported and unknown are four different things', () => {
assert.equal(listLoad([st({ last_updated: '', rule_count: 0 })], true).tone, 'crit')
assert.equal(listLoad([st({ last_updated: '', rule_count: 0 })], true).tag, 'not loaded')
assert.equal(listLoad([st({ last_updated: '2026-07-26T10:00:00Z', rule_count: 0 })], true).tone, 'warn')
assert.equal(listLoad([st({ last_updated: '2026-07-26T10:00:00Z', rule_count: 0 })], true).tag, 'empty')
// Nothing reported is UNKNOWN, never green.
assert.equal(listLoad([], true).tone, 'unknown')
assert.equal(listLoad(null, true).tone, 'unknown')
assert.notEqual(listLoad(null, true).tone, 'on')
// A device pointing at a list the config no longer has.
assert.equal(listLoad([st({ rule_count: 9 })], false).tone, 'crit')
assert.equal(listLoad([st({ rule_count: 9 })], false).tag, 'no such list')
})
test('a geo list cannot hide a category that never arrived behind a healthy sibling', () => {
const r = listLoad(
[
st({ category: 'youtube', last_updated: '2026-07-26T10:00:00Z', rule_count: 1_284 }),
st({ category: 'google', last_updated: '', rule_count: 0 }),
],
true,
)
assert.equal(r.tone, 'crit')
assert.equal(r.tag, 'not loaded')
})
test('the chip count stays chip-sized', () => {
assert.equal(compactCount(3), '3')
assert.equal(compactCount(9_999), '9,999')
assert.equal(compactCount(218_431), '218k')
})
// ---------------------------------------------------------------------------
// 3. What the DNS page may claim
// ---------------------------------------------------------------------------
const dev = (o: Partial<Device>): Device => ({ Name: 'd', Enabled: true, ...o })
test('with nothing attached the DNS page keeps its old off-text', () => {
assert.equal(dnsFilterOffNote(null), null)
assert.equal(dnsFilterOffNote([]), null)
assert.equal(dnsFilterOffNote([dev({ Name: 'Max', Block: ['a.com'] })]), null)
})
test('one attached list makes "configured but inactive" a lie, and the replacement names who', () => {
const one = dnsFilterOffNote([dev({ Name: 'Kids iPad', Blocklists: ['family-extra'] })])
assert.ok(one)
assert.match(one, /1 device — Kids iPad —/)
assert.match(one, /still apply to it/)
assert.doesNotMatch(one, /configured but inactive/)
const many = dnsFilterOffNote([
dev({ Name: 'Max laptop', Blocklists: ['oisd-basic'] }),
dev({ Name: "Lena's phone", Allowlists: ['school-allow'] }),
dev({ Name: 'Kids iPad', Blocklists: ['family-extra'] }),
dev({ Name: 'Tablet', Blocklists: ['family-extra'] }),
])
assert.ok(many)
assert.match(many, /4 devices/)
assert.match(many, /Max laptop, Lena's phone, Kids iPad and 1 more/)
})
test('a paused device is not counted — the engine emits nothing for it', () => {
const devices = [
dev({ Name: 'Kids iPad', Enabled: false, Blocklists: ['family-extra'] }),
dev({ Name: 'Max laptop', Blocklists: ['oisd-basic'] }),
]
assert.deepEqual(devicesWithAttachedLists(devices), ['Max laptop'])
assert.deepEqual(attachedDeviceNames(devices, 'blocklist', 'family-extra'), [])
assert.deepEqual(attachedDeviceNames(devices, 'blocklist', 'oisd-basic'), ['Max laptop'])
// The kinds do not bleed into each other.
assert.deepEqual(attachedDeviceNames(devices, 'allowlist', 'oisd-basic'), [])
})
test('nameList caps the names and then counts', () => {
assert.equal(nameList([]), '')
assert.equal(nameList(['A']), 'A')
assert.equal(nameList(['A', 'B']), 'A and B')
assert.equal(nameList(['A', 'B', 'C']), 'A, B and C')
assert.equal(nameList(['A', 'B', 'C', 'D', 'E']), 'A, B, C and 2 more')
})
// ---------------------------------------------------------------------------
// 3b. The blocklist row's own state tag
// ---------------------------------------------------------------------------
const rowIn = (o: Partial<Parameters<typeof listRowState>[0]> = {}) => ({
enabled: true,
filterOn: true,
remote: true,
hasStatus: true,
neverAny: false,
ruleCount: 100,
attached: 0,
...o,
})
test('with nothing attached the row keeps exactly the readings it had', () => {
assert.deepEqual(listRowState(rowIn({ enabled: false })), { text: 'off', tone: 'off' })
assert.deepEqual(listRowState(rowIn({ filterOn: false })), { text: 'inactive', tone: 'off' })
assert.deepEqual(listRowState(rowIn({ remote: false })), { text: 'filtering', tone: 'on' })
assert.deepEqual(listRowState(rowIn({ hasStatus: false })), { text: 'load not reported', tone: 'off' })
assert.deepEqual(listRowState(rowIn({ neverAny: true })), {
text: 'not loaded — nothing blocked',
tone: 'warn',
})
assert.deepEqual(listRowState(rowIn({ ruleCount: 0 })), {
text: 'loaded empty — nothing blocked',
tone: 'warn',
})
assert.deepEqual(listRowState(rowIn()), { text: 'filtering', tone: 'on' })
})
test('a list off for the network but attached to devices is neither "off" nor "filtering"', () => {
// Switched off for the network, attached to two devices, and loaded: it works,
// for exactly those two.
assert.deepEqual(listRowState(rowIn({ enabled: false, attached: 2 })), {
text: '2 devices only',
tone: 'on',
})
assert.deepEqual(listRowState(rowIn({ filterOn: false, attached: 1 })), {
text: '1 device only',
tone: 'on',
})
// Attached but never fetched — the count must not upgrade a broken list.
assert.deepEqual(listRowState(rowIn({ enabled: false, attached: 2, neverAny: true })), {
text: 'not loaded — nothing blocked',
tone: 'warn',
})
// Attached and unreported stays unknown, and says who is relying on it.
assert.deepEqual(listRowState(rowIn({ enabled: false, attached: 1, hasStatus: false })), {
text: '1 device only · load not reported',
tone: 'off',
})
})
+411
View File
@@ -0,0 +1,411 @@
// Per-device domain policy: the parts of the Devices / DNS pages that must be
// TESTED, and therefore cannot live inside a .tsx file — `npm test` is
// `node --test src/*.test.ts`, plain modules, no JSX and no DOM. Same reasoning
// as `dnsListEdit.ts`, `ruleset.ts` and `planeState.ts`; the imports here are
// type-only, so the tests run against the real code with nothing stubbed.
//
// Three separate jobs live here, and they share one theme — a per-device rule is
// the one place in this panel where a control that LOOKS applied and is not gets
// read as "the kid is protected".
//
// 1. parseDomainEntry — what the form is allowed to accept. The panel used to
// demand /^[a-z0-9.-]+$/, so a colon was impossible to type and the engine's
// own `full:` / `suffix:` / `keyword:` vocabulary could not be reached from
// the UI at all. Opening the charset would have been worse: the engine drops
// an unrecognised `word:` prefix, drops a marker with no value, and an empty
// `keyword:` is strings.Contains(host, "") — TRUE FOR EVERY HOST, i.e. a
// lone "keyword:" in a child's Block list takes that device off the internet
// with no error anywhere (generate/devices.go:145-175). So the prefixes are
// a closed positive list and a marker without a value is refused by name.
//
// 2. listLoad — whether an attached named list actually materialised, read off
// /api/ruleset/status rather than off the fact that someone attached it. A
// list whose URL is unreachable is a parental control that does not work,
// and it has to look like one.
//
// 3. attachedDeviceNames / dnsFilterOffNote / listRowState — the DNS page's
// sentences, which stop being true the moment a list is attached to a
// device: `Blocklist.Enabled` means "participates in the NETWORK filter",
// not "this object is alive", and a device that references a list runs it
// whatever the network switch says.
import type { Device, RulesetStatus } from './api'
// ---------------------------------------------------------------------------
// 1. Domain entries
// ---------------------------------------------------------------------------
/**
* The `word:` prefixes a DOMAIN list understands — positive and closed, mirroring
* `recognisedDomainMarkers` in generate/dnsfilter.go. Anything else is provably
* unmatchable (a domain name cannot contain ":") and the engine drops it, so the
* form must refuse it instead of storing something that can only ever do nothing.
*
* `regexp:` and `geosite:` are absent ON PURPOSE, not by oversight: the DNS-filter
* path validates neither, and both are already expressed properly elsewhere (an
* inline rule-set, and a geosite-sourced list with category chips).
*/
export const DOMAIN_MARKERS = ['full', 'suffix', 'keyword'] as const
export type DomainMarker = (typeof DOMAIN_MARKERS)[number]
const MARKER_SET: ReadonlySet<string> = new Set<string>(DOMAIN_MARKERS)
/** How a stored entry matches. `suffix` is what a bare entry means in a list. */
export type DomainEntryKind = DomainMarker
export type DomainParse =
| { ok: true; value: string; kind: DomainEntryKind }
| { ok: false; reason: string }
/** A leading `word:` marker. The word must START WITH A LETTER, so a numeric IPv6
* group can never look like one (generate/dnsfilter.go domainMarkerRe). */
const MARKER_RE = /^([A-Za-z][A-Za-z0-9_-]*)\s*:\s*([\s\S]*)$/
/** The characters a hostname (or a substring of one) can be built from. */
const HOST_RE = /^[a-z0-9.-]+$/
const IPV4_RE = /^\d{1,3}(?:\.\d{1,3}){3}(?:\/\d{1,2})?$/
const IPV6_RE = /^[0-9a-f]{0,4}(?::[0-9a-f]{0,4}){2,7}(?:\/\d{1,3})?$/i
/** Strip the wildcard/dot decorations a person types around a domain. A LEADING
* DOT IS A SYNONYM OF `suffix:`, not a narrower matcher — the engine strips it
* too, so ".example.com" and "example.com" are the same entry. */
function bareDomain(s: string): string {
return s
.replace(/^\*\./, '')
.replace(/^\.+/, '')
.replace(/\.+$/, '')
}
/**
* Parse one typed entry for a device's Block/Allow list, or explain the refusal.
*
* Accepts: a bare domain (matches it and its subdomains), `full:<domain>` (that
* exact name), `suffix:<domain>` (identical to bare — canonicalised to bare, as
* the engine does), `keyword:<text>` (substring of the hostname). A leading `*.`
* or `.` is stripped.
*
* Refuses, by name: an IP address, an unknown `word:` prefix, and a marker with
* no value — the three shapes the engine silently discards.
*/
export function parseDomainEntry(raw: string): DomainParse {
const trimmed = raw.trim()
if (!trimmed) return { ok: false, reason: 'Enter a domain like example.com.' }
// Addresses first: they are not domains, and an IPv6 literal is full of colons
// so it would otherwise be mistaken for a prefixed entry.
if (IPV4_RE.test(trimmed) || IPV6_RE.test(trimmed)) {
return {
ok: false,
reason: 'That is an IP address. A device list matches names — route addresses with a Routing rule instead.',
}
}
const m = MARKER_RE.exec(trimmed)
if (m) {
const marker = m[1].toLowerCase()
const value = m[2].trim()
if (!MARKER_SET.has(marker)) {
return {
ok: false,
reason: `“${m[1]}:” is not a matcher, and a domain name cannot contain “:” — this entry could never match. Use full:, suffix: or keyword:.`,
}
}
if (!value) {
return {
ok: false,
reason:
marker === 'keyword'
? 'keyword: needs text after it. An empty keyword matches EVERY host — it would take this device off the internet entirely.'
: `${marker}: needs a domain after it. The engine drops an empty matcher, and an empty domain token stops it from starting.`,
}
}
if (marker === 'keyword') {
const kw = value.toLowerCase()
if (!HOST_RE.test(kw)) {
return {
ok: false,
reason: 'A keyword is a piece of a hostname — letters, digits, dots and dashes only.',
}
}
return { ok: true, value: `keyword:${kw}`, kind: 'keyword' }
}
const dom = bareDomain(value.toLowerCase())
if (!dom || !HOST_RE.test(dom)) {
return { ok: false, reason: `Enter a domain after ${marker}:, like ${marker}:example.com.` }
}
// `suffix:` and a bare entry produce the SAME matcher, so store one shape —
// otherwise the two spellings look like different rules and dedupe misses.
return marker === 'suffix'
? { ok: true, value: dom, kind: 'suffix' }
: { ok: true, value: `full:${dom}`, kind: 'full' }
}
const dom = bareDomain(trimmed.toLowerCase())
if (!dom || !HOST_RE.test(dom)) {
return {
ok: false,
reason: 'Enter a domain like example.com, or full: / suffix: / keyword: with a value after it.',
}
}
return { ok: true, value: dom, kind: 'suffix' }
}
/** Read a STORED entry back, for the chip label and its hover text. Never
* refuses: a value already in the config is shown as what the engine will make
* of it, not hidden. */
export function describeDomainEntry(stored: string): { label: string; kind: DomainEntryKind; detail: string } {
const m = MARKER_RE.exec(stored.trim())
if (m && MARKER_SET.has(m[1].toLowerCase())) {
const marker = m[1].toLowerCase() as DomainMarker
const value = m[2].trim()
if (marker === 'keyword') {
return { label: stored, kind: 'keyword', detail: `any hostname containing “${value}”` }
}
if (marker === 'full') {
return { label: stored, kind: 'full', detail: `exactly ${value} — subdomains are not matched` }
}
return { label: value, kind: 'suffix', detail: `${value} and its subdomains` }
}
return { label: stored, kind: 'suffix', detail: `${stored} and its subdomains` }
}
// ---------------------------------------------------------------------------
// 2. Did the attached list actually load?
// ---------------------------------------------------------------------------
/** on = materialised; warn = present but matches nothing; crit = not working;
* unknown = the engine has not said, which is not the same as fine. */
export type LoadTone = 'on' | 'warn' | 'crit' | 'unknown'
export interface ListLoad {
tone: LoadTone
/** The short tag printed on the chip. */
tag: string
/** One sentence, for the chip's title and its aria description. */
detail: string
ruleCount: number
}
/** 218431 -> "218k"; 1284 -> "1,284". Chips are narrow and a device card holds
* several of them. */
export function compactCount(n: number): string {
if (n < 10_000) return n.toLocaleString('en-US')
return `${Math.round(n / 1000).toLocaleString('en-US')}k`
}
/**
* What the running engine reports about one named list, grouped by name (a
* geosite list with N categories emits N records).
*
* `known` is whether the config still HAS a list by that name: a device pointing
* at a deleted list is the loudest possible failure and gets said first.
*
* The order of the remaining checks matters. "never fetched" outranks "no rules",
* so a geo list whose second category never arrived cannot hide behind a healthy
* sibling; "nothing reported" is its own state and is never drawn as healthy.
*/
export function listLoad(
statuses: RulesetStatus[] | null | undefined,
known: boolean,
): ListLoad {
if (!known) {
return {
tone: 'crit',
tag: 'no such list',
detail: 'This device points at a list that is not in the config — it filters nothing.',
ruleCount: 0,
}
}
const recs = statuses ?? []
if (recs.length === 0) {
return {
tone: 'unknown',
tag: 'load unknown',
detail: 'The engine has not reported this list — it may not have been applied yet. Unknown, not fine.',
ruleCount: 0,
}
}
let ruleCount = 0
let neverAny = false
// How many of these records CAN report a count. Only remote sets do: the engine
// fills RuleSetStat.RuleCount from (*rule.RemoteRuleSet).RuleCount(), and a
// local set has no such method at all — `LocalRuleSet.RuleCount` does not exist
// in the tree. So a local record's zero is the field never being written, not a
// list with nothing in it, and summing the two kinds together silently turns
// "not published" into "empty".
let counted = 0
for (const s of recs) {
if (s.remote) {
counted++
ruleCount += s.rule_count
}
if (s.remote && !s.last_updated) neverAny = true
}
if (neverAny) {
return {
tone: 'crit',
tag: 'not loaded',
detail: 'The engine has never fetched this list, so it matches nothing for this device.',
ruleCount,
}
}
// NOTHING HERE COULD BE COUNTED. An inline or file-backed list is reported by
// the engine — so it exists in the running box — but the daemon publishes no
// size for it, ever. This used to read "empty · nothing matches", which is a
// verdict about a working list built entirely out of a field that is never
// filled in. It is an unknown, and it is drawn as one: unlit, never green and
// never the amber that says something is wrong.
if (counted === 0) {
return {
tone: 'unknown',
tag: 'size unknown',
detail:
'This list is in the running engine, but the daemon publishes an entry count only for lists it fetches from a URL — so how many entries an inline or file list holds is not reported. Not a fault, and not a confirmation that it matches anything.',
ruleCount: 0,
}
}
if (ruleCount === 0) {
return {
tone: 'warn',
tag: 'empty',
detail: 'Loaded, but it holds no entries — nothing matches.',
ruleCount: 0,
}
}
// A group whose records are a MIX (a geosite name resolving to several sets,
// one of them local) can only be counted in part, so the number is a floor and
// says so rather than presenting a partial sum as the whole.
const partial = counted < recs.length
return {
tone: 'on',
tag: partial ? `${compactCount(ruleCount)}+` : compactCount(ruleCount),
detail: partial
? `At least ${ruleCount.toLocaleString('en-US')} entries loaded and matching; the inline part of this list is not counted by the daemon.`
: `${ruleCount.toLocaleString('en-US')} entries loaded and matching.`,
ruleCount,
}
}
// ---------------------------------------------------------------------------
// 3. What the DNS page may claim once a list is attached to a device
// ---------------------------------------------------------------------------
export type ListKind = 'blocklist' | 'allowlist'
/** Which of a device's list slots a kind reads. */
function slotOf(d: Device, kind: ListKind): string[] {
const raw = kind === 'blocklist' ? d.Blocklists : d.Allowlists
return raw ?? []
}
/**
* The ENABLED devices attaching `name`, in config order.
*
* Enabled-only because generate/devices.go resolves only enabled devices — a
* paused device emits no rule at all, so counting it would put a number on the
* DNS page that nothing on the router agrees with.
*/
export function attachedDeviceNames(
devices: Device[] | null | undefined,
kind: ListKind,
name: string,
): string[] {
const out: string[] = []
for (const d of devices ?? []) {
if (d.Enabled === false) continue
if (slotOf(d, kind).includes(name)) out.push(d.Name)
}
return out
}
/** Every enabled device attaching at least one list of either kind, in config
* order, deduped by name. */
export function devicesWithAttachedLists(devices: Device[] | null | undefined): string[] {
const out: string[] = []
const seen = new Set<string>()
for (const d of devices ?? []) {
if (d.Enabled === false) continue
if (slotOf(d, 'blocklist').length === 0 && slotOf(d, 'allowlist').length === 0) continue
const nm = d.Name || '(unnamed device)'
if (seen.has(nm)) continue
seen.add(nm)
out.push(nm)
}
return out
}
/** "A, B and C" / "A, B, C and 2 more" — at most three names, then a count. */
export function nameList(names: string[], max = 3): string {
if (names.length === 0) return ''
if (names.length === 1) return names[0]
const head = names.slice(0, max)
const rest = names.length - head.length
if (rest > 0) return `${head.join(', ')} and ${rest} more`
return `${head.slice(0, -1).join(', ')} and ${head[head.length - 1]}`
}
/**
* The DNS page's third sentence for the master switch.
*
* The two it had were "on: lists are answering" and "off: the lists below are
* configured but inactive". The second becomes a lie the moment a device attaches
* one, because attaching is itself the switch for that device. Returns null when
* no device attaches anything, so the caller keeps the plain off-text.
*/
export function dnsFilterOffNote(devices: Device[] | null | undefined): string | null {
const names = devicesWithAttachedLists(devices)
if (names.length === 0) return null
const who = nameList(names)
return names.length === 1
? `Network-wide filtering is off, but 1 device — ${who} — has lists attached on the Devices page, and those still apply to it. Turn this on to filter DNS for every device.`
: `Network-wide filtering is off, but ${names.length} devices — ${who} — have lists attached on the Devices page, and those still apply to them. Turn this on to filter DNS for every device.`
}
/**
* The state tag on a blocklist/allowlist row.
*
* Extracted from DNS.tsx so it can be tested, and extended with one fact the
* inline version could not know: `attached`. A list switched off for the network
* but referenced by a device is neither "off" nor "filtering" — it runs, for
* exactly those devices, and only if it actually loaded.
*/
export interface ListRowStateIn {
/** Blocklist/Allowlist.Enabled — participates in the NETWORK filter. */
enabled: boolean
/** Globals.DNSFilter. */
filterOn: boolean
/** url/geosite lists are fetched; inline/file are read straight from config. */
remote: boolean
hasStatus: boolean
neverAny: boolean
ruleCount: number
/** How many ENABLED devices attach this list. */
attached: number
}
export function listRowState(i: ListRowStateIn): { text: string; tone: 'on' | 'off' | 'warn' } {
const networkOn = i.enabled && i.filterOn
if (!networkOn && i.attached === 0) {
return i.enabled ? { text: 'inactive', tone: 'off' } : { text: 'off', tone: 'off' }
}
const load: 'ok' | 'unknown' | 'never' | 'empty' = !i.remote
? 'ok'
: !i.hasStatus
? 'unknown'
: i.neverAny
? 'never'
: i.ruleCount === 0
? 'empty'
: 'ok'
// Who is still using it, when the network is not.
const who = networkOn ? '' : `${i.attached} device${i.attached === 1 ? '' : 's'} only`
switch (load) {
case 'unknown':
return { text: who ? `${who} · load not reported` : 'load not reported', tone: 'off' }
case 'never':
return { text: 'not loaded — nothing blocked', tone: 'warn' }
case 'empty':
return { text: 'loaded empty — nothing blocked', tone: 'warn' }
case 'ok':
return { text: who || 'filtering', tone: 'on' }
}
}
+273
View File
@@ -0,0 +1,273 @@
// The DNS page's edit-after-create merges.
//
// Run with `npm test` (node's built-in test runner + native type stripping).
// dnsListEdit.ts has no runtime imports, so this runs against the real module.
//
// THE CASE THESE WERE WRITTEN FOR: before this, a blocklist, an allowlist and a
// resolver could only be configured at CREATION. A typo in a URL meant deleting
// the list and building it again; a resolver's address could not be corrected at
// all — only its detour. Adding the editors introduces the failure `ruleset.ts`
// already documents: a form that REBUILDS an object drops every field it does
// not render, silently, at the moment someone renames something.
//
// So each test below fixes one field the form must carry through, and one field
// switching source must genuinely clear. The compile-time half is `Complete<T>`
// on the return types: a new field in api.ts fails the BUILD in the function that
// has to decide about it, which no runtime test can do.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
allowlistToForm,
blocklistToForm,
isRemoteListSource,
nextAllowlist,
nextBlocklist,
nextResolver,
parseDomains,
resolverNeedsAddress,
resolverToForm,
} from './dnsListEdit.ts'
import type { Allowlist, Blocklist, Resolver } from './api.ts'
// ---- blocklists -------------------------------------------------------------
test('renaming a url blocklist keeps the refresh interval it was given', () => {
// A list written over SSH with a fast cadence, because it is the one that keeps
// breaking a site. The panel pinned every list to 24h and showed no control, so
// this value had nowhere to come back from once dropped.
const stored: Blocklist = {
Name: 'oisd',
Enabled: true,
Source: 'url',
URL: 'https://big.oisd.nl',
Response: 'nxdomain',
UpdateInterval: '1h',
}
const form = blocklistToForm(stored)
const out = nextBlocklist({ ...form, name: 'oisd-big' })
assert.equal(out.Name, 'oisd-big')
assert.equal(out.UpdateInterval, '1h', 'the interval must survive a rename')
assert.equal(out.URL, 'https://big.oisd.nl')
assert.equal(out.Response, 'nxdomain')
})
test('a blocklist keeps its non-default reply through an unrelated edit', () => {
const stored: Blocklist = {
Name: 'ads',
Enabled: true,
Source: 'url',
URL: 'https://example.org/ads.txt',
Response: 'zero',
UpdateInterval: '24h',
}
const out = nextBlocklist({ ...blocklistToForm(stored), url: 'https://example.org/ads-v2.txt' })
assert.equal(out.Response, 'zero', 'a 0.0.0.0 list must not silently become NXDOMAIN')
assert.equal(out.URL, 'https://example.org/ads-v2.txt')
})
test('a file blocklist keeps its path — the source the add form never offered', () => {
// `file` cannot be created in the panel, only over SSH, so a form that dropped
// Path would make every such list unopenable without destroying it.
const stored: Blocklist = {
Name: 'local',
Enabled: true,
Source: 'file',
Path: '/etc/shater/blocked.lst',
Response: 'nxdomain',
}
const out = nextBlocklist({ ...blocklistToForm(stored), name: 'local-hosts' })
assert.equal(out.Path, '/etc/shater/blocked.lst')
assert.equal(out.Source, 'file')
})
test('a geosite blocklist keeps every category chip', () => {
const stored: Blocklist = {
Name: 'trackers',
Enabled: true,
Source: 'geosite',
Categories: ['category-ads-all', 'win-spy'],
Response: 'nxdomain',
UpdateInterval: '12h',
}
const out = nextBlocklist({ ...blocklistToForm(stored), name: 'trackers-2' })
assert.deepEqual(out.Categories, ['category-ads-all', 'win-spy'])
assert.equal(out.UpdateInterval, '12h')
})
test('switching a blocklist from url to inline really does drop the url', () => {
// The other half of the contract. A merge that only ever preserves would leave
// the config carrying a URL the new source ignores.
const stored: Blocklist = {
Name: 'ads',
Enabled: true,
Source: 'url',
URL: 'https://example.org/ads.txt',
UpdateInterval: '6h',
}
const out = nextBlocklist({
...blocklistToForm(stored),
source: 'inline',
entriesText: 'ads.example.com',
})
assert.equal(out.URL, undefined)
assert.deepEqual(out.Entries, ['ads.example.com'])
assert.equal(
out.UpdateInterval,
undefined,
'an inline list is never fetched, so it must not carry a refresh cadence',
)
})
test('an emptied inline list writes no Entries key rather than an empty one', () => {
const stored: Blocklist = { Name: 'manual', Enabled: true, Source: 'inline', Entries: ['a.test'] }
const out = nextBlocklist({ ...blocklistToForm(stored), entriesText: ' ' })
assert.equal(out.Entries, undefined)
})
// ---- allowlists -------------------------------------------------------------
test('an allowlist carries an update interval at all — the field it never had', () => {
// The asymmetry this closes: an allowlist is HOW a blocklist's false positive
// is corrected. Pinned to 24h while the blocklist that broke the site refreshes
// faster, the fix lands up to a day after the breakage.
const stored: Allowlist = {
Name: 'work',
Enabled: true,
Source: 'url',
URL: 'https://example.org/allow.txt',
UpdateInterval: '30m',
}
const form = allowlistToForm(stored)
assert.equal(form.interval, '30m', 'the stored interval must reach the form')
const out = nextAllowlist({ ...form, name: 'work-allow' })
assert.equal(out.Name, 'work-allow')
assert.equal(out.UpdateInterval, '30m', 'and survive a rename')
})
test('an allowlist rebuild emits no Response — the field is not on the shape', () => {
const stored: Allowlist = { Name: 'work', Enabled: true, Source: 'inline', Entries: ['ok.test'] }
const out = nextAllowlist(allowlistToForm(stored))
assert.equal(
'Response' in out,
false,
'PUT /api/config decodes with DisallowUnknownFields — an invented key fails the whole write',
)
})
test('a disabled allowlist stays disabled through an edit', () => {
const stored: Allowlist = { Name: 'work', Enabled: false, Source: 'inline', Entries: ['ok.test'] }
const out = nextAllowlist({ ...allowlistToForm(stored), name: 'work2' })
assert.equal(out.Enabled, false)
})
// ---- resolvers --------------------------------------------------------------
test('renaming a fake-IP resolver keeps its pool', () => {
// The pool was settable only at creation. Losing it is not cosmetic: a fake-IP
// resolver without one hands out addresses from whatever the engine defaults to.
const stored: Resolver = { Name: 'fake', Type: 'fakeip', Pool: '198.18.0.0/15' }
const out = nextResolver({ ...resolverToForm(stored, 'direct'), name: 'fakeip' })
assert.equal(out.Name, 'fakeip')
assert.equal(out.Pool, '198.18.0.0/15')
})
test('editing a DoH resolver keeps the detour its row set', () => {
const stored: Resolver = {
Name: 'cf',
Type: 'doh',
Address: 'https://1.1.1.1/dns-query',
Detour: 'group:eu',
}
const out = nextResolver({
...resolverToForm(stored, 'group:eu'),
address: 'https://1.0.0.1/dns-query',
})
assert.equal(out.Address, 'https://1.0.0.1/dns-query')
assert.equal(out.Detour, 'group:eu', 'the DNS path is set on the row and must not be lost here')
})
test('switching a resolver to fake-IP drops the address and the path', () => {
// Neither means anything to fake-IP: it answers out of the pool and sends
// nothing, so a kept detour would describe a route no packet takes.
const stored: Resolver = {
Name: 'cf',
Type: 'doh',
Address: 'https://1.1.1.1/dns-query',
Detour: 'group:eu',
}
const out = nextResolver({
...resolverToForm(stored, 'group:eu'),
type: 'fakeip',
pool: '198.18.0.0/15',
})
assert.equal(out.Address, undefined)
assert.equal(out.Detour, undefined)
assert.equal(out.Pool, '198.18.0.0/15')
})
test('“direct” is the absence of a detour, not a value to store', () => {
const stored: Resolver = { Name: 'q9', Type: 'plain', Address: '9.9.9.9', Detour: 'group:eu' }
const out = nextResolver({ ...resolverToForm(stored, 'group:eu'), detour: 'direct' })
assert.equal(out.Detour, undefined)
})
test('a local resolver carries no address — there is nothing to dial', () => {
const out = nextResolver({
name: 'system',
type: 'local',
address: 'left over from doh',
pool: '',
detour: 'direct',
})
assert.equal(out.Address, undefined)
})
// ---- the small predicates the forms branch on -------------------------------
test('only url and geosite are refreshed on a cadence', () => {
assert.equal(isRemoteListSource('url'), true)
assert.equal(isRemoteListSource('geosite'), true)
assert.equal(isRemoteListSource('inline'), false)
assert.equal(isRemoteListSource('file'), false)
})
test('the four resolver types that dial a server are the ones needing an address', () => {
for (const t of ['doh', 'dot', 'plain', 'tcp']) assert.equal(resolverNeedsAddress(t), true, t)
for (const t of ['local', 'fakeip']) assert.equal(resolverNeedsAddress(t), false, t)
})
test('the domain textarea drops comments, lower-cases, and unwraps a hosts-file star', () => {
// A bare domain is a SUFFIX match in the engine, so `*.` would only produce an
// entry that matches nothing — stripping it is the difference between a working
// list and a silently dead one.
assert.deepEqual(
parseDomains('# ads\nAds.Example.com\n*.tracker.test\nads.example.com\n!comment'),
['ads.example.com', 'tracker.test'],
)
assert.deepEqual(parseDomains(' \n\n'), [])
assert.deepEqual(parseDomains('a.test, b.test'), ['a.test', 'b.test'])
})
test('a comment is cut at the line, not at the token', () => {
// The whole line after `#` is a note. Splitting the textarea on whitespace and
// dropping only tokens that start with `#` turned this into three entries —
// `ads`, `and`, `trackers` — all valid, all silently blocked.
assert.deepEqual(parseDomains('# ads and trackers\nreal.test'), ['real.test'])
assert.deepEqual(parseDomains('real.test # keep this one'), ['real.test'])
assert.deepEqual(parseDomains('! AdGuard style note\nreal.test'), ['real.test'])
})
test('a stored inline list round-trips through the textarea unchanged', () => {
const stored: Blocklist = {
Name: 'manual',
Enabled: true,
Source: 'inline',
Entries: ['a.test', 'b.test'],
}
const out = nextBlocklist(blocklistToForm(stored))
assert.deepEqual(out.Entries, ['a.test', 'b.test'])
})
+220
View File
@@ -0,0 +1,220 @@
import type { Allowlist, Blocklist, Complete, Resolver } from './api'
/**
* The DNS page's edit-after-create merges: what a Save writes back for a
* blocklist, an allowlist and a resolver.
*
* Lives outside `pages/DNS.tsx` because it is the part that must be TESTED, and
* the panel's runner is `node --test src/*.test.ts` — plain modules, no JSX, no
* DOM. Same reasoning as `egressEdit.ts` and `ruleset.ts`.
*
* # The rule these three functions exist to keep
*
* A save may only CLEAR a field the editor was in a position to SHOW. That rule
* was written after a rule-set rename silently dropped its `Format` (see
* `ruleset.ts`), and it cuts both ways here:
*
* - every one of the three shapes below is rebuilt from the form rather than
* spread onto the stored object, so switching a list from `url` to `inline`
* really does drop the stale URL instead of leaving a value the new source
* ignores;
* - which is only safe because the editor renders a control for EVERY field of
* the shape. So the return type is `Complete<T>`: a field added to `Blocklist`
* / `Allowlist` / `Resolver` in api.ts stops the build HERE, in the function
* that has to decide, instead of vanishing on the next rename.
*
* Writing `undefined` for a field the chosen source has no use for is therefore a
* statement, not an omission.
*/
/** The four `Source` values a DNS list can carry (model.go Blocklist.Source). */
export type ListSource = 'inline' | 'file' | 'url' | 'geosite'
/** What a blocklist answers with when it matches. */
export type BlockResponse = 'nxdomain' | 'zero'
/** Sources fetched on a cadence, and therefore the only ones an interval means anything for. */
const REMOTE_SOURCES: ReadonlySet<string> = new Set<ListSource>(['url', 'geosite'])
/** True when `source` is re-fetched by the engine, so an update interval applies. */
export function isRemoteListSource(source: string): boolean {
return REMOTE_SOURCES.has(source)
}
/**
* Split a textarea of domains: line by line, comments dropped, then whitespace or
* comma separated, lower-cased, a leading `*.` stripped, deduped.
*
* A bare domain is a SUFFIX match in the engine, so `example.com` already covers
* its subdomains and the `*.` a hosts-file habit adds would only make the entry
* fail to match anything.
*
* COMMENTS ARE CUT PER LINE, not per token. The version this replaces split the
* whole textarea on whitespace and dropped only the tokens that THEMSELVES began
* with `#`/`!` — so pasting `# ads and trackers` contributed `ads`, `and` and
* `trackers` as three real blocked domains, from a line the operator wrote as a
* note. Nothing reported it: they are valid entries that simply match nothing,
* until one of them is a name someone needs.
*/
export function parseDomains(text: string): string[] {
const seen = new Set<string>()
const out: string[] = []
for (const line of text.split(/\r?\n/)) {
// `#` (hosts files) and `!` (AdGuard/uBlock) both start a comment; neither is
// legal in a hostname, so cutting at the first one cannot eat a real entry.
const body = line.split(/[#!]/, 1)[0]
for (const raw of body.split(/[\s,]+/)) {
const d = raw.trim().toLowerCase().replace(/^\*\./, '')
if (!d || seen.has(d)) continue
seen.add(d)
out.push(d)
}
}
return out
}
/**
* The editor's state for a blocklist or an allowlist, exactly as the controls
* hold it: `entriesText` is the raw textarea (parsed on save by
* {@link parseDomains}), `categories` is the chip widget's already-validated
* array, and the rest are raw strings.
*/
export interface ListForm {
name: string
enabled: boolean
source: ListSource
url: string
path: string
entriesText: string
categories: string[]
/** Blocklists only — an allowlist has no reply to choose. */
response: BlockResponse
/** Go duration, e.g. `6h`. Blank ⇒ the daemon's 24h default. */
interval: string
}
const trimmed = (v: string): string | undefined => (v.trim() ? v.trim() : undefined)
const list = (v: string[]): string[] | undefined => (v.length ? v : undefined)
const entriesOf = (text: string): string[] | undefined => list(parseDomains(text))
/**
* Rebuild a blocklist from the edit form.
*
* `UpdateInterval` is carried only for the two remote sources. That is not
* tidiness: on an inline list the engine fetches nothing, so a stored interval
* would sit in the config describing a refresh that never happens — and the row
* reads the interval the ENGINE reports, so the two would disagree on screen.
*/
export function nextBlocklist(f: ListForm): Complete<Blocklist> {
const remote = isRemoteListSource(f.source)
return {
Name: f.name.trim(),
Enabled: f.enabled,
Source: f.source,
URL: f.source === 'url' ? trimmed(f.url) : undefined,
Path: f.source === 'file' ? trimmed(f.path) : undefined,
Entries: f.source === 'inline' ? entriesOf(f.entriesText) : undefined,
Categories: f.source === 'geosite' ? list(f.categories) : undefined,
Response: f.response,
UpdateInterval: remote ? trimmed(f.interval) : undefined,
}
}
/** Rebuild an allowlist from the edit form. Same contract as {@link nextBlocklist},
* minus `Response` — an allowlist does not answer, it exempts. */
export function nextAllowlist(f: ListForm): Complete<Allowlist> {
const remote = isRemoteListSource(f.source)
return {
Name: f.name.trim(),
Enabled: f.enabled,
Source: f.source,
URL: f.source === 'url' ? trimmed(f.url) : undefined,
Path: f.source === 'file' ? trimmed(f.path) : undefined,
Entries: f.source === 'inline' ? entriesOf(f.entriesText) : undefined,
Categories: f.source === 'geosite' ? list(f.categories) : undefined,
UpdateInterval: remote ? trimmed(f.interval) : undefined,
}
}
/** Seed the edit form from a stored blocklist. */
export function blocklistToForm(b: Blocklist): ListForm {
return {
name: b.Name,
enabled: b.Enabled,
source: b.Source,
url: b.URL ?? '',
path: b.Path ?? '',
entriesText: (b.Entries ?? []).join('\n'),
categories: b.Categories ?? [],
response: b.Response === 'zero' ? 'zero' : 'nxdomain',
interval: b.UpdateInterval ?? '',
}
}
/** Seed the edit form from a stored allowlist. `response` is a placeholder the
* allowlist branch never reads — {@link nextAllowlist} does not emit it. */
export function allowlistToForm(a: Allowlist): ListForm {
return {
name: a.Name,
enabled: a.Enabled,
source: a.Source,
url: a.URL ?? '',
path: a.Path ?? '',
entriesText: (a.Entries ?? []).join('\n'),
categories: a.Categories ?? [],
response: 'nxdomain',
interval: a.UpdateInterval ?? '',
}
}
// ---- resolvers --------------------------------------------------------------
/** Resolver kinds that dial a remote server and therefore need an address. */
const REMOTE_RESOLVERS: ReadonlySet<string> = new Set(['doh', 'dot', 'plain', 'tcp'])
/** True when this resolver type needs an `Address`. */
export function resolverNeedsAddress(type: string): boolean {
return REMOTE_RESOLVERS.has(type)
}
/** The resolver editor's state. */
export interface ResolverForm {
name: string
type: string
address: string
pool: string
/** Canonical detour value: `direct` | `group:x` | `chain:x` | `egress:x` | `node:x`. */
detour: string
}
/**
* Rebuild a resolver from the edit form.
*
* A fake-IP resolver answers out of a local pool and never sends a packet, so it
* carries neither an address nor a detour — the editor renders no control for
* either, and this writes `undefined` rather than leaving a path the engine
* throws away. `direct` is the absence of a detour, not a value.
*/
export function nextResolver(f: ResolverForm): Complete<Resolver> {
const fakeip = f.type === 'fakeip'
return {
Name: f.name.trim(),
Type: f.type,
Address: resolverNeedsAddress(f.type) ? trimmed(f.address) : undefined,
Pool: fakeip ? trimmed(f.pool) : undefined,
Detour: fakeip || f.detour === 'direct' ? undefined : trimmed(f.detour),
}
}
/** Seed the resolver editor from a stored resolver. `detour` must already be
* canonicalised by the caller (DNS.tsx owns the catalog that resolves a bare
* legacy name), so an unresolvable value passes through and stays visible. */
export function resolverToForm(r: Resolver, canonDetour: string): ResolverForm {
return {
name: r.Name,
type: r.Type,
address: r.Address ?? '',
pool: r.Pool ?? '',
detour: canonDetour,
}
}
+298
View File
@@ -0,0 +1,298 @@
// Tests for the DNS row's OUTCOME reading — logRoute.dnsRowMark.
//
// The defect this file exists to keep dead: a lookup that FAILED was drawn from
// `action` alone. A query that left through a detour and then timed out came back
// as an accent-coloured `proxy` tag and nothing else, which made the row that
// describes the tunnel breaking the healthiest-looking line in the whole log.
// `actionTag()` also had an open `return 'pass'`, so every value the panel did not
// recognise — including every value a future daemon might add — came out green.
//
// Four things are defended, and the CONTROL is the load-bearing one:
//
// 1. the four situations a row can be in (answered / blocked / failed / not
// recorded) are drawn PAIRWISE DIFFERENTLY. An instrument whose failure looks
// like its success is useless; so is one where everything looks like failure;
// 2. `action=proxy` + `status=failed` is distinguishable BOTH from a healthy
// proxy AND from a failed pass — the pair is the whole diagnosis;
// 3. every fallback is closed and lands on the recoverable side;
// 4. a failure with no recorded cause SAYS so.
//
// Run: npm test (node --test src/*.test.ts)
import test from 'node:test'
import assert from 'node:assert/strict'
import { readFileSync } from 'node:fs'
import { fileURLToPath } from 'node:url'
import { LOG_SEARCH_HINT, dnsRowMark, logSearchFields, rowMatches } from './logRoute.ts'
import type { DnsRowMark } from './logRoute.ts'
import type { QueryLogEntry } from './api.ts'
// ---- fixtures ---------------------------------------------------------------
function query(over: Partial<QueryLogEntry> = {}): QueryLogEntry {
return {
time: '12:00:00',
unix: 1_700_000_000,
domain: 'youtube.com',
qtype: 'A',
rcode: 0,
blocked: false,
server: 'cloudflare-doh',
action: 'pass',
device: 'laptop',
seq: 7,
status: 'answered',
outbound_kind: 'default',
outbound: '',
...over,
}
}
/** The four situations, exactly as the daemon can produce them. */
const ANSWERED = query({ status: 'answered', blocked: false, action: 'pass' })
const BLOCKED = query({ status: 'answered', blocked: true, action: 'block', rcode: 3 })
const FAILED = query({
status: 'failed',
blocked: false,
action: 'proxy',
error: 'dial udp 1.1.1.1:53: i/o timeout',
rcode: -1,
})
// An old row out of the persistent store: the outcome was never written.
const UNRECORDED = query({ status: '', blocked: false, action: 'pass' })
/**
* Everything of the reading that reaches the screen, as one string. Two rows are
* "drawn the same" exactly when these match: the row's own marking, both chips'
* classes and texts, and the failure lines.
*/
function drawn(m: DnsRowMark): string {
return JSON.stringify([m.state, m.outcome, m.label, m.path, m.pathLabel, m.cause, m.rcodeNote])
}
// ---- 1 · THE CONTROL: no two of the four may look alike ----------------------
test('CONTROL: the four situations are drawn pairwise differently', () => {
const cases: [string, QueryLogEntry][] = [
['answered', ANSWERED],
['blocked', BLOCKED],
['failed', FAILED],
['not recorded', UNRECORDED],
]
const seen = new Map<string, string>()
for (const [name, row] of cases) {
const key = drawn(dnsRowMark(row))
const clash = seen.get(key)
assert.equal(
clash,
undefined,
`“${name}” is drawn exactly like “${clash}” — ${key}. A reader cannot tell them apart.`,
)
seen.set(key, name)
}
assert.equal(seen.size, 4)
})
// The other half of the control. The test above passes just as happily if every
// row is marked as a failure, which would be the opposite lie: an instrument that
// reads "broken" on healthy traffic is no more useful than one that reads "fine"
// on broken traffic. So the healthy readings are pinned positively.
test('CONTROL: an answered row is not marked as a failure', () => {
const m = dnsRowMark(ANSWERED)
assert.equal(m.state, 'answered')
assert.equal(m.outcome, 'answered')
assert.equal(m.cause, '', 'a row that did not fail has no failure cause to show')
assert.equal(m.rcodeNote, '')
const b = dnsRowMark(BLOCKED)
assert.equal(b.state, 'blocked', 'the filter answering is not a malfunction')
assert.equal(b.outcome, 'answered')
assert.equal(b.cause, '')
})
// ---- 2 · the pair: which path failed ----------------------------------------
test('a failure in the tunnel is distinguishable from a healthy tunnel', () => {
const healthy = dnsRowMark(query({ status: 'answered', action: 'proxy' }))
const broken = dnsRowMark(FAILED)
assert.equal(healthy.path, 'proxy')
assert.equal(broken.path, 'proxy', 'the PATH must survive the failure — it names what broke')
assert.notEqual(drawn(healthy), drawn(broken))
assert.equal(broken.state, 'failed')
assert.equal(broken.label, 'failed')
})
test('a failure in the tunnel is distinguishable from a failure on the direct path', () => {
const inTunnel = dnsRowMark(FAILED)
const direct = dnsRowMark(
query({ status: 'failed', action: 'pass', error: 'rejected', rcode: 2 }),
)
assert.equal(inTunnel.state, direct.state, 'both are failures — the outcome axis agrees')
assert.notEqual(inTunnel.path, direct.path)
assert.notEqual(drawn(inTunnel), drawn(direct))
})
test('the outcome leads: a failed row never keeps a healthy outcome word', () => {
for (const action of ['pass', 'proxy', 'block', '', 'a-word-from-2027']) {
const m = dnsRowMark(query({ status: 'failed', action }))
assert.equal(m.state, 'failed', `action=${action || '(empty)'} must not overrule the outcome`)
assert.equal(m.outcome, 'failed')
}
})
// ---- 3 · every fallback is closed -------------------------------------------
test('an unrecognised status falls to NOT RECORDED, never to answered', () => {
for (const status of ['', 'ANSWERED', 'ok', 'partial', 'tomorrows-word']) {
const m = dnsRowMark(query({ status }))
assert.equal(m.outcome, 'unrecorded', `status=${status || '(empty)'}`)
assert.equal(m.state, 'unrecorded')
assert.equal(m.label, 'not recorded')
}
})
// This is the open `return 'pass'` that used to live in Insights.actionTag: any
// value the panel did not know came out as the green, everything-is-fine path.
test('an unrecognised action falls to UNKNOWN, never to pass', () => {
for (const action of ['', 'PASS', 'reject', 'hijack-dns', 'tomorrows-word']) {
const m = dnsRowMark(query({ action }))
assert.equal(m.path, 'unknown', `action=${action || '(empty)'}`)
assert.equal(m.pathLabel, 'unknown', 'an unknown path is named, not left blank')
}
})
test('an unknown path names the value it could not read', () => {
assert.match(dnsRowMark(query({ action: 'hijack-dns' })).pathTitle, /hijack-dns/)
assert.match(dnsRowMark(query({ action: '' })).pathTitle, /not recorded/i)
})
test('a not-recorded outcome is not a claim that the lookup succeeded', () => {
const m = dnsRowMark(UNRECORDED)
assert.match(m.title, /NOT a claim/)
assert.equal(m.cause, '', 'an unknown outcome has no failure to explain either')
})
// `blocked` is a fact from a different field, and it is only allowed to REFINE a
// recorded answer. A build that never wrote the outcome does not get to have its
// other fields stand in for one.
test('blocked does not manufacture an outcome on a row that has none', () => {
const m = dnsRowMark(query({ status: '', blocked: true, action: 'block' }))
assert.equal(m.outcome, 'unrecorded')
assert.equal(m.state, 'unrecorded')
})
// ---- 4 · a failure with no recorded cause says so ---------------------------
test('a failure shows its cause verbatim, uncategorised', () => {
assert.equal(
dnsRowMark(FAILED).cause,
'cause dial udp 1.1.1.1:53: i/o timeout',
'the resolver text is passed through, not graded into a category',
)
assert.equal(dnsRowMark(query({ status: 'failed', error: 'loopback' })).cause, 'cause loopback')
assert.equal(
dnsRowMark(query({ status: 'failed', error: 'rejected (cached)' })).cause,
'cause rejected (cached)',
)
})
test('a failure with NO cause says the cause was not recorded, not nothing', () => {
for (const e of [undefined, '', ' ']) {
const m = dnsRowMark(query({ status: 'failed', error: e }))
assert.equal(m.cause, 'cause not recorded', `error=${JSON.stringify(e)}`)
}
})
test('rcode separates "never answered" from "answered with a refusal"', () => {
assert.equal(dnsRowMark(query({ status: 'failed', rcode: -1 })).rcodeNote, 'no response')
assert.equal(dnsRowMark(query({ status: 'failed', rcode: 2 })).rcodeNote, 'rcode 2')
assert.equal(dnsRowMark(query({ status: 'failed', rcode: 5 })).rcodeNote, 'rcode 5')
// CONTROL: the rcode line belongs to failures only — on a healthy row it would
// be noise dressed as a diagnosis.
assert.equal(dnsRowMark(query({ status: 'answered', rcode: 3 })).rcodeNote, '')
})
// ---- 5 · the search box, and what it promises -------------------------------
test('search FINDS a failure by its cause', () => {
assert.equal(rowMatches(logSearchFields(FAILED), 'timeout'), true)
assert.equal(rowMatches(logSearchFields(FAILED), 'I/O TIMEOUT'), true)
})
// CONTROL for the assertion above: a field list that returned everything would
// pass it. `status` is a vocabulary word — q=failed must not sweep up every
// failure while the operator is looking for a domain with "failed" in its name.
test('CONTROL search DOES NOT FIND: status is not searched', () => {
assert.equal(
rowMatches(logSearchFields(query({ status: 'failed', error: '', domain: 'a.example', action: 'pass', server: 'r', device: 'd', outbound: '' })), 'failed'),
false,
)
assert.equal(rowMatches(logSearchFields(ANSWERED), 'answered'), false)
})
test('the search hint names error, because the daemon searches it', () => {
assert.match(LOG_SEARCH_HINT, /error/)
})
// ---- 6 · the page actually renders the reading ------------------------------
//
// The reading is pure and testable; the JSX around it is not, under `node --test`
// (no DOM, no JSX transform). So the seam between them is checked as source text:
// a component that quietly stopped drawing `cause`, or that kept its own opinion
// about what an unknown action means, would pass every assertion above.
const INSIGHTS = readFileSync(
fileURLToPath(new URL('./pages/Insights.tsx', import.meta.url)),
'utf8',
)
// Anchored at the RENDER SITE, not at "the identifier appears somewhere": a first
// attempt only checked that the string `mark.cause` was in the file, and it
// survived being rewritten to `{false && mark.cause ? …}`. A grep that a disabled
// render passes is not an instrument.
test('Insights renders every part of the reading', () => {
const sites: [string, RegExp][] = [
['the row-state marking', /st-\$\{mark\.state\}/],
['the outcome chip class', /s-\$\{mark\.outcome\}/],
['the outcome chip text', /\{mark\.label\}/],
['the outcome tooltip', /title=\{mark\.title\}/],
['the path tag class', /tag \$\{mark\.path\}/],
['the path tag text', /\{mark\.pathLabel\}/],
['the failure cause', /\{mark\.cause \?/],
['the rcode note', /\{mark\.rcodeNote \?/],
]
for (const [what, at] of sites) {
assert.match(INSIGHTS, at, `Insights.tsx does not render ${what}`)
}
})
test('Insights keeps no open fallback of its own', () => {
assert.ok(
!/function actionTag/.test(INSIGHTS),
'the open actionTag fallback is back — unknown values are green again',
)
assert.ok(!/return 'pass'/.test(INSIGHTS))
})
// Each selector must open a RULE — `.st-failed` followed by `{` or `,`. Matching
// the bare substring would be satisfied by `.st-failed-DISABLED`, which is exactly
// how one earlier version of this test let a deleted rule through.
test('the four row states each get their own marking in the stylesheet', () => {
const css = readFileSync(
fileURLToPath(new URL('./pages/Insights.css', import.meta.url)),
'utf8',
)
for (const sel of ['st-blocked', 'st-failed', 'st-unrecorded', 's-answered', 's-failed', 's-unrecorded']) {
assert.match(
css,
new RegExp(`\\.${sel}\\s*[,{]`),
`Insights.css opens no rule for .${sel}`,
)
}
// The two chips that must never share a look: a failure and an unrecorded row
// are drawn from different declarations, not one shared block.
assert.ok(
!/\.s-failed\s*,[^{]*\.s-unrecorded|\.s-unrecorded\s*,[^{]*\.s-failed/.test(css),
'the failed and not-recorded chips share one rule — they would look identical',
)
})
+199
View File
@@ -0,0 +1,199 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { Egress } from './api.ts'
import {
DPI_TYPES,
EGRESS_TYPES,
RETIRED_EGRESS_TYPES,
UNKNOWN_EGRESS_TYPE_HINT,
isKnownEgressType,
nextEgress,
retiredEgressType,
} from './egressEdit.ts'
// The egress editor's save merge. The defect these pin: the submit handler
// cleared Interface, Port and DPI for every type it did not have a branch for —
// including types it renders no field for at all — so opening an egress the
// panel calls "(unknown)" and changing only its NAME deleted its interface. On a
// `tunnel` egress that was live damage: the data plane routed it for real, and
// `untunnelable_egress` resolves through exactly that field, so the ESP/AH/GRE/
// IGMP/SCTP carrier silently stopped existing and those protocols fell back to
// the untunnelable policy.
test('an unknown type keeps the fields the editor never showed', () => {
const initial: Egress = { Name: 'vpn', Type: 'wireguard', Interface: 'wg0', DPI: 'fragment' }
// The form as the editor would hold it for an unknown type: no Interface
// input is rendered, no port input, no DPI select. Only the name was touched.
const out = nextEgress(initial, {
name: 'vpn-renamed',
type: 'wireguard',
iface: 'wg0',
dpi: 'off',
})
assert.equal(out.Name, 'vpn-renamed')
assert.equal(out.Type, 'wireguard')
assert.equal(
out.Interface,
'wg0',
'renaming an egress whose type this panel does not know must not delete its interface',
)
assert.equal(out.DPI, 'fragment', 'nor any other field the form declined to display')
})
test('an unknown type with a blank form state still keeps what was stored', () => {
// The stricter version: the editor's `iface` state is seeded from `initial`,
// so a test that passes the same value back could pass on a broken merge too.
// Blank the form and the stored value must still survive.
const initial: Egress = { Name: 'vpn', Type: 'wireguard', Interface: 'wg0', DPI: 'spoof' }
const out = nextEgress(initial, { name: 'vpn', type: 'wireguard', iface: '', dpi: '' })
assert.equal(out.Interface, 'wg0')
assert.equal(out.DPI, 'spoof')
})
test('the inputs are not mutated — the caller keeps a usable `initial`', () => {
const initial: Egress = { Name: 'vpn', Type: 'wireguard', Interface: 'wg0' }
nextEgress(initial, { name: 'other', type: 'interface', iface: 'wan2', dpi: 'off' })
assert.deepEqual(initial, { Name: 'vpn', Type: 'wireguard', Interface: 'wg0' })
})
test('a known type still clears the fields it does not use', () => {
// The other half of the contract: for a type the editor DOES render, stale
// settings from the previous type must go, or the config keeps a value the new
// type ignores and the panel shows a setting that does nothing.
const initial: Egress = { Name: 'e', Type: 'interface', Interface: 'wan2', DPI: 'fragment' }
const out = nextEgress(initial, { name: 'e', type: 'direct', iface: 'wan2', dpi: 'record' })
assert.equal(out.Type, 'direct')
assert.equal(out.Interface, undefined, 'a direct egress binds no interface')
assert.equal(out.DPI, 'record', 'but it does carry the native preset')
})
test('an interface egress carries its interface and its DPI preset', () => {
const out = nextEgress(undefined, {
name: ' wan-direct ',
type: 'interface',
iface: ' wan2 ',
dpi: 'record',
})
assert.equal(out.Name, 'wan-direct', 'the name is trimmed')
assert.equal(out.Interface, 'wan2', 'the interface is trimmed')
assert.equal(out.DPI, 'record')
})
test('the type list is the closed set the daemon builds outbounds for', () => {
// model.KnownEgressTypes. `tunnel` must NOT be here: the daemon folds it to
// `interface` on read (model.NormalizeEgressTypes), so the panel receives the
// canonical spelling and a second entry would put the split back into the UI.
assert.deepEqual(
EGRESS_TYPES.map((t) => t.id),
['interface', 'direct'],
)
assert.equal(isKnownEgressType('tunnel'), false)
assert.equal(isKnownEgressType('interface'), true)
assert.equal(isKnownEgressType('direct'), true)
assert.deepEqual([...DPI_TYPES].sort(), ['direct', 'interface'])
})
test('the unknown-type hint describes what actually happens, both halves of it', () => {
// It has to name the ROUTING as well as the outbound — the old text said only
// "this engine builds no outbound", which was false for the one unknown type
// anybody had, because the data plane was building that egress a real routing
// table at the same time. And it has to promise what nextEgress now keeps.
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /routing rule or table/)
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /no outbound/)
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /blocked/)
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /leaves this egress’s other settings/)
assert.doesNotMatch(
UNKNOWN_EGRESS_TYPE_HINT,
/This engine builds no outbound for that type/,
'the superseded sentence claimed the engine was the only half involved',
)
})
// ---- the RETIRED type -------------------------------------------------------
//
// `byedpi` was a working egress kind. It handed traffic to a separate ciadpi
// process because the engine's own TLS fragmentation was not getting through
// DPI; the cause turned out to be a defect in that fragmentation — the cut
// always landed inside the FIRST label of the name, so the blocked word
// travelled intact — and with that fixed the external process was weight. The
// type is gone from the Go model.
//
// What these pin is the SECOND half of removing it. Routers still carry
// `option type 'byedpi'` in /etc/config/shater, and the cheap thing to do is let
// the type fall into the generic unknown branch. That branch tells an operator
// the router "does not recognise this type", which reads as a typo — so they go
// looking for a misspelling that is not there, while every rule bound to that
// egress is blocked right now.
test('a retired type is named as removed, not as unrecognised', () => {
const hint = retiredEgressType('byedpi')
assert.ok(hint, 'a stored byedpi egress must get a sentence of its own')
// (a) removed from the product — and explicitly NOT a misspelling, because
// that is the wrong hunt to send somebody on.
assert.match(hint, /REMOVED from this product/)
assert.match(hint, /not a misspelling/)
// (b) what is happening RIGHT NOW: fail-closed, not a quiet WAN leak.
assert.match(hint, /BLOCKED/)
assert.match(hint, /fail-closed/)
assert.match(hint, /never quietly sent out over the plain WAN/)
// (c) the replacement, in the words of the two controls this form has.
assert.match(hint, /Direct/)
assert.match(hint, /Interface/)
assert.match(hint, /record/)
// (d) no promise about the operator's own ISP. The daemon does not make one,
// and a caption that quietly did would be the panel contradicting it.
assert.match(hint, /property of your ISP and is not promised here/)
})
test('the retired sentence promises nothing about getting through', () => {
// The failure mode this guards is a rewrite that "helps" by upgrading the
// replacement instruction into a claim. `record` is a preset, not an outcome.
const hint = retiredEgressType('byedpi') ?? ''
for (const claim of [/will get through/i, /works with your/i, /restores/i, /fixes your/i]) {
assert.doesNotMatch(hint, claim, `the caption must not claim an outcome: ${claim}`)
}
})
test('a retired type is spelled loosely — a hand-edited UCI file is the source', () => {
assert.equal(retiredEgressType(' ByeDPI '), retiredEgressType('byedpi'))
assert.equal(retiredEgressType('BYEDPI'), retiredEgressType('byedpi'))
})
test('only the retired spellings match — no alias, no prototype, no live type', () => {
// The control for the test above: a lookup that says yes to everything would
// pass "byedpi gets a sentence" and be worthless. These must all be undefined.
for (const live of ['interface', 'direct', 'tunnel', 'wireguard', '', ' ', 'byedpi2', 'bye dpi']) {
assert.equal(retiredEgressType(live), undefined, `${JSON.stringify(live)} is not retired`)
}
// `RETIRED_EGRESS_TYPES[k]` answers these off Object.prototype; the lookup
// must not, or a config typo becomes "[object Object]" under the Type select.
for (const proto of ['toString', 'constructor', 'hasOwnProperty', '__proto__']) {
assert.equal(retiredEgressType(proto), undefined, proto)
}
assert.deepEqual(Object.keys(RETIRED_EGRESS_TYPES), ['byedpi'])
})
test('a retired type is retired — the engine list must not take it back', () => {
// The other direction, and the one that would make the sentence above a lie:
// if `byedpi` ever reads as a live type again, the editor renders fields for
// it and nextEgress starts rewriting egresses the daemon builds nothing for.
assert.equal(isKnownEgressType('byedpi'), false)
assert.ok(!EGRESS_TYPES.some((t) => t.id === 'byedpi'), 'the type list must not offer it')
assert.ok(!DPI_TYPES.has('byedpi'), 'and no preset is stamped on it either')
// A type cannot be both, or the editor would show a blurb and a removal notice.
for (const t of EGRESS_TYPES) {
assert.equal(retiredEgressType(t.id), undefined, `${t.id} is offered AND retired`)
}
})
test('a retired egress survives a rename with every stored field intact', () => {
// The whole reason `if (!isKnownEgressType(type)) return base` exists, now that
// a real router carries a type in exactly that position. Blanking the form
// state too, so a merge that echoed the inputs back could not pass this.
const initial: Egress = { Name: 'ciadpi-exit', Type: 'byedpi', Interface: 'wan', DPI: 'fragment' }
const out = nextEgress(initial, { name: 'old-desync', type: 'byedpi', iface: '', dpi: 'off' })
assert.equal(out.Name, 'old-desync')
assert.equal(out.Type, 'byedpi', 'the type is not silently rewritten either')
assert.equal(out.Interface, 'wan', 'renaming a retired egress must not delete its interface')
assert.equal(out.DPI, 'fragment')
})
+199
View File
@@ -0,0 +1,199 @@
import type { Egress } from './api'
/**
* The egress editor's data half — the closed type list and the one function that
* decides which fields a save writes.
*
* It lives outside `pages/Targets.tsx` because it is the part that must be
* TESTED, and the panel's test runner is `node --test src/*.test.ts`: plain
* modules only, no JSX, no DOM. Extracting it is not tidiness — the bug below
* shipped precisely because "which fields does Save write?" was three lines
* buried in a submit handler that nothing could call.
*/
/**
* The two egress kinds that produce a real way out, in the order the editor
* offers them. This is the panel's copy of `model.KnownEgressTypes` and must
* stay equal to it: the daemon builds no outbound for anything else, and the
* router installs no mark, no `ip rule` and no routing table for it either, so
* every node, group and rule bound to such an egress is blocked.
*
* The list is CLOSED and POSITIVE. There is no fallback entry and no "other":
* a type that is not spelled here has no fields in this editor, and
* {@link nextEgress} therefore refuses to rewrite it.
*
* `tunnel` is deliberately NOT here. It is an accepted spelling in
* `/etc/config/shater`, but the daemon folds it to `interface` on read
* (model.NormalizeEgressTypes), so an egress written that way arrives at this
* panel already saying `interface` — with its Interface field rendered, its
* blurb correct and no "(unknown)" label. Adding a third entry here would put
* the second spelling back into a UI that has to agree with two backend halves.
*
* `proxy` and `block` were removed: neither ever created an outbound, so
* everything bound to them fell through to the plain WAN with the real address.
* Send traffic through a proxy by routing it at a group/node/chain, and drop it
* with the `block` target on a rule.
*
* A third type was RETIRED for the opposite reason — it worked, and then the
* defect it was compensating for got fixed, so it became weight. It is not
* dropped into the generic unknown branch on the way out; it is named, once, in
* {@link RETIRED_EGRESS_TYPES}, which is the only place its spelling and its
* story live.
*/
export const EGRESS_TYPES: ReadonlyArray<{ id: string; label: string; blurb: string }> = [
{
id: 'interface',
label: 'Interface — out a specific WAN or tunnel',
blurb: 'Binds to one device (wan, wg0, …) so this traffic leaves over that uplink.',
},
{
id: 'direct',
label: 'Direct — straight out, with an optional DPI preset',
blurb: 'Uses the normal route. Its point is the DPI preset below, applied to what you route here.',
},
]
/** Lookup by id, or undefined when the stored type is not one this panel knows. */
export function egressTypeInfo(type: string): (typeof EGRESS_TYPES)[number] | undefined {
return EGRESS_TYPES.find((t) => t.id === type)
}
/** Whether this panel has a definition — and therefore its own fields — for the type. */
export function isKnownEgressType(type: string): boolean {
return egressTypeInfo(type) !== undefined
}
/**
* Types whose native DPI-bypass preset applies — which is now every type this
* panel knows. It is kept as its own set rather than folded into
* {@link EGRESS_TYPES} because the two lists answer different questions ("what
* may be created" vs "what carries a preset"), and they have already come apart
* once.
*/
export const DPI_TYPES: ReadonlySet<string> = new Set(['interface', 'direct'])
/**
* Types that WERE built and are not any more, with the sentence the editor shows
* for each. A CLOSED, positive table: nothing lands here by accident, and a type
* absent from it falls to {@link UNKNOWN_EGRESS_TYPE_HINT}.
*
* It exists so a router carrying `option type 'byedpi'` in /etc/config/shater is
* told what happened rather than handed the generic "not recognised" — the
* operator did not mistype anything, the kind was taken out from under them, and
* the difference decides what they do next.
*
* The daemon says the same thing at length (`model.RetiredEgressTypes`); this is
* the caption-length version and must not contradict it.
*/
export const RETIRED_EGRESS_TYPES: Readonly<Record<string, string>> = {
byedpi:
'The type “byedpi” was REMOVED from this product — this is not a misspelling, it is a kind ' +
'that no longer exists. Nothing is built for it: no outbound, no mark, no routing rule or ' +
'table, so every node, group and rule bound to this egress is BLOCKED right now — ' +
'fail-closed, never quietly sent out over the plain WAN. It existed to hand traffic to a ' +
'separate ciadpi process because the engine’s own TLS fragmentation was not getting through; ' +
'that turned out to be a defect in the fragmentation, and it is fixed. Move this egress onto ' +
'the built-in desync: set the type to Direct with the DPI preset “record”, or to Interface ' +
'with the same preset plus the interface this traffic should leave through, then apply. ' +
'Which preset gets through is a property of your ISP and is not promised here — “fragment” ' +
'and “spoof” are the other two.',
}
/**
* The retired-type sentence for a stored type, or undefined when the type is not
* a retired one.
*
* Case and surrounding space are folded, because that is what a hand-edited
* `/etc/config/shater` produces. Nothing else is: there is deliberately no alias
* or normalisation step in front of this lookup, so only the exact retired
* spellings match and no live type can ever be routed into a "this was removed"
* message.
*
* The own-property check is not ceremony: a bare `obj[key]` answers `toString`
* and `constructor` out of the prototype, and a `type` comes off a config file
* this panel does not control.
*/
export function retiredEgressType(type: string): string | undefined {
const key = type.trim().toLowerCase()
return Object.prototype.hasOwnProperty.call(RETIRED_EGRESS_TYPES, key)
? RETIRED_EGRESS_TYPES[key]
: undefined
}
/**
* What the editor shows under the Type select when the stored type is neither a
* known one nor a RETIRED one — the retired table above answers first, because
* "you mistyped something" and "we took this kind away" send an operator to
* different places. It has to describe what the router actually does with such
* an egress, and what THIS FORM does to it on save — both halves were wrong
* before.
*
* It used to read: "This engine builds no outbound for that type, so everything
* routed here is blocked. Pick one above." Two problems. It said "this engine",
* as if only the engine were involved, at a time when the data plane happily
* built an `ip rule`, a routing table and a mark bypass for a `tunnel` egress and
* `untunnelable_egress` carried live ESP/GRE out of it — so the sentence was
* flatly false for the one unknown type anybody had. And it stayed silent about
* the thing this form was doing to the egress: saving it wiped `Interface`,
* because the field is only rendered for `type === 'interface'` and the submit
* handler cleared every field it did not render. That is fixed in nextEgress
* below, and the text now says so, because a promise about saving is only worth
* making next to the code that keeps it.
*/
export const UNKNOWN_EGRESS_TYPE_HINT =
'The router does not recognise this type: it builds no outbound for it and installs no ' +
'routing rule or table, so everything routed here is blocked — never sent out over the ' +
'plain WAN. Pick a type above to fix it; saving leaves this egress’s other settings ' +
'exactly as they are until you do.'
/** The editor's form state, as strings straight out of the inputs. */
export interface EgressForm {
name: string
type: string
iface: string
dpi: string
}
/**
* Merge the form back onto the egress being edited.
*
* # The rule, and why it is the rule
*
* A save may only CLEAR a field the editor was in a position to show. For the
* known types the editor renders every field that type uses, so clearing the
* others is right: switching `interface` → `direct` must drop the stale
* interface name, or the config keeps a setting the new type ignores.
*
* For a type this panel has no definition for, the editor renders NONE of those
* fields — and used to clear all three anyway:
*
* base.Interface = type === 'interface' ? iface.trim() : undefined
*
* So opening an egress the panel calls "(unknown)", changing nothing but its
* name, and pressing Save silently deleted its `interface`. That was not
* hypothetical damage. `tunnel` was such a type, the data plane routed it for
* real, and `untunnelable_egress` pointing at it is resolved by
* netplane.UntunnelableEgressBinding through exactly that field: with the
* interface gone the binding fails, the ESP/AH/GRE/IGMP/SCTP protection quietly
* stops existing, and those protocols fall back to the untunnelable policy —
* from one rename, with no message anywhere.
*
* The daemon no longer hands this panel a `tunnel` (it is folded to `interface`
* on read), so that particular type is gone. The rule stays, and the RETIRED type
* is why it earns its keep today: routers still carry that spelling in their
* config, the editor renders none of its fields, and renaming such an egress must
* leave every stored setting on it exactly as saved.
*
* `initial` is never mutated — the caller keeps a usable object if the save
* fails.
*/
export function nextEgress(initial: Egress | undefined, form: EgressForm): Egress {
const type = form.type
const base: Egress = { ...(initial ?? ({} as Egress)), Name: form.name.trim(), Type: type }
// There is no `Target` to clear — the field is not in the Go model, so GET
// never delivers one and the spread above cannot produce one.
if (!isKnownEgressType(type)) return base
base.Interface = type === 'interface' ? form.iface.trim() : undefined
base.DPI = DPI_TYPES.has(type) ? form.dpi : undefined
return base
}
+127
View File
@@ -0,0 +1,127 @@
// findings.ts — which apply-time finding is shown where.
//
// Run with `npm test` (node's built-in test runner + native TypeScript
// stripping; no test dependency is added to the SPA, which ships inside the
// daemon binary).
//
// Two defects are pinned here.
//
// 1. THE TRUNCATION NOTE WAS UNREACHABLE. The daemon caps Status.warnings at 50
// and overwrites the last slot with an `info` note counting what it dropped.
// Overview filtered `info` away wholesale, and the settings-page route keys on
// a section (`generate`) that no page owns — so the single line telling the
// operator "you are not seeing all of it" reached no screen at all.
//
// 2. FINDINGS ABOUT AN ENTITY NEVER REACHED THAT ENTITY'S PAGE. The generator
// drops a node it cannot build and names it; the Nodes page rendered that node
// as an ordinary row with a green toggle, because it never read the findings.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
attentionFindings,
entityFindings,
findingsByName,
sectionNotes,
truncationNote,
worstSeverity,
} from './findings.ts'
import type { StatusWarning } from './api.ts'
const crit = (section: string, name: string, message = 'broken'): StatusWarning => ({
severity: 'critical',
section,
name,
message,
})
const warn = (section: string, name: string, message = 'degraded'): StatusWarning => ({
severity: 'warning',
section,
name,
message,
})
const info = (section: string, name: string, message: string): StatusWarning => ({
severity: 'info',
section,
name,
message,
})
/** Verbatim from apply/warnings.go finalizeWarnings. */
const SUPPRESSED = info(
'generate',
'',
'7 further warning(s) suppressed; run `logread -e shater` for the full list',
)
// --- the truncation note ----------------------------------------------------
test('the truncation note is found, whatever else is in the list', () => {
const note = truncationNote([crit('rule', 'a'), warn('node', 'b'), SUPPRESSED])
assert.notEqual(note, null)
assert.match(note!.message, /7 further warning/)
})
test('a whole list has no truncation note', () => {
assert.equal(truncationNote([crit('rule', 'a'), warn('node', 'b')]), null)
assert.equal(truncationNote([]), null)
assert.equal(truncationNote(undefined), null)
})
test('an ordinary info note is not mistaken for the truncation note', () => {
const notes = [info('untunnelable', 'block', 'Ping and traceroute do not work…')]
assert.equal(truncationNote(notes), null)
})
test('the truncation note is kept out of the settings-page notes it would pollute', () => {
const all = [info('generate', '', 'cache: moved to /overlay'), SUPPRESSED]
const notes = sectionNotes(all, 'generate')
assert.equal(notes.length, 1)
assert.match(notes[0].message, /cache:/)
})
test('the attention list still carries only critical and warning', () => {
const all = [crit('rule', 'a'), warn('node', 'b'), info('untunnelable', 'block', 'x'), SUPPRESSED]
const attention = attentionFindings(all)
assert.equal(attention.length, 2)
assert.ok(attention.every((w) => w.severity !== 'info'))
})
// --- per-entity findings ----------------------------------------------------
test('a page takes only the sections it owns', () => {
const all = [
crit('node', 'tokyo-01', 'parse share-link: bad scheme (skipped)'),
warn('subscription', 'qomar', 'fetch failed'),
crit('rule', 'default', 'never applies'),
info('generate', '', 'cache: x'),
]
const mine = entityFindings(all, ['node', 'subscription'])
assert.deepEqual(
mine.map((w) => w.name),
['tokyo-01', 'qomar'],
)
})
test('entity findings never include info notes', () => {
const all = [info('node', 'tokyo-01', 'just a note'), SUPPRESSED]
assert.equal(entityFindings(all, ['node', 'generate']).length, 0)
})
test('findings index by name, and global (unnamed) ones are left out', () => {
const all = [
crit('node', 'tokyo-01', 'first'),
warn('node', 'tokyo-01', 'second'),
crit('node', '', 'global to the section'),
]
const byName = findingsByName(entityFindings(all, ['node']))
assert.equal(byName.size, 1)
assert.equal(byName.get('tokyo-01')!.length, 2)
})
test('one lamp per row takes the loudest severity', () => {
assert.equal(worstSeverity([warn('node', 'a'), crit('node', 'a')]), 'critical')
assert.equal(worstSeverity([warn('node', 'a')]), 'warning')
assert.equal(worstSeverity([]), null)
})
+91 -2
View File
@@ -6,7 +6,8 @@
//
// critical / warning — something needs attention: a protection promise is
// broken, or something you configured isn't in effect. These belong on
// Overview, where the operator looks first.
// Overview, where the operator looks first — and, when they name an entity,
// ALSO on the page that owns that entity (see `entityFindings`).
//
// info — a statement ABOUT the configuration, not a problem. It never clears,
// because nothing is wrong: it is simply describing a choice that was made.
@@ -16,9 +17,49 @@
// page that never goes away and never asks for anything trains people to skim
// the list — which is exactly how a real critical finding gets missed. Anything
// standing in the findings list should be something you could act on.
//
// The one exception is carved out below: the daemon's own note that it dropped
// findings to fit the cap. It is `info` by severity and unactionable by nature,
// and it is the single most important line in the list, because it is the list
// telling you it is not the whole list.
import type { StatusWarning } from './api'
/**
* The daemon's truncation disclosure, verbatim from apply/warnings.go
* finalizeWarnings:
*
* "%d further warning(s) suppressed; run `logread -e shater` for the full list"
*
* Matched on the stable clause rather than the whole sentence so a reworded tail
* still registers. If this ever stops matching, the failure mode is a list that
* silently claims to be complete — which is why `truncationNote` is tested.
*/
const SUPPRESSED_RE = /further warning\(s\) suppressed/
/**
* The daemon's "this list is incomplete" note, or null when the list is whole.
*
* Status.warnings is capped at 50, sorted critical-first, and the last slot is
* REPLACED by an `info` note counting what was dropped. That note therefore
* arrives on the one channel the panel filtered away wholesale: `info` never
* reached Overview, and the settings-page route (`sectionNotes`) keys on
* section `generate`, which no page owns. So the single line saying "there are
* findings you are not being shown" was the only one guaranteed to be invisible.
*
* Callers must render this WITH the attention list, not instead of it.
*/
export function truncationNote(warnings: StatusWarning[] | undefined): StatusWarning | null {
return (
(warnings ?? []).find((w) => w.severity === 'info' && SUPPRESSED_RE.test(w.message)) ?? null
)
}
/** Is this the truncation disclosure rather than an ordinary note? */
function isTruncationNote(w: StatusWarning): boolean {
return w.severity === 'info' && SUPPRESSED_RE.test(w.message)
}
/** Findings that need attention — the Overview list. Info notes are excluded. */
export function attentionFindings(warnings: StatusWarning[] | undefined): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'critical' || w.severity === 'warning')
@@ -29,10 +70,58 @@ export function attentionFindings(warnings: StatusWarning[] | undefined): Status
* (e.g. `untunnelable` → the Networks page's "Other traffic" section). Only info:
* a critical/warning is an attention item and stays on Overview, so it can't be
* quietly buried on a settings page instead.
*
* The truncation note is excluded: it is about the LIST, not about any section,
* and it has its own home beside the list ({@link truncationNote}).
*/
export function sectionNotes(
warnings: StatusWarning[] | undefined,
section: string,
): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'info' && w.section === section)
return (warnings ?? []).filter(
(w) => w.severity === 'info' && w.section === section && !isTruncationNote(w),
)
}
/**
* The attention findings about entities ONE page owns — for that page to show
* beside the entities themselves.
*
* Overview is where you look when you already suspect something; a page like
* Nodes is where you look when you don't. The generator drops a node it cannot
* build — an unparseable share link, a WireGuard key materialised twice — and
* says so by name ("node \"x\": parse share-link: … (skipped)"), yet that node
* kept rendering as an ordinary row with a green toggle, because the page never
* read the findings at all. The switch says on; the engine has no such outbound.
*
* This does NOT move anything off Overview: the same finding appears in both
* places, which is correct — one list is "what is wrong with this router", the
* other is "what is wrong with this node".
*/
export function entityFindings(
warnings: StatusWarning[] | undefined,
sections: readonly string[],
): StatusWarning[] {
const want = new Set(sections)
return attentionFindings(warnings).filter((w) => want.has(w.section))
}
/** Index attention findings by entity name, for badging a row directly. Entries
* with an empty `name` are global to their section and are left out. */
export function findingsByName(findings: StatusWarning[]): Map<string, StatusWarning[]> {
const out = new Map<string, StatusWarning[]>()
for (const f of findings) {
if (!f.name) continue
const list = out.get(f.name)
if (list) list.push(f)
else out.set(f.name, [f])
}
return out
}
/** The loudest severity in a set — for a row badge that has room for one lamp. */
export function worstSeverity(findings: StatusWarning[]): 'critical' | 'warning' | null {
if (findings.some((f) => f.severity === 'critical')) return 'critical'
if (findings.length > 0) return 'warning'
return null
}
+94
View File
@@ -0,0 +1,94 @@
// The geo-data source contract the Settings page renders.
//
// Run with `npm test` (node's built-in test runner + native type stripping).
//
// THE CASE THIS FILE WAS WRITTEN FOR: five real Globals fields — GeoProvider and
// the four URLs — were consumed by generate.SetGeoProvider and had NO panel
// control at all, while the panel happily used the geo data those settings pick
// (the geosite/geoip pickers, the category suggestions). The only way to change
// where a `geoip:us` list came from was to edit /etc/config/shater over SSH.
//
// The trap in exposing them is that three of the five are conditional, and the
// daemon never errors: an unknown provider and a broken template both degrade to
// the built-in Auto chain with a warning and the router comes up. So a control
// that looked applied while the router used something else would be exactly the
// kind of lie the panel is not allowed to tell.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
GEO_CATEGORY_PLACEHOLDER,
GEO_PROVIDERS,
GEO_PROVIDER_DEFAULT,
checkGeoTemplate,
isKnownGeoProvider,
normGeoProvider,
usesCustomTemplates,
} from './geoProvider.ts'
test('the provider list is the closed five the daemon knows', () => {
assert.deepEqual(
GEO_PROVIDERS.map((p) => p.id),
['auto', 'sagernet', 'loyalsoldier', 'metacubex', 'custom'],
'generate/geosource.go — a sixth here would read as configured and run as auto',
)
})
test('an unset provider reads as auto, because that is what the router runs', () => {
assert.equal(normGeoProvider(undefined), GEO_PROVIDER_DEFAULT)
assert.equal(normGeoProvider(''), 'auto')
assert.equal(normGeoProvider(' '), 'auto')
assert.equal(normGeoProvider('AUTO'), 'auto', 'the daemon lower-cases before matching')
assert.equal(normGeoProvider(' MetaCubeX '), 'metacubex')
})
test('an unknown provider is preserved, not silently retargeted to auto', () => {
// Opening the page must not be an edit. The select marks the value instead, so
// the operator sees what is stored and chooses whether to replace it.
assert.equal(normGeoProvider('mirror-42'), 'mirror-42')
assert.equal(isKnownGeoProvider('mirror-42'), false)
assert.equal(isKnownGeoProvider('sagernet'), true)
assert.equal(isKnownGeoProvider(''), true, 'empty is auto, which is known')
})
test('only custom reads the URL templates', () => {
// On any other provider they are stored and ignored, so the editor hides them
// rather than offering a control the router does not consult.
assert.equal(usesCustomTemplates('custom'), true)
assert.equal(usesCustomTemplates(''), false)
assert.equal(usesCustomTemplates('auto'), false)
assert.equal(usesCustomTemplates('metacubex'), false)
})
test('a template without {category} is refused — the daemon will not append it', () => {
// Appending would build a plausible-looking wrong URL that 404s at fetch time,
// rather than failing where someone can read it.
const bad = checkGeoTemplate('https://mirror.test/geosite.srs')
assert.equal(bad.ok, false)
if (!bad.ok) assert.match(bad.reason, /\{category\}/)
})
test('an empty template falls that source back to the built-in chain, and says so', () => {
const empty = checkGeoTemplate(' ')
assert.equal(empty.ok, false)
if (!empty.ok) assert.match(empty.reason, /Auto chain/)
})
test('a template carrying the placeholder passes', () => {
assert.deepEqual(
checkGeoTemplate(`https://mirror.test/geosite/${GEO_CATEGORY_PLACEHOLDER}.srs`),
{ ok: true },
)
assert.deepEqual(checkGeoTemplate(' https://mirror.test/{category}.srs '), { ok: true })
})
test('every provider carries a blurb — the select is where the cost gap is said', () => {
// SagerNet's geoip is country codes only; Loyalsoldier's country lists are about
// twice the size. Picking a provider without that in front of you is picking
// blind, and `us` is ~159 000 prefixes against netflix's ~108.
for (const p of GEO_PROVIDERS) {
assert.ok(p.label.length > 0, `${p.id} needs a label`)
assert.ok(p.blurb.length > 40, `${p.id} needs a blurb that says something`)
}
})
+130
View File
@@ -0,0 +1,130 @@
/**
* The geo-data source settings' data half — the closed provider list, which of
* the four URL fields a given provider actually reads, and the validator for the
* two templates.
*
* It lives outside `pages/Settings.tsx` for the same reason `egressEdit.ts` lives
* outside `pages/Targets.tsx`: this is the part that must be TESTED, and the
* panel's runner is `node --test src/*.test.ts` — plain modules, no JSX, no DOM.
*
* Everything here mirrors `shater/generate/geosource.go`. Where the two could
* drift, this file states the daemon's behaviour rather than the panel's wish:
* the panel does not decide any of it, it only has to stop describing it wrongly.
*/
/**
* The five providers, in the order the Settings select offers them.
*
* A CLOSED, POSITIVE list on purpose. The daemon degrades an unknown value to
* `auto` with a warning and keeps running, so a free-text field here would let a
* typo look configured while the router quietly used something else.
*
* `blurb` is what the operator reads under the select, and every one of them is a
* claim about coverage and cost that the Go side backs:
* - SagerNet's geoip is COUNTRY CODES ONLY — 238 two-letter files. There is no
* `netflix`, no `google`, no `telegram` in it.
* - which is why `auto` is a split and not a single upstream: a country code
* stays on SagerNet (small, official, what every existing config already
* resolves to) and everything else goes to Loyalsoldier.
* - and why `auto` is NOT "Loyalsoldier for everything": their country lists are
* roughly twice SagerNet's size, so a blanket switch inflates the single most
* expensive case. `netflix` is ~108 prefixes; `us` is ~159 000.
*/
export const GEO_PROVIDERS: ReadonlyArray<{ id: string; label: string; blurb: string }> = [
{
id: 'auto',
label: 'Auto — country codes from SagerNet, the rest from Loyalsoldier',
blurb:
'The default, and a per-category choice rather than one upstream: a two-letter country code comes from SagerNet, every other geoip category from Loyalsoldier, and geosite from SagerNet. Existing lists resolve to exactly the same files as before.',
},
{
id: 'sagernet',
label: 'SagerNet — official, country codes only',
blurb:
'One upstream for both planes. Its geoip publishes country codes and nothing else, so a named category such as netflix or telegram does not exist here — express it as a country list, or pick another provider.',
},
{
id: 'loyalsoldier',
label: 'Loyalsoldier — all geoip, one dataset',
blurb:
'Sends every geoip category here, country codes included: one consistent dataset, at roughly twice SagerNet’s size for a country. Geosite still comes from SagerNet — Loyalsoldier publishes no .srs domain lists.',
},
{
id: 'metacubex',
label: 'MetaCubeX — the widest catalogue',
blurb:
'MetaCubeX/meta-rules-dat republishes both planes as .srs: about 1900 geosite categories and 260 geoip ones. Its geoip is Loyalsoldier’s data, so the sizes match.',
},
{
id: 'custom',
label: 'Custom — your own mirror',
blurb:
'Takes the URL templates below instead of a built-in list. For a private mirror, or a repo the built-ins do not cover.',
},
]
/** The value an empty/unset `GeoProvider` means. */
export const GEO_PROVIDER_DEFAULT = 'auto'
/**
* Normalise a stored provider for DISPLAY. Empty ⇒ `auto`, because that is what
* the daemon runs. Anything else is returned verbatim — including a value this
* list does not know, so the select can mark it rather than silently retarget the
* setting to `auto` the moment the page is opened.
*/
export function normGeoProvider(raw: string | undefined): string {
const v = (raw ?? '').trim().toLowerCase()
return v === '' ? GEO_PROVIDER_DEFAULT : v
}
/** Whether this panel has a definition — and a blurb — for the provider. */
export function isKnownGeoProvider(provider: string): boolean {
return GEO_PROVIDERS.some((p) => p.id === normGeoProvider(provider))
}
/**
* Whether the two URL TEMPLATES are read at all. Only `custom` reads them; on
* every other provider they are stored and ignored, so the editor hides them
* instead of offering a control the router does not consult.
*/
export function usesCustomTemplates(provider: string): boolean {
return normGeoProvider(provider) === 'custom'
}
/** The placeholder the daemon splices a category into. */
export const GEO_CATEGORY_PLACEHOLDER = '{category}'
export type GeoTemplateVerdict =
| { ok: true }
/** The template is unusable and that SOURCE falls back to the built-in `auto`
* chain. The daemon warns and carries on — it never fails the config. */
| { ok: false; reason: string }
/**
* Validate one `{category}` URL template the way `ValidateGeoProviderConfig`
* does, and say what the daemon will DO about a bad one.
*
* Two rules, both from the Go side:
* - empty ⇒ that source keeps using the built-in `auto` chain;
* - no `{category}` ⇒ likewise, because the daemon refuses to append the
* category itself: appending would build a plausible-looking wrong URL that
* 404s at fetch time instead of failing here where it can be read.
*
* Note the verdict is never fatal. `SetGeoProvider` returns warnings and never an
* error, so the panel must not tell the operator their router will refuse to
* start — it will start, using different data than they asked for.
*/
export function checkGeoTemplate(template: string): GeoTemplateVerdict {
const t = template.trim()
if (t === '') {
return { ok: false, reason: 'Empty — this source keeps using the built-in Auto chain.' }
}
if (!t.includes(GEO_CATEGORY_PLACEHOLDER)) {
return {
ok: false,
reason:
'No {category} placeholder, so the category cannot be spliced in — this source keeps using the built-in Auto chain.',
}
}
return { ok: true }
}
+53
View File
@@ -0,0 +1,53 @@
// Is the interception the config describes actually happening?
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// What these protect: the Networks coverage board used to light `lan` GREEN and
// say "Through the tunnel via lan" on a fresh install with the service switched
// off — because coverage was computed from /etc/config/shater and never once
// looked at the running daemon. Nothing inactive may look healthy, and nothing
// unknown may look healthy either.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { Status } from './api'
import { interceptLive } from './intercept.ts'
const st = (over: Partial<Status>): Status => ({ running: true, ...over }) as Status
test('green needs BOTH: the engine up and the full plane installed', () => {
assert.equal(interceptLive(st({ engine_running: true, plane: 'full' })), 'running')
})
test('a stopped engine is never green, whatever the config says', () => {
// The case that shipped: fresh install, service off.
assert.equal(interceptLive(st({ running: false })), 'stopped')
assert.equal(interceptLive(st({ engine_running: false, plane: 'full' })), 'stopped')
})
test('no plane, or the fail-closed hold plane, is not interception', () => {
// `none` — nothing is marked into the listener, so nothing reaches it.
assert.equal(interceptLive(st({ engine_running: true, plane: 'none' })), 'stopped')
// `hold` — the backstop is in the path and BLOCKS this traffic. Protection
// engaged is not the same as traffic carried, and the board must not say it is.
assert.equal(interceptLive(st({ engine_running: true, plane: 'hold' })), 'stopped')
})
test('unknown is its own answer and is never green', () => {
// No status at all yet.
assert.equal(interceptLive(null), 'unknown')
// A daemon that predates `plane`: running, engine up, plane unreported. This is
// the branch where `undefined !== 'none'` has burned this panel before.
assert.equal(interceptLive(st({ engine_running: true })), 'unknown')
// Daemon alive, engine liveness not reported.
assert.equal(interceptLive(st({ plane: 'full' })), 'unknown')
})
test('a negative wins without needing the other half to agree', () => {
// Engine down but plane still reported full (a stale or half-applied reading):
// the proof that nothing is intercepted is complete on its own.
assert.equal(interceptLive(st({ running: true, engine_running: false, plane: 'full' })), 'stopped')
// Plane gone but engine reported up.
assert.equal(interceptLive(st({ running: true, engine_running: true, plane: 'none' })), 'stopped')
})
+48
View File
@@ -0,0 +1,48 @@
import type { Status } from './api'
// Explicit `.ts` because this is a RUNTIME import and `npm test` runs the file
// through node's type-stripping loader, which does no extension resolution.
// tsconfig sets allowImportingTsExtensions and Vite resolves it unchanged.
import { engineState } from './planeState.ts'
/**
* Is the interception a config DESCRIBES actually happening right now?
*
* The Networks page's coverage board answered a different question than the one it
* appeared to answer. It is computed from `/etc/config/shater` alone — an enabled
* tproxy inbound naming this network — and never looked at the running daemon, so
* on a fresh install, with the service switched off and no data plane installed at
* all, it lit `lan` GREEN and said "Through the tunnel via lan". Nothing was going
* through anything.
*
* That is the failure this project keeps paying for: something inactive drawn as
* healthy, on the very page an operator opens to find out why their traffic is not
* being carried. Same rule as planeState.engineState, whose three-answer shape
* this mirrors — unknown is an unlit socket, never green.
*
* Testable on its own because it is the whole judgement: `node --test src/*.test.ts`
* cannot load a page component (JSX, CSS imports), and a truth table nothing can
* call is how the green lamp survived this long.
*/
export type InterceptLive = 'running' | 'stopped' | 'unknown'
/**
* running — the engine is up AND the full data plane is installed. Only this
* combination diverts a packet: a listener with nothing marked into it
* captures nothing, and an engine that never started has no listener.
* stopped — the engine is down, or the plane is `none` (nothing installed) or
* `hold` (the fail-closed backstop is in the path, which BLOCKS this
* traffic rather than carrying it). Configured, not happening.
* unknown — no status yet, or a daemon that reports no `plane`. Not a claim.
*
* A NEGATIVE ALWAYS WINS, and the negatives are checked first: a stopped engine or
* an absent plane each prove on their own that nothing is being intercepted, and
* neither needs the other's agreement. "Running" is the only verdict that needs
* both, because it is the only one that asserts something good.
*/
export function interceptLive(status: Status | null): InterceptLive {
const engine = engineState(status)
if (engine === 'down') return 'stopped'
if (status?.plane === 'none' || status?.plane === 'hold') return 'stopped'
if (engine === 'up' && status?.plane === 'full') return 'running'
return 'unknown'
}
+135
View File
@@ -0,0 +1,135 @@
// Rule.Kill — the per-rule policy for a target that cannot be built.
//
// Run with `npm test`. Plain module, no React, no DOM.
//
// What these protect, in the order they matter:
//
// 1. A RULE THAT FAILS OPEN IS VISIBLE. `open` sends the rule's traffic out
// direct — around the kill-switch, with the real address — when its target
// does not resolve. It was typed, round-tripped and drawn nowhere, so such a
// rule was indistinguishable from one that fails closed. killMark is what
// the row draws; a null here is an invisible bypass.
// 2. THE FAIL-CLOSED DEFAULT DRAWS NOTHING. If every row wore a badge the two
// states would be priced the same, which is the reading the mark exists to
// prevent.
// 3. AN UNREADABLE VALUE IS ITS OWN STATE. It blocks (the daemon's `default:`
// branch), but silently folding it into `closed` would hide a typo AND let
// the editor rewrite a value it never showed.
// 4. THE POLICY SURVIVES AN EDIT. carryKill is what the edit form writes back;
// it must not lose `open` and must not churn `''`/`default` into `closed`.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
carryKill,
killFallbackTarget,
killMark,
killPolicy,
killSelectValue,
KILL_OPTIONS,
} from './killPolicy.ts'
// --- 1. classification mirrors generate/route.go ruleKillFallback -------------
test('every spelling the daemon treats as fail-closed reads as closed here', () => {
for (const raw of ['', ' ', 'default', 'closed', 'CLOSED', ' Default ']) {
assert.equal(killPolicy(raw), 'closed', `${JSON.stringify(raw)} must be closed`)
assert.equal(killFallbackTarget(raw), 'block')
}
assert.equal(killPolicy(undefined), 'closed')
assert.equal(killPolicy(null), 'closed')
})
test('open is open in any casing, and it is the only value that routes direct', () => {
for (const raw of ['open', 'OPEN', ' Open ']) {
assert.equal(killPolicy(raw), 'open')
assert.equal(killFallbackTarget(raw), 'direct')
}
})
test('an unrecognised value is its own state and still blocks', () => {
assert.equal(killPolicy('opne'), 'unknown')
assert.equal(killPolicy('allow'), 'unknown')
// The daemon: "an unreadable policy must not open a bypass".
assert.equal(killFallbackTarget('opne'), 'block')
})
// --- 2. the row mark ---------------------------------------------------------
test('a rule that fails OPEN is marked on the row, and the mark says it leaves direct', () => {
const mark = killMark({ Kill: 'open' })
assert.notEqual(mark, null, 'kill=open must draw a mark — an unmarked bypass is invisible')
assert.match(mark!.badge, /open/i)
assert.match(mark!.badge, /direct/i)
// The note has to price it, not just name it: around the tunnel, real address.
assert.match(mark!.note, /direct/i)
assert.match(mark!.note, /real IP/i)
assert.match(mark!.note, /kill-switch/i)
})
test('the fail-closed default draws NO mark', () => {
for (const raw of ['', 'default', 'closed', undefined]) {
assert.equal(killMark({ Kill: raw }), null, `${JSON.stringify(raw)} must draw nothing`)
}
})
test('an unreadable policy is marked, quotes the value, and says it blocks', () => {
const mark = killMark({ Kill: 'opne' })
assert.notEqual(mark, null)
assert.match(mark!.badge, /closed/i)
assert.ok(mark!.note.includes('opne'), 'the note must quote what was actually written')
assert.match(mark!.note, /blocked/i)
})
// --- 3. the editor can show it, and offers only what the daemon reads ---------
test('the picker offers exactly the two outcomes the daemon has', () => {
assert.deepEqual(
KILL_OPTIONS.map((o) => o.value),
['closed', 'open'],
)
// The dangerous one must say so in its own label, not only in a warning below.
const open = KILL_OPTIONS.find((o) => o.value === 'open')!
assert.match(open.label, /kill-switch/i)
})
test('the editor preselects the stored policy, and re-surfaces an unknown one verbatim', () => {
assert.equal(killSelectValue(''), 'closed')
assert.equal(killSelectValue('default'), 'closed')
assert.equal(killSelectValue('CLOSED'), 'closed')
assert.equal(killSelectValue('OPEN'), 'open')
// Verbatim, so the <option> the form renders carries the operator's own text.
assert.equal(killSelectValue(' opne '), 'opne')
})
// --- 4. a save must not lose it ----------------------------------------------
test('editing a rule that fails open keeps it failing open', () => {
assert.equal(carryKill('open', 'open'), 'open')
})
test('a no-op edit does not rewrite the stored spelling', () => {
// model/uci.go hands the panel "default" for a rule with no `option kill` at
// all. Rewriting that into "closed" would put an explicit option on every rule
// anyone ever opened the form for, and show a no-op edit as a config change.
assert.equal(carryKill('default', 'closed'), 'default')
assert.equal(carryKill('', 'closed'), '')
assert.equal(carryKill(undefined, 'closed'), '')
assert.equal(carryKill('closed', 'closed'), 'closed')
})
test('a real change of policy is written, in both directions', () => {
assert.equal(carryKill('default', 'open'), 'open')
assert.equal(carryKill('open', 'closed'), 'closed')
})
test('an unknown policy is not silently normalised by a save that did not touch it', () => {
// The form shows "opne" as its own option, so leaving it alone must leave it
// alone. Folding it to `closed` here would be the editor rewriting a value it
// did show, on a save the operator made about something else entirely.
assert.equal(carryKill('opne', 'opne'), 'opne')
// …but picking a real option out of that state does change it.
assert.equal(carryKill('opne', 'closed'), 'closed')
assert.equal(carryKill('opne', 'open'), 'open')
})
+131
View File
@@ -0,0 +1,131 @@
import type { Rule } from './api'
/**
* `Rule.Kill` — what a rule does when its target cannot be built.
*
* WHY THIS MODULE EXISTS. The field was typed, round-tripped and completely
* invisible: no editor, and no mark on the rule row. A rule that fails OPEN —
* i.e. sends its traffic out with the real address when its group/chain/node
* cannot be resolved, deliberately around the kill-switch — was drawn exactly
* like one that fails closed. The single most consequential per-rule setting on
* the page was the one thing the page did not draw.
*
* The behaviour mirrored here is `generate/route.go ruleKillFallback`, which is
* consulted only when the target does not resolve (a dead group, a chain that
* would not assemble, a missing egress/node). The rule is ALWAYS still emitted —
* its traffic never falls through to the default route — so the only question is
* which of two outbounds it gets:
*
* "" | "default" | "closed" → block (fail-closed)
* "open" → direct (a warned, deliberate kill-switch bypass)
* anything else → block, and the daemon warns that the policy was
* unreadable ("an unreadable policy must not open
* a bypass")
*
* It lives outside `pages/Routing.tsx` because it is the part that must be
* TESTED, and the panel's runner is `node --test src/*.test.ts`: plain modules
* only, no JSX, no DOM (same reason as ruleset.ts and egressEdit.ts).
*/
/**
* The three states the panel draws. Positive and closed on purpose: `unknown` is
* a state of its own rather than folded into `closed`, because those two look
* identical to the router and completely different to the operator — one is a
* choice, the other is a typo that happens to land on the safe side.
*/
export type KillPolicy = 'closed' | 'open' | 'unknown'
/** Classify a stored `Kill` exactly as ruleKillFallback switches on it. */
export function killPolicy(raw: string | undefined | null): KillPolicy {
const v = (raw ?? '').trim().toLowerCase()
if (v === '' || v === 'default' || v === 'closed') return 'closed'
if (v === 'open') return 'open'
return 'unknown'
}
/** Where this policy actually sends the traffic. `unknown` blocks, like the daemon. */
export function killFallbackTarget(raw: string | undefined | null): 'block' | 'direct' {
return killPolicy(raw) === 'open' ? 'direct' : 'block'
}
/** The two values the editor offers. `unknown` is surfaced separately (see killSelectValue). */
export const KILL_OPTIONS: ReadonlyArray<{ value: 'closed' | 'open'; label: string }> = [
{ value: 'closed', label: 'Block it — fail closed (default)' },
{ value: 'open', label: 'Send it direct — bypasses the kill-switch' },
]
/**
* The `<select>` value that represents this stored policy.
*
* An unrecognised value comes back VERBATIM so the editor can offer it as its own
* option (the way TargetOptions re-surfaces a target pointing at a since-removed
* node). Folding it to `closed` here would mean the picker silently rewrote a
* value it never showed — the one thing a save is not allowed to do — and would
* also erase the evidence of the typo the daemon is warning about.
*/
export function killSelectValue(raw: string | undefined | null): string {
const p = killPolicy(raw)
if (p === 'unknown') return (raw ?? '').trim()
return p
}
/**
* What a save writes back, given what was stored and what the operator picked.
*
* `""`, `"default"` and `"closed"` are the SAME policy, and the picker shows them
* as one option, so choosing that option on a rule that already had one of them
* must leave the stored spelling alone. Rewriting `default` (what model/uci.go
* hands the panel for a rule with no `option kill` at all) into `closed` would
* put an explicit option on every rule anyone ever opened the editor for, and
* make a no-op edit show up as a config change.
*
* Every other transition writes the picked value: it is a real change of policy.
*/
export function carryKill(stored: string | undefined | null, picked: string): string {
if (picked === 'closed' && killPolicy(stored) === 'closed') return (stored ?? '').trim()
return picked
}
/** One mark on a rule row: the pill, and the sentence under it. */
export interface KillMark {
/** Pill text. Uppercased by the stylesheet — keep it short. */
badge: string
/** The sentence that says what happens and why it matters. */
note: string
}
/**
* The row mark for a rule's kill policy, or `null` when there is nothing to say.
*
* `closed` draws NOTHING. It is the default, it is the safe side, and a badge on
* every row would price the two states the same — which is exactly the reading
* this mark exists to prevent. Only a rule that has been moved off the safe side,
* or one whose policy cannot be read, earns a mark.
*/
export function killMark(rule: Pick<Rule, 'Kill'>): KillMark | null {
const raw = (rule.Kill ?? '').trim()
switch (killPolicy(raw)) {
case 'closed':
return null
case 'open':
return {
badge: 'fails open · direct',
note:
'If this rule’s target can’t be built — a dead group, a missing node, a chain that won’t ' +
'assemble — its traffic leaves direct instead of being blocked: around the tunnel, with ' +
'your real IP address. That is a deliberate kill-switch bypass for this rule alone.',
}
case 'unknown':
return {
badge: 'fails closed · unreadable',
note:
`“${raw}” is not a policy this router reads, so if this rule’s target can’t be built its ` +
'traffic is blocked — the safe side, but not a setting anyone chose. Edit the rule and ' +
'pick Block or Direct.',
}
}
}
/** The flag shown in the rule form while `open` is selected. */
export const KILL_OPEN_FORM_WARN =
'fails open — if this target breaks, the traffic leaves direct with your real IP'
+319
View File
@@ -0,0 +1,319 @@
// Tests for logRoute.ts — the readings that keep the two logs from lying.
//
// Three things are being defended here, and each has an explicit CONTROL:
//
// 1. the empty string is "not recorded", never "no rule" / never "local";
// 2. the search finds what it should AND does not find what it must not
// (a filter that matched everything would pass a one-directional test);
// 3. "the scan stopped" and "the data ran out" never render the same.
//
// Run: npm test (node --test src/*.test.ts)
import test from 'node:test'
import assert from 'node:assert/strict'
import {
CONN_SEARCH_HINT,
LOG_SEARCH_HINT,
chainPath,
connRule,
connSearchFields,
dnsOutbound,
historyExhausted,
logCountLabel,
logEndNote,
logSearchFields,
resumeCursor,
rowMatches,
} from './logRoute.ts'
import type { ConnLogEntry, QueryLogEntry } from './api.ts'
// ---- fixtures ---------------------------------------------------------------
function conn(over: Partial<ConnLogEntry> = {}): ConnLogEntry {
return {
unix: 1_700_000_000,
src_ip: '192.168.1.50',
src_name: 'laptop',
dest: 'youtube.com',
dest_ip: '142.250.74.238',
port: 443,
network: 'tcp',
proto: 'tls',
outbound: 'nl-reality-1',
seq: 7,
rule_kind: 'matched',
rule: 'protocol=tls domain_suffix=youtube.com',
chain: ['nl-reality-1', 'auto'],
...over,
}
}
function query(over: Partial<QueryLogEntry> = {}): QueryLogEntry {
return {
time: '12:00:00',
unix: 1_700_000_000,
domain: 'youtube.com',
qtype: 'A',
rcode: 0,
blocked: false,
server: 'cloudflare-doh',
action: 'proxy',
device: 'laptop',
seq: 7,
status: 'answered',
outbound_kind: 'detour',
outbound: 'nl-reality-1',
...over,
}
}
// ---- 1 · rule_kind: three states, three readings -----------------------------
test('connRule: a matched rule shows the engine match condition', () => {
const r = connRule(conn())
assert.equal(r.kind, 'matched')
assert.equal(r.text, 'protocol=tls domain_suffix=youtube.com')
})
test('connRule: "no rule matched" is a recorded fact with its own words', () => {
const r = connRule(conn({ rule_kind: 'default', rule: '' }))
assert.equal(r.kind, 'default')
assert.match(r.text, /default route/)
assert.match(r.title, /recorded fact/)
})
test('connRule: an empty rule_kind is NOT RECORDED, not "no rule"', () => {
const r = connRule(conn({ rule_kind: '', rule: '' }))
assert.equal(r.kind, 'unrecorded')
assert.match(r.label, /not recorded/)
assert.match(r.title, /does NOT mean no rule matched/)
})
// CONTROL for the three above: it is not enough that each reads sensibly on its
// own — the two rule-less states must be DISTINGUISHABLE. A panel that printed
// an empty cell for both would pass every assertion above.
test('CONTROL connRule: unrecorded and default never render alike', () => {
const unrecorded = connRule(conn({ rule_kind: '', rule: '' }))
const noRule = connRule(conn({ rule_kind: 'default', rule: '' }))
const matched = connRule(conn())
assert.notEqual(unrecorded.kind, noRule.kind)
assert.notEqual(unrecorded.label, noRule.label)
assert.notEqual(unrecorded.title, noRule.title)
const kinds = new Set([unrecorded.kind, noRule.kind, matched.kind])
assert.equal(kinds.size, 3, 'rule_kind must produce three distinct readings')
})
test('connRule: an unrecognised rule_kind falls to unrecorded, the recoverable side', () => {
assert.equal(connRule(conn({ rule_kind: 'tomorrows-word' })).kind, 'unrecorded')
})
test('chainPath: reads rule-named tag first, dialling outbound last', () => {
assert.equal(chainPath(['nl-reality-1', 'auto']), 'auto → nl-reality-1')
assert.equal(chainPath([]), '')
assert.equal(chainPath(undefined), '')
})
// ---- 2 · outbound_kind: four states, four readings ---------------------------
test('dnsOutbound: detour names the tag the lookup left through', () => {
const r = dnsOutbound(query())
assert.equal(r.kind, 'detour')
assert.equal(r.tag, 'nl-reality-1')
})
test('dnsOutbound: default says the lookup went out past the tunnel', () => {
const r = dnsOutbound(query({ outbound_kind: 'default', outbound: '' }))
assert.equal(r.kind, 'default')
assert.match(r.label, /default WAN/)
assert.match(r.title, /outside the tunnel/)
})
test('dnsOutbound: local means nothing egressed at all', () => {
const r = dnsOutbound(query({ outbound_kind: 'local', outbound: '' }))
assert.equal(r.kind, 'local')
assert.match(r.title, /Nothing left the box/)
})
test('dnsOutbound: an empty outbound_kind is NOT RECORDED, not "local"', () => {
const r = dnsOutbound(query({ outbound_kind: '', outbound: '' }))
assert.equal(r.kind, 'unrecorded')
assert.match(r.title, /does NOT mean it stayed on the router/)
})
// CONTROL: all four must be mutually distinguishable. The dangerous collapses are
// default↔unrecorded (both have no tag) and local↔unrecorded (both "went nowhere"
// if you squint), so assert the whole set at once rather than one pair.
test('CONTROL dnsOutbound: four states, four distinct labels', () => {
const rs = ['detour', 'default', 'local', ''].map((k) =>
dnsOutbound(query({ outbound_kind: k, outbound: k === 'detour' ? 'nl-reality-1' : '' })),
)
assert.equal(new Set(rs.map((r) => r.kind)).size, 4)
assert.equal(new Set(rs.map((r) => r.label)).size, 4)
assert.equal(new Set(rs.map((r) => r.title)).size, 4)
})
test('dnsOutbound: an unrecognised outbound_kind falls to unrecorded', () => {
assert.equal(dnsOutbound(query({ outbound_kind: 'tomorrows-word' })).kind, 'unrecorded')
})
// ---- 3 · the search: finds, and does not find --------------------------------
test('search FINDS: a domain in the query log', () => {
assert.equal(rowMatches(logSearchFields(query()), 'YOUTUBE'), true)
})
test('search FINDS: a device, a resolver and an exit tag', () => {
const f = logSearchFields(query())
assert.equal(rowMatches(f, 'laptop'), true)
assert.equal(rowMatches(f, 'cloudflare'), true)
assert.equal(rowMatches(f, 'reality'), true)
})
test('search FINDS: a domain through the rule that routed it', () => {
const f = connSearchFields(conn({ dest: '142.250.74.238', dest_ip: '142.250.74.238' }))
assert.equal(rowMatches(f, 'youtube.com'), true, 'the rule text carries the domain')
})
test('search FINDS: a chain hop', () => {
assert.equal(rowMatches(connSearchFields(conn()), 'auto'), true)
})
// CONTROL for every assertion above: a matcher that returned true unconditionally
// passes all of them. These are the searches that MUST come back empty — the two
// vocabulary words and the port, excluded by the daemon on purpose.
test('CONTROL search DOES NOT FIND: outbound_kind is not searched', () => {
const row = query({ outbound_kind: 'default', outbound: '', server: 'router-local', action: 'pass', device: 'tv', domain: 'example.org', qtype: 'A' })
assert.equal(
rowMatches(logSearchFields(row), 'default'),
false,
'q=default must not match every default-egress row',
)
})
test('CONTROL search DOES NOT FIND: rule_kind is not searched', () => {
const row = conn({ rule_kind: 'matched', rule: 'port=443', dest: 'a.example', dest_ip: '1.2.3.4', outbound: 'wan', chain: undefined, src_name: 'tv', network: 'tcp', proto: 'tls' })
assert.equal(rowMatches(connSearchFields(row), 'matched'), false)
})
test('CONTROL search DOES NOT FIND: the port number is not searched', () => {
const row = conn({ port: 8443, dest: 'a.example', dest_ip: '10.0.0.1', rule: 'domain=a.example', chain: undefined, outbound: 'wan', src_ip: '192.168.1.9', src_name: 'tv' })
assert.equal(rowMatches(connSearchFields(row), '8443'), false)
})
test('CONTROL search DOES NOT FIND: a string that is in no searched field', () => {
assert.equal(rowMatches(logSearchFields(query()), 'facebook'), false)
})
test('search: the field lists are exactly the daemon\'s (filter.go matchLog/matchConn)', () => {
assert.deepEqual(logSearchFields(query({ status: 'failed', error: 'i/o timeout' })), [
'youtube.com',
'A',
'cloudflare-doh',
'proxy',
'laptop',
'nl-reality-1',
'i/o timeout',
])
assert.deepEqual(connSearchFields(conn()), [
'192.168.1.50',
'laptop',
'youtube.com',
'142.250.74.238',
'tcp',
'tls',
'nl-reality-1',
'protocol=tls domain_suffix=youtube.com',
'nl-reality-1',
'auto',
])
})
test('search hints name what is searched', () => {
assert.match(LOG_SEARCH_HINT, /domain/)
assert.match(CONN_SEARCH_HINT, /rule/)
})
// ---- 4 · a truncated page is not the end of the log --------------------------
test('historyExhausted: a short unfiltered page really is the end', () => {
assert.equal(historyExhausted({ returned: 3, page: 100, truncated: false }), true)
})
test('historyExhausted: a short TRUNCATED page is NOT the end', () => {
assert.equal(
historyExhausted({ returned: 3, page: 100, truncated: true }),
false,
'the scan budget ran out — there is more log behind it',
)
})
test('historyExhausted: an empty truncated page is still not the end', () => {
assert.equal(historyExhausted({ returned: 0, page: 100, truncated: true }), false)
})
// CONTROL: the two situations must reach DIFFERENT verdicts from the same page
// shape. A hook that ignored `truncated` would answer "exhausted" to both, hide
// the paging control and call a partial search complete.
test('CONTROL historyExhausted: budget-ended and data-ended differ on identical pages', () => {
const shape = { returned: 3, page: 100 }
assert.notEqual(
historyExhausted({ ...shape, truncated: true }),
historyExhausted({ ...shape, truncated: false }),
)
})
test('resumeCursor: the daemon scan cursor wins over the oldest row', () => {
assert.equal(resumeCursor({ cursor: 41, rows: [{ seq: 90 }, { seq: 88 }] }), 41)
})
test('resumeCursor: an empty truncated page can still be continued', () => {
assert.equal(resumeCursor({ cursor: 41, rows: [] }), 41, 'no row seq exists — only the cursor can carry on')
})
test('resumeCursor: falls back to the oldest row, then to nowhere', () => {
assert.equal(resumeCursor({ cursor: 0, rows: [{ seq: 90 }, { seq: 88 }] }), 88)
assert.equal(resumeCursor({ cursor: 0, rows: [] }), null)
})
test('logEndNote: a completed empty search says nothing matched', () => {
const n = logEndNote({ filter: 'vk.com', truncated: false, matches: 0 })
assert.equal(n.complete, true)
assert.match(n.text, /Nothing in the log matches/)
})
test('logEndNote: a truncated empty search says it STOPPED, not that nothing exists', () => {
const n = logEndNote({ filter: 'vk.com', truncated: true, matches: 0 })
assert.equal(n.complete, false)
assert.match(n.text, /stopped/)
assert.match(n.text, /not the end of the log/)
assert.doesNotMatch(n.text, /Nothing in the log matches/)
})
// CONTROL: same filter, same zero rows, one bit of difference — the two must not
// produce the same sentence or the same `complete` verdict. This is the assertion
// that fails if the panel draws "budget ran out" and "data ran out" identically.
test('CONTROL logEndNote: zero matches reads differently when the scan stopped early', () => {
const ended = logEndNote({ filter: 'vk.com', truncated: false, matches: 0 })
const stopped = logEndNote({ filter: 'vk.com', truncated: true, matches: 0 })
assert.notEqual(ended.text, stopped.text)
assert.notEqual(ended.complete, stopped.complete)
})
test('logCountLabel: never claims "all loaded" on a truncated scan', () => {
const t = logCountLabel({ rows: 12, showMore: false, truncated: true, filtered: true })
assert.doesNotMatch(t, /all loaded/)
assert.match(t, /scan incomplete/)
})
// CONTROL: the un-truncated version of the very same page DOES say "all loaded",
// so the assertion above is about truncation and not about the wording never
// appearing at all.
test('CONTROL logCountLabel: the same page says "all loaded" when the scan completed', () => {
const shape = { rows: 12, showMore: false, filtered: true }
assert.match(logCountLabel({ ...shape, truncated: false }), /all loaded/)
assert.notEqual(
logCountLabel({ ...shape, truncated: false }),
logCountLabel({ ...shape, truncated: true }),
)
})
+560
View File
@@ -0,0 +1,560 @@
// Reading the two log streams honestly: WHY a connection went where it went,
// WHERE a DNS lookup left the box, what the search box does and does not cover,
// and — the one that matters most — whether a short page is the end of the log
// or a scan that stopped early.
//
// Everything here is pure so it can be tested without a DOM. The page renders
// these readings; it never re-derives them, because each of the three closed
// vocabularies below has the same trap in it: the EMPTY STRING IS RESERVED FOR
// "NOT RECORDED", and a reader that folds it into the neighbouring value turns
// "we don't know" into a confident wrong answer.
//
// Backend contract: shater/stats/stats.go (ConnLogEntry.RuleKind,
// LogEntry.OutboundKind), shater/stats/filter.go (matchLog/matchConn),
// shater/stats/store.go (LogPage.Truncated/ScanCursor).
import type { ConnLogEntry, QueryLogEntry } from './api'
// ---- why a connection went where it went -----------------------------------
/** The closed vocabulary of `ConnLogEntry.rule_kind`, normalised. */
export type ConnRuleKind = 'matched' | 'default' | 'unrecorded'
export interface ConnRuleReading {
kind: ConnRuleKind
/** Short chip label — what KIND of answer this is. */
label: string
/** The detail line: the engine's match condition, or the statement itself. */
text: string
/** Tooltip: the full sentence, including what the state does NOT mean. */
title: string
}
/**
* How to render one connection's routing record.
*
* The list is POSITIVE and CLOSED: anything the daemon might send that is not
* one of the two recorded words lands on `unrecorded`, the recoverable side. An
* open `default:` here would quietly file an unknown future value under "no rule
* matched" — a sentence about the config that nobody wrote.
*/
export function connRule(
e: Pick<ConnLogEntry, 'rule_kind' | 'rule'>,
): ConnRuleReading {
const rule = (e.rule ?? '').trim()
switch (e.rule_kind) {
case 'matched':
return {
kind: 'matched',
label: 'rule',
// A matched row without condition text is degenerate, but saying so beats
// printing an empty cell that reads as "no rule".
text: rule || 'condition not recorded',
title: rule
? `A routing rule matched this connection: ${rule}. This is the engine's match condition, not the rule name from the config.`
: 'A routing rule matched this connection, but its match condition was not recorded.',
}
case 'default':
return {
kind: 'default',
label: 'no rule',
text: 'default route',
title:
'No routing rule matched — the connection took the default route. A recorded fact, not missing data.',
}
default:
return {
kind: 'unrecorded',
label: 'not recorded',
text: '',
title:
'Nothing was recorded about routing for this row. It does NOT mean no rule matched — that state says so by name.',
}
}
}
/**
* The outbound path a connection took, written the way a person reads it: the
* tag the rule named first, the outbound that actually dialled last.
*
* The wire order is the opposite (`chain[0]` is the dialler), so this reverses
* it. '' when there is no path to show — a single-hop route has none, and an
* absent key is not an empty path to draw.
*/
export function chainPath(chain?: string[] | null): string {
if (!chain || chain.length === 0) return ''
return [...chain].reverse().join(' → ')
}
// ---- did the connection go anywhere at all? ---------------------------------
/**
* The closed vocabulary of a connection's FATE — a second axis, orthogonal to
* {@link connRule}, and the one the operator is actually sorting rows by.
*
* killed — the route sent this connection to the engine's `block`
* outbound. Nothing left the router for it.
* carried — it left through some named outbound. Says which WAY it went,
* never that it worked; the log records the route decision, not
* the result of the flow.
* unrecorded — no outbound tag on the row. NOT a claim that it was carried,
* and NOT a claim that it was killed.
*/
export type ConnFate = 'killed' | 'carried' | 'unrecorded'
export interface ConnFateReading {
fate: ConnFate
/** Short chip label. The three states must not share one. */
label: string
title: string
}
/**
* `block` is the engine's own reserved outbound tag (generate/generate.go
* tagBlock), and generate/outbound.go + generate/group.go SKIP any node or group
* whose name collides with it — so the tag on a connection row can only be the
* block outbound. That reservation is an EXACT string comparison in the
* generator, so the comparison here is exact too: a node someone named `Block`
* is emitted under that tag and is a real destination, not a kill.
*
* WHY this axis exists. `outbound:"block"` is a legitimate rule target and the
* value of route.Final on a fail-closed box, so it is what a kill-switch drop
* looks like in the log — and it used to be drawn as plain mono text,
* indistinguishable from `outbound:"nl-reality-1"`, on the same page where a DNS
* row for the same host gets a crit rail and a BLOCK mark. The connection log is
* the first place a "this site does not open" report is taken to, and the row
* that IS the answer must not read as traffic that was carried.
*
* It deliberately does NOT say WHICH decision killed it — a rule with target
* Block, or the default route on a box with no matching rule. That is the
* {@link connRule} chip beside it, and duplicating it here could only disagree
* with it.
*
* POSITIVE and CLOSED, like {@link connRule}: anything that is not the reserved
* tag and not empty is `carried` — a named outbound this build does not have to
* recognise to report honestly.
*/
export function connFate(e: Pick<ConnLogEntry, 'outbound'>): ConnFateReading {
const tag = (e.outbound ?? '').trim()
if (tag === 'block') {
return {
fate: 'killed',
label: 'killed',
title:
'The router did NOT carry this connection: the route sent it to the block outbound, so nothing left the box for it. The chip beside this says which decision did it — a routing rule, or the default route.',
}
}
if (tag === '') {
return {
fate: 'unrecorded',
label: 'no exit recorded',
title:
'No outbound was recorded for this connection. It is NOT a claim that it was carried, and NOT a claim that it was blocked — nothing about its exit is known.',
}
}
return {
fate: 'carried',
label: 'carried',
title: `The route sent this connection out through ${tag}. That is which WAY it went — the log records the routing decision, not whether the flow then succeeded.`,
}
}
// ---- where a DNS lookup left the box ---------------------------------------
/** The closed vocabulary of `QueryLogEntry.outbound_kind`, normalised. */
export type DnsOutKind = 'detour' | 'default' | 'local' | 'unrecorded'
export interface DnsOutReading {
kind: DnsOutKind
/** Short chip label — the four states must not share one. */
label: string
/** The tag, when there is one to name; '' otherwise. */
tag: string
title: string
}
/**
* How to render where one DNS lookup went out.
*
* Four states, four labels. `default` is the one an operator has to be able to
* see at a glance: it means the resolver names no detour, so the lookup left
* over the router's own WAN — past the tunnel — which is exactly the leak an
* anti-leak `detour` is configured to prevent. It is still not an ERROR (a box
* with no anti-leak intent reads `default` all day), so it is marked, not alarmed.
*
* POSITIVE and CLOSED, same as {@link connRule}: an unrecognised value is
* `unrecorded`.
*/
export function dnsOutbound(
e: Pick<QueryLogEntry, 'outbound_kind' | 'outbound'>,
): DnsOutReading {
const tag = (e.outbound ?? '').trim()
switch (e.outbound_kind) {
case 'detour':
return {
kind: 'detour',
label: 'via',
tag,
title: tag
? `The resolver that answered is bound to a detour: this lookup's own packets left through ${tag}. It is the configured binding, so a group tag names the group, not the member that was live.`
: 'The resolver that answered is bound to a detour, but the outbound tag was not recorded.',
}
case 'default':
return {
kind: 'default',
label: 'default WAN',
tag: '',
title:
"The resolver that answered names no detour, so this lookup left over the router's own WAN — outside the tunnel. A recorded fact: on a box configured for DNS anti-leak it is the one to look at.",
}
case 'local':
return {
kind: 'local',
label: 'local',
tag: '',
title:
'Answered on the router — a cache hit, an optimistic answer, or a filter block. Nothing left the box for this query.',
}
default:
return {
kind: 'unrecorded',
label: 'not recorded',
tag: '',
title:
'Where this lookup went out was not recorded. It does NOT mean it stayed on the router — that state says so by name.',
}
}
}
// ---- what came of a DNS lookup, and which way it went ----------------------
/** The OUTCOME axis — the closed vocabulary of `QueryLogEntry.status`. */
export type DnsOutcomeKind = 'answered' | 'failed' | 'unrecorded'
/** The PATH axis — the closed vocabulary of `QueryLogEntry.action`, plus the
* recoverable side for anything this build does not recognise. */
export type DnsPathKind = 'block' | 'proxy' | 'pass' | 'unknown'
/**
* The FOUR SITUATIONS one DNS row can be in, as one word for the row itself.
* They are what the reader is actually sorting rows into, and no two of them may
* ever be drawn the same way:
*
* answered — an answer came back and the filter did not make it.
* blocked — the filter answered it on purpose. Intended, not a fault.
* failed — no usable answer. A fault, whichever path it took.
* unrecorded — the outcome is not known. Neither of the three above.
*/
export type DnsRowState = 'answered' | 'blocked' | 'failed' | 'unrecorded'
/**
* Everything one DNS row draws about outcome and path. The page renders this and
* derives nothing of its own, so the four situations cannot quietly collapse into
* three in one component while the tests check another.
*/
export interface DnsRowMark {
/** The row's own state — drives the row-level marking (`st-<state>`). */
state: DnsRowState
/** The outcome chip: what came of the lookup (`s-<outcome>`). */
outcome: DnsOutcomeKind
label: string
title: string
/** The path chip: which way it went (`.tag <path>`). Kept VISIBLE in every
* state — "it failed" and "it failed in the tunnel" are different reports. */
path: DnsPathKind
pathLabel: string
pathTitle: string
/** The failure cause, already phrased. '' unless the outcome is `failed`, and
* never empty when it is: a failure with no recorded cause says so. */
cause: string
/** What the rcode adds to a failure: whether the server answered at all.
* '' unless the outcome is `failed`. */
rcodeNote: string
}
/** The outcome axis alone. POSITIVE and CLOSED: an unrecognised value — including
* a future one — lands on `unrecorded`, which is the side you can recover from.
* The old reading had an open `return 'pass'` here, so every unknown value came
* out green. */
function dnsOutcomeKind(status: string | undefined): DnsOutcomeKind {
switch (status) {
case 'answered':
return 'answered'
case 'failed':
return 'failed'
default:
return 'unrecorded'
}
}
/** The path axis alone. Same discipline, same reason. */
function dnsPathKind(action: string | undefined): DnsPathKind {
switch (action) {
case 'block':
return 'block'
case 'proxy':
return 'proxy'
case 'pass':
return 'pass'
default:
return 'unknown'
}
}
/**
* How to render one DNS row's outcome and path.
*
* The two are separate axes on purpose. `action` says WHICH WAY the lookup went;
* `status` says WHAT CAME OF IT. A lookup that left through a detour and then
* timed out is `proxy` AND `failed`, and that pair is the single most useful row
* in the log — it names the tunnel as the thing that broke. When the outcome axis
* did not exist, that row was drawn from `action` alone: an accent-coloured
* `proxy` tag, no other marking, the healthiest-looking row on the page.
*
* So the OUTCOME leads and the PATH stays beside it, never the other way round.
*/
export function dnsRowMark(
e: Pick<QueryLogEntry, 'status' | 'error' | 'rcode' | 'action' | 'blocked'>,
): DnsRowMark {
const outcome = dnsOutcomeKind(e.status)
const path = dnsPathKind(e.action)
const rawAction = (e.action ?? '').trim()
const pathTitle =
path === 'unknown'
? rawAction
? `This build does not recognise the recorded path “${rawAction}”, so nothing is claimed about which way this lookup went.`
: 'Which way this lookup went was not recorded.'
: path === 'block'
? 'The DNS filter answered this lookup itself — nothing was carried out of the box for it.'
: path === 'proxy'
? 'This lookup took the proxied path. It says which WAY it went, not whether it worked — read the outcome beside it.'
: 'This lookup took the direct path. It says which WAY it went, not whether it worked — read the outcome beside it.'
// The row state folds in the one fact the outcome axis deliberately does not
// carry: a blocked lookup IS answered, but "the filter answered it" and "the
// resolver answered it" are not the same report and must not look alike.
// `blocked` is only trusted to REFINE `answered` — on an unrecorded row the
// outcome stays unrecorded, because a build that never wrote the outcome is not
// a build whose other fields get to stand in for it.
const state: DnsRowState =
outcome === 'answered' ? (e.blocked ? 'blocked' : 'answered') : outcome
if (outcome === 'failed') {
const err = (e.error ?? '').trim()
return {
state,
outcome,
label: 'failed',
title:
'No usable answer was produced: a timeout or network error, a resolver refusal, a loopback, or a cached rejection. The path beside it says which way the lookup went before it failed.',
path,
pathLabel: dnsPathLabel(path),
pathTitle,
// A failure whose cause was not recorded is still a failure, and saying so
// is the whole point — an empty cell here would read as "nothing wrong".
cause: err ? `cause ${err}` : 'cause not recorded',
// -1 is the daemon's "no response at all"; anything else is a code the
// server really sent back (2 SERVFAIL, 5 REFUSED).
rcodeNote: e.rcode === -1 ? 'no response' : `rcode ${e.rcode}`,
}
}
return {
state,
outcome,
label: outcome === 'answered' ? 'answered' : 'not recorded',
title:
outcome === 'answered'
? 'An answer was produced — by the resolver, from the cache, or synthesized by the filter when it blocked. It does not claim the answer was useful: an upstream NXDOMAIN is answered too.'
: 'What came of this lookup was not recorded — an older row, or a producer this build does not recognise. It is NOT a claim that the lookup succeeded.',
path,
pathLabel: dnsPathLabel(path),
pathTitle,
cause: '',
rcodeNote: '',
}
}
/** The path chip's text. `unknown` is named, not blanked: a path nobody recorded
* and a path this build cannot read are both "we do not know", and an empty chip
* would read as "direct". */
function dnsPathLabel(k: DnsPathKind): string {
return k === 'unknown' ? 'unknown' : k
}
// ---- what the search box covers --------------------------------------------
/**
* The fields the daemon searches in a QUERY-log row — the mirror of
* stats/filter.go matchLog, and the same POSITIVE, CLOSED list.
*
* It exists so the panel can say what it searches (and the `?mock` backend can
* behave like the real one) without either of them drifting into searching
* something the daemon doesn't. `error` is in because a failure cause is how the
* failures are found at all — `q=timeout` is the search somebody actually runs.
* Deliberately absent: seq/unix/time/rcode (numbers, where a substring is noise),
* `blocked` (a boolean), and the two vocabulary words `outbound_kind` and
* `status` — `q=default` would match every default-egress row while the operator
* was looking for a tag by that name, and `q=failed` would match every failure
* while they were looking for a domain with "failed" in it.
*/
export function logSearchFields(e: QueryLogEntry): string[] {
return [e.domain, e.qtype, e.server, e.action, e.device, e.outbound, e.error ?? '']
}
/**
* The fields the daemon searches in a CONNECTION-log row — the mirror of
* stats/filter.go matchConn. `rule` is in because the engine's match condition
* is how a domain is found through the rule that routed it. Deliberately absent:
* seq/unix, `port` (a bare "443" would also match any address containing 443,
* reading as a port filter while being something else) and `rule_kind`.
*/
export function connSearchFields(e: ConnLogEntry): string[] {
const out = [e.src_ip, e.src_name, e.dest, e.dest_ip, e.network, e.proto, e.outbound, e.rule]
return e.chain ? out.concat(e.chain) : out
}
/** Human-readable naming of {@link logSearchFields}, for the search box hint. */
export const LOG_SEARCH_HINT = 'domain · type · resolver · action · device · exit · error'
/**
* Human-readable naming of {@link connSearchFields}.
*
* It must name EVERY field in that list. A hint that stops short is a promise the
* search over-keeps: `proto` and the chain hops WERE searched (stats/filter.go
* matchConn walks e.Proto and every e.Chain element) while this line said they
* were not, so a hit on either read as the filter misbehaving.
*/
export const CONN_SEARCH_HINT =
'device · destination · network · proto · exit · rule · path'
/**
* ASCII-only lowercase — A-Z and nothing else, the exact fold stats/filter.go
* performs (lowerASCII on the needle, matchAtFold on the haystack byte).
*
* `String.prototype.toLowerCase` is Unicode-aware and therefore WRONG here: it
* case-folds non-ASCII letters too, which the daemon does not do — pinned on that
* side by the non-ASCII case of stats/logfilter_test.go TestContainsFold, where an
* upper-case non-ASCII needle is asserted NOT to match its lower-case haystack.
* The panel folding wider than the daemon is not a harmless nicety: the `?mock`
* backend exists so a search behaves in mock mode exactly as it does on hardware,
* and a needle that finds rows in the panel and none on the router is the one
* outcome that makes the filter untrustworthy.
*/
function lowerAscii(s: string): string {
let out = ''
for (let i = 0; i < s.length; i++) {
const c = s.charCodeAt(i)
out += c >= 65 && c <= 90 ? String.fromCharCode(c + 32) : s[i]
}
return out
}
/**
* Case-insensitive substring match over a row's searched fields — the client-side
* twin of stats/filter.go containsFold, used by the `?mock` backend so a search
* in mock mode finds exactly what a search on hardware finds.
*
* ASCII folding only, on BOTH sides, exactly like the daemon: non-ASCII compares
* code unit for code unit, so no Unicode case pair matches. See {@link lowerAscii}.
*/
export function rowMatches(fields: readonly string[], needle: string): boolean {
const q = lowerAscii(needle.trim())
if (!q) return true
return fields.some((f) => lowerAscii(f ?? '').includes(q))
}
// ---- a short page is not always the end ------------------------------------
/**
* Is the history behind this page exhausted?
*
* A short page normally means the log ran out — but NOT when the daemon reported
* that its scan budget ended the walk. Reading a truncated page as exhausted is
* how the panel would hide the rest of the log behind a "nothing found": the
* paging control disappears and the operator is told the search is complete when
* it stopped early.
*/
export function historyExhausted(p: {
returned: number
page: number
truncated: boolean
}): boolean {
if (p.truncated) return false
return p.returned < p.page
}
/**
* Where to resume paging from. The daemon's scan cursor is authoritative — on a
* complete page it already equals the last row's seq, and on a truncated page it
* is the ONLY thing that can carry the search on, because a page with no matches
* has no row seq to page from. 0/absent ⇒ fall back to the oldest row we hold;
* null ⇒ there is nowhere to resume.
*/
export function resumeCursor(p: { cursor: number; rows: readonly { seq: number }[] }): number | null {
if (p.cursor > 0) return p.cursor
const last = p.rows[p.rows.length - 1]
return last ? last.seq : null
}
/**
* The sentence under a log that has nothing (more) to show.
*
* `complete` is the load-bearing part: it says whether the panel is claiming to
* have reached the end of the data. It is false whenever the scan stopped early,
* and the text says so in those words — "stopped" and "not the end", never
* "nothing found".
*/
export interface LogEndNote {
/** True only when the panel can honestly claim the search reached the end. */
complete: boolean
text: string
}
export function logEndNote(p: {
filter: string
truncated: boolean
matches: number
}): LogEndNote {
const q = p.filter.trim()
if (p.truncated) {
return {
complete: false,
text:
p.matches === 0
? `The search stopped at the daemon's scan limit before it matched anything. This is not the end of the log — there may be matches further back.`
: `The search stopped at the daemon's scan limit. These are the matches found so far, not all of them.`,
}
}
if (q) {
return {
complete: true,
text:
p.matches === 0
? `Nothing in the log matches “${q}”.`
: `All ${p.matches === 1 ? 'match' : `${p.matches} matches`} for “${q}” are loaded.`,
}
}
return {
complete: true,
text: p.matches === 0 ? 'Nothing logged yet.' : 'The whole log is loaded.',
}
}
/**
* The count read-out under a log. It exists as a function because the string it
* must NEVER produce is "all loaded" on a truncated scan — the panel would be
* signing for data it never looked at.
*/
export function logCountLabel(p: {
rows: number
showMore: boolean
truncated: boolean
filtered: boolean
}): string {
const n = p.rows.toLocaleString('en-US')
const noun = p.filtered ? (p.rows === 1 ? 'match' : 'matches') : 'loaded'
if (p.truncated) return `${n} ${p.filtered ? noun : 'rows'} · scan incomplete`
if (p.showMore) return p.filtered ? `${n} ${noun}` : `${n} loaded`
return p.filtered ? `${n} ${noun} · all loaded` : `${n} · all loaded`
}
+25 -9
View File
@@ -2,17 +2,33 @@ import { StrictMode } from 'react'
import { createRoot } from 'react-dom/client'
import './tokens.css'
import { App } from './App'
import { initMockBackend } from './api'
import { ConfirmProvider } from './components'
const rootEl = document.getElementById('root')
if (!rootEl) throw new Error('#root not found')
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
// Settle the fixture question BEFORE the first render: pages read `MOCK` while
// they render, so a backend that arrives afterwards would paint half a screen
// from the daemon and half from fixtures. In a production build this resolves
// immediately and to `false` — the fixtures are not in the bundle to load (see
// api.ts initMockBackend and the assertNoMockFixtures plugin in vite.config.ts).
function mount() {
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl!).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
}
// A fixture module that fails to load is a broken dev checkout, not a reason to
// hand the operator a blank plate — mount anyway and let the shell report that it
// cannot reach a daemon, which by then is the truth.
void initMockBackend().then(mount, (e) => {
console.error('mock backend failed to load; continuing against the real API', e)
mount()
})
+651 -59
View File
@@ -6,7 +6,39 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import type { ApplyResult, ChainHealth, ChainHopHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, Profile, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning, Traffic } from './api'
import { killSwitchClosed } from './planeState'
// The SAME field lists and matcher the page documents and the daemon implements —
// so a search in `?mock` cannot quietly be more (or less) generous than the real one.
import { connSearchFields, logSearchFields, rowMatches } from './logRoute'
import type { ApplyResult, ChainHealth, ChainHopHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, Profile, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning, TestKind, Traffic } from './api'
/** One URL knob, safe to read before `location` exists (SSR-less builds/tests). */
function mockParam(name: string): string | null {
if (typeof location === 'undefined') return null
return new URLSearchParams(location.search).get(name)
}
/**
* `?subleak=draft|applied|divergent` — which half of the subscription
* fetch-route reconciliation to show (see subFetch.ts). Anything else, or
* absent, leaves the fixture's subscription fetching directly, which is the
* state that must show NO badge at all.
*/
type SubLeakMode = '' | 'draft' | 'applied' | 'divergent' | 'paused' | 'pausedapplied' | 'nourl'
const SUB_LEAK_MODES: readonly SubLeakMode[] = [
'draft',
'applied',
'divergent',
'paused',
'pausedapplied',
'nourl',
]
function mockSubLeak(): SubLeakMode {
const want = (mockParam('subleak') ?? '').trim()
return (SUB_LEAK_MODES as readonly string[]).includes(want) ? (want as SubLeakMode) : ''
}
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
@@ -33,7 +65,19 @@ const CONFIG: Model = {
ActiveProfile: 'mobile-uplink',
// Policy for traffic TPROXY physically can't carry (non-TCP/UDP). Override
// from the URL — ?mock&untun=icmp / &untun=direct — to see all three states.
Untunnelable: new URLSearchParams(typeof location === 'undefined' ? '' : location.search).get('untun') ?? 'block',
Untunnelable: mockParam('untun') ?? 'block',
// The two settings that decide part of the untunnelable traffic BEFORE the
// policy above is consulted, so the Networks copy has to change shape for
// them: ?mock&l3=1 (ping rides the tunnel) and ?mock&uegress=wg0 (the kernel
// routes ESP/GRE/SCTP out that interface). Both off by default, as shipped.
L3Tunnel: mockParam('l3') === '1',
UntunnelableEgress: mockParam('uegress') ?? '',
// Network-wide DNS filter, OFF by default here. That is the interesting
// state: the devices below attach lists, which keep running for them with
// this switch off, so the master switch has to say so instead of calling the
// lists "configured but inactive". ?mock&dnsfilter=1 turns it on to see the
// other sentence.
DNSFilter: mockParam('dnsfilter') === '1',
DNSIntercept: true, // force ALL LAN plaintext DNS (:53) through the engine
BlockDoH: false, // block known public DoH resolvers so clients fall back to plaintext :53
GroupHealth: true, // observatory: background probing of used groups/chains + Targets health stats (default on)
@@ -42,7 +86,7 @@ const CONFIG: Model = {
StatsMaxDomains: 5000, // fixed cap — shows the "limit" rendering (5000)
StatsRetentionDisabled: false,
StatsBackend: 'memory', // logging backend: off | memory | sqlite
StatsDiskLimitMB: 64, // SQLite-only disk cap (MB); shows once backend=sqlite (0 ⇒ Unlimited)
StatsDiskLimitMB: 64, // disk-backend-only cap (MB); shows once backend=sqlite (0 ⇒ Unlimited)
// Daemon operational log (shaterd's own log): both destinations on, file in
// tmpfs (the default), 2 MB cap — the defaults a fresh install ships with.
LogToSyslog: true,
@@ -76,10 +120,32 @@ const CONFIG: Model = {
// of a 200 GB plan used, expiring in ~24 days. Drives the quota bar + countdown.
{
Name: 'primary',
Enabled: true,
URL: 'https://sub.example.net/link',
// SWITCHED OFF in the `paused` mode, and that mode exists because the
// switch does not do what the badge used to claim: apply.go fetches a
// subscription BY NAME without reading Enabled, and the row's own
// "Fetch now" is not gated on it either. Off + proxy + no route was drawn
// quiet, with "nothing is disclosed yet" — over a button that discloses.
Enabled: mockSubLeak() !== 'paused' && mockSubLeak() !== 'pausedapplied',
// `nourl` is the other half of the same split, and the only genuinely quiet
// one: subscribe/fetch.go refuses an empty URL before it builds a request.
URL: mockSubLeak() === 'nourl' ? '' : 'https://sub.example.net/link',
Format: 'auto',
UpdateInterval: '12h',
// The fetch route, driven from the URL so the ONE badge can be seen in each
// of the states it has to reconcile (see subFetch.ts and mockSubLeak):
// ?subleak=draft — proxy, no route: the panel PREDICTS the disclosure
// ?subleak=applied — the same config, and the daemon has now REPORTED it
// ?subleak=divergent — a route IS named, and the daemon reports the leak
// anyway. The dangerous direction: the panel's own
// predicate is content and must not be believed over
// the applied verdict.
// ?subleak=paused — switched off, proxy, no route: amber, because the
// switch stops the SCHEDULED refresh and nothing else
// ?subleak=pausedapplied — the same, and the daemon has REPORTED it, with
// its own disabled-subscription sentence
// ?subleak=nourl — no URL: the one state that really is quiet
FetchVia: mockSubLeak() ? 'proxy' : undefined,
FetchDetour: mockSubLeak() === 'divergent' ? 'group:auto' : undefined,
UserUpload: 8_142_336_512,
UserDownload: 60_293_117_952,
UserTotal: 214_748_364_800, // 200 GiB
@@ -155,8 +221,6 @@ const CONFIG: Model = {
// Goes straight out the default WAN, but splits the TLS ClientHello on the way
// — a direct egress exists to carry a native DPI preset.
{ Name: 'tunnel', Type: 'direct', DPI: 'fragment' },
// The local ciadpi desync proxy (the only type that dials a port).
{ Name: 'ciadpi', Type: 'byedpi', Port: 1080 },
],
Rules: [
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
@@ -196,6 +260,19 @@ const CONFIG: Model = {
{ Name: 'StevenBlack', Enabled: true, Source: 'url', URL: 'https://raw.githubusercontent.com/StevenBlack/hosts/master/hosts', Response: 'nxdomain', UpdateInterval: '24h' },
{ Name: 'oisd-basic', Enabled: true, Source: 'url', URL: 'https://big.oisd.nl/domainswild', Response: 'nxdomain', UpdateInterval: '24h' },
{ Name: 'telegram-block', Enabled: false, Source: 'geosite', Categories: ['telegram'], Response: 'nxdomain', UpdateInterval: '24h' },
// OFF for the network and ATTACHED to a device — the state the panel used to
// draw as dead. It is inline, so its content is the config and it loads at
// apply; the status record below reports it so the chip can be honestly green.
{ Name: 'family-extra', Enabled: false, Source: 'inline', Entries: ['roblox.com', 'discord.com', 'twitch.tv'], Response: 'zero' },
],
// Allow lists a device can attach. Three readings, all reachable offline:
// school-allow is loaded, work-allow has never been fetched (crit), and
// holidays-allow has NO status record at all — "load unknown", which must not
// be drawn as healthy.
Allowlists: [
{ Name: 'school-allow', Enabled: true, Source: 'inline', Entries: ['school.example.edu', 'classroom.google.com'] },
{ Name: 'work-allow', Enabled: false, Source: 'url', URL: 'https://lists.example.net/work-allow.txt?key=secret', UpdateInterval: '6h' },
{ Name: 'holidays-allow', Enabled: true, Source: 'geosite', Categories: ['github'], UpdateInterval: '24h' },
],
Resolvers: [
{ Name: 'cloudflare-doh', Type: 'doh', Address: 'https://1.1.1.1/dns-query', Detour: 'tunnel' },
@@ -216,22 +293,37 @@ const CONFIG: Model = {
// getDevices() below — without these the Devices page showed every client as
// unmanaged while discovery claimed configured:true, which is not a state the
// real backend can produce (it derives `configured` from exactly this list).
//
// Between them the three rows reach every chip state the picker can draw:
// loaded (oisd-basic), not loaded (StevenBlack — never fetched), load unknown
// (holidays-allow — no status record), no such list (`old-adblock`, a name the
// config no longer has), and a list that is OFF for the network but attached
// here and working (family-extra).
Devices: [
{ Name: 'Max laptop', MAC: 'a4:83:e7:11:22:33', Enabled: true },
{
Name: 'Max laptop',
MAC: 'a4:83:e7:11:22:33',
Enabled: true,
Blocklists: ['oisd-basic'],
},
{
Name: "Lena's phone",
MAC: 'f0:18:98:aa:bb:cc',
Enabled: true,
Block: ['tiktok.com', 'ads.doubleclick.net', 'telemetry.example'],
Block: ['tiktok.com', 'ads.doubleclick.net', 'keyword:telemetry'],
Blocklists: ['StevenBlack'],
Allowlists: ['school-allow'],
},
// Paused policy + an allow that overrides a network blocklist — exercises the
// "paused" badge and the allowlist chips in one row.
// "paused" badge, both chip lanes, and the two failure readings in one row.
{
Name: 'Kids iPad',
MAC: '3c:22:fb:44:55:66',
Enabled: false,
Block: ['youtube.com', 'roblox.com', 'discord.com'],
Block: ['youtube.com', 'roblox.com', 'full:discord.com'],
Allow: ['school.example.edu'],
Blocklists: ['family-extra', 'old-adblock'],
Allowlists: ['holidays-allow', 'work-allow'],
},
],
Alerts: [
@@ -299,7 +391,13 @@ const RULESET_STATUS: RulesetStatus[] = [
rule_count: 14_203,
},
{ tag: 'rs-ad-hosts', name: 'ad-hosts', category: '', kind: 'ruleset', remote: true, last_updated: '', interval_seconds: 43_200, rule_count: 0 },
{ tag: 'rs-private-nets', name: 'private-nets', category: '', kind: 'ruleset', remote: false, last_updated: '', interval_seconds: 0, rule_count: 3 },
// LOCAL (inline) — and its rule_count is 0 because the daemon has no other
// answer to give: engine.go fills RuleSetStat.RuleCount from
// (*rule.RemoteRuleSet).RuleCount(), and there is no LocalRuleSet.RuleCount in
// the tree at all. This fixture used to say 3, a number no router can produce,
// which is how the panel's "empty — nothing matches" verdict over a working
// inline list went unnoticed: the test data disagreed with the daemon.
{ tag: 'rs-private-nets', name: 'private-nets', category: '', kind: 'ruleset', remote: false, last_updated: '', interval_seconds: 0, rule_count: 0 },
// Live geosite/geoip rule-sets — remote, so they show freshness + Update-now.
// yt-geosite has TWO category entries sharing name 'yt-geosite': the row groups
// them, shows the OLDEST last_updated (5h, from google), and sums the counts.
@@ -342,6 +440,19 @@ const RULESET_STATUS: RulesetStatus[] = [
// Disabled in CONFIG, so the row reads "off" whatever this says — it exists to
// prove the row does not start claiming things the moment a status appears.
{ tag: 'bl-telegram-block-telegram', name: 'telegram-block', category: 'telegram', kind: 'blocklist', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
// Inline: nothing to fetch, so remote:false, no timestamp — and NO COUNT, which
// is the whole point. The daemon publishes an entry count only for the sets it
// fetches, so an inline list that holds three entries and one that holds none
// are the same bytes on the wire. The device chip therefore reads "size
// unknown", not "empty": see deviceLists.listLoad.
{ tag: 'bl-family-extra', name: 'family-extra', category: '', kind: 'blocklist', remote: false, last_updated: '', interval_seconds: 0, rule_count: 0 },
// Allowlists report the same way under `al-<name>`. school-allow is INLINE, so
// it is present in the engine with no count to publish ("size unknown");
// work-allow is remote and has NEVER been fetched, so a device that attached it
// is not getting the exception it thinks it is. holidays-allow is deliberately
// ABSENT from this array — that is the third reading, "load unknown".
{ tag: 'al-school-allow', name: 'school-allow', category: '', kind: 'allowlist', remote: false, last_updated: '', interval_seconds: 0, rule_count: 0 },
{ tag: 'al-work-allow', name: 'work-allow', category: '', kind: 'allowlist', remote: true, last_updated: '', interval_seconds: 21_600, rule_count: 0 },
]
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
@@ -444,7 +555,11 @@ export async function updateRuleset(tag: string): Promise<RulesetStatus> {
const entry = RULESET_STATUS.find((r) => r.tag === tag)
if (!entry) throw new Error(`unknown ruleset ${tag}`)
entry.last_updated = new Date().toISOString()
entry.rule_count = entry.rule_count > 0 ? entry.rule_count + 7 : 15_734
// Only a REMOTE set gains a count: the daemon reads it off RemoteRuleSet and a
// local set has none to read. Fabricating one here would put a number on screen
// that no router can produce — the same fixture lie that hid the inline-list
// "empty" verdict.
if (entry.remote) entry.rule_count = entry.rule_count > 0 ? entry.rule_count + 7 : 15_734
return { ...entry }
}
@@ -500,14 +615,28 @@ export async function getRulesetCategories(source: string): Promise<RulesetCateg
// the field case the readout used to call "Protected" (one
// rule, `default → direct`); `unknown` is a daemon too old to
// report. Default: tunnel.
function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSwitch: string } {
// ?mock&plane=unreported → a daemon that sends NO `plane` field. The panel then
// knows nothing about what is installed, which is the state
// the Kill-switch module used to render as a green "ARMED"
// (`undefined !== 'none'` is true).
function mockPlane(): {
plane: 'full' | 'hold' | 'none' | undefined
engine: boolean
killSwitch: string
} {
const q = typeof location === 'undefined' ? '' : location.search
const params = new URLSearchParams(q)
const killSwitch = params.get('ks') === 'open' ? 'open' : 'closed'
// Passed through VERBATIM, because that is what the daemon does: apply.go sets
// `s.KillSwitch = m.Globals.KillSwitch` with no normalisation, so `?ks=Closed`,
// `?ks=%20closed%20` and `?ks=` are all reachable readings of a router that
// BLOCKS. The mock used to fold everything that wasn't "open" to "closed",
// which made the panel's own `=== 'closed'` bug unreproducible here.
const killSwitch = params.get('ks') ?? 'closed'
const p = params.get('plane')
if (p === 'hold') return { plane: 'hold', engine: false, killSwitch: 'closed' }
if (p === 'none') return { plane: 'none', engine: false, killSwitch }
if (p === 'open') return { plane: 'none', engine: false, killSwitch: 'open' }
if (p === 'unreported') return { plane: undefined, engine: true, killSwitch }
return { plane: 'full', engine: true, killSwitch }
}
@@ -515,7 +644,7 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
// meaningful with the plane installed: with the engine down there is no running
// config to judge, and the daemon reports the unknown/zero value — so do the same
// here rather than leaving a stale "tunnel" behind a dead engine.
function mockTraffic(plane: 'full' | 'hold' | 'none'): Traffic | undefined {
function mockTraffic(plane: 'full' | 'hold' | 'none' | undefined): Traffic | undefined {
if (plane !== 'full') return { verdict: '', default: '', tunnel_rules: 0 }
const params = new URLSearchParams(typeof location === 'undefined' ? '' : location.search)
switch (params.get('traffic')) {
@@ -560,6 +689,29 @@ const MOCK_WARNINGS: StatusWarning[] = [
name: 'fakeip-pool',
message: 'fake-IP resolver cannot be used as a fallback; the failover chain was not built',
},
// Two findings the generator attributes to a NODE by name — the class that the
// Nodes page never showed, leaving a node the engine threw away rendered as an
// ordinary row with a green toggle. Both name real fixture nodes so the row
// badge, the collapsed-bucket "N flagged" count and the per-row strip all fire.
{
severity: 'warning',
section: 'node',
name: 'fi-trojan',
message: 'parse share-link: unsupported scheme "trojan+ws" (skipped)',
},
{
severity: 'warning',
section: 'node',
name: 'home-wg',
message:
'this WireGuard node is materialised twice in the engine config — as "home-wg" and as "group-stealth-m1-home-wg" — and traffic can reach both. A WireGuard peer keeps ONE session per public key, so two devices built from one private key evict each other continuously and NEITHER tunnel passes traffic. Only "home-wg" is kept; everything that routed through "group-stealth-m1-home-wg" is fail-closed (blocked) instead of leaving over the plain WAN',
},
{
severity: 'warning',
section: 'subscription',
name: 'backup',
message: 'fetch failed: dial tcp 203.0.113.9:443: i/o timeout — serving the nodes cached earlier',
},
{
severity: 'info',
section: 'generate',
@@ -568,6 +720,34 @@ const MOCK_WARNINGS: StatusWarning[] = [
},
]
/**
* The daemon's truncation disclosure, exactly as apply/warnings.go writes it when
* the published set overflows the 50-entry cap. Served under `?mock&trunc` so the
* "this list is incomplete" rendering is exercisable — it used to be dropped
* wholesale by the panel's `info` filter and reached no screen at all.
*/
/**
* The daemon's own critical finding when it cannot read the configuration
* (apply.go, section "config" / name "unreadable"). Copied close to verbatim: the
* sentence about NOT switching anything off is the load-bearing one — the instinct
* in front of a dead LAN is to turn things off, and that is the single action that
* makes this worse.
*/
const CONFIG_UNREADABLE_WARNING: StatusWarning = {
severity: 'critical',
section: 'config',
name: 'unreadable',
message:
"the router's configuration could NOT be read (uci show shater: exit status 1), so this status cannot say whether shater is switched on, whether the kill switch is closed, or which port this panel is served on — enabled, kill_switch and panel_port are placeholders here, not readings. If traffic is being blocked, that is the fail-closed plane doing its job and NOT the service being switched off: do not turn anything off to fix it. The usual causes are a full /overlay and a `uci commit` interrupted part-way; free space, check /etc/config/shater, then restart shaterd.",
}
const MOCK_TRUNCATION: StatusWarning = {
severity: 'info',
section: 'generate',
name: '',
message: '7 further warning(s) suppressed; run `logread -e shater` for the full list',
}
/**
* The standing `untunnelable` note the daemon reports. It is INFO, never a
* problem: it states a correct, chosen configuration. Two shapes, mirroring the
@@ -575,7 +755,36 @@ const MOCK_WARNINGS: StatusWarning[] = [
* is inert entirely while the kill-switch is open.
*/
function untunnelableNote(mode: string, killSwitch: string): StatusWarning[] {
if (killSwitch === 'open') {
const g = CONFIG.Globals as { L3Tunnel?: boolean; UntunnelableEgress?: string }
const egress = (g.UntunnelableEgress ?? '').trim()
// The daemon's own precedence: the egress carrier owns the whole story, then
// the L3 ingress, then the kill-switch, then the policy (apply/warnings.go).
if (egress) {
return [
{
severity: 'info',
section: 'untunnelable',
name: egress,
message:
(g.L3Tunnel
? 'ping and Windows tracert travel THROUGH the tunnel; everything else the tunnel cannot carry — IPsec (ESP/AH), PPTP/GRE, SCTP — now leaves'
: 'ping, Windows tracert, IPsec (ESP/AH), PPTP/GRE, SCTP and every other protocol that is neither TCP nor UDP now leave') +
` through egress "${egress}": the kernel routes them out that interface with that interface's own NAT, and none of it follows your routing rules. Multicast IPTV does not pass this router under any setting, and carrying IGMP out an egress cannot change that.`,
},
]
}
if (g.L3Tunnel) {
return [
{
severity: 'info',
section: 'untunnelable',
name: '',
message:
'ping and Windows tracert work and travel THROUGH the tunnel, toward every address your rules send to an outbound that can carry plain IP (WireGuard/AmneziaWG); addresses your rules send anywhere else cannot be pinged at all, deliberately. Raw VPN passthrough (IPsec ESP/AH, PPTP/GRE) cannot enter the tunnel and stays with the untunnelable policy. Multicast IPTV does not pass this router on any setting; the L3 ingress does not change that.',
},
]
}
if (!killSwitchClosed(killSwitch)) {
return [
{
severity: 'info',
@@ -609,14 +818,57 @@ function untunnelableNote(mode: string, killSwitch: string): StatusWarning[] {
return []
}
/**
* The daemon's own critical finding for a proxy fetch that travels through
* nothing (apply/warnings.go subFetchDirectMessage), abridged to its first
* sentences. It is published only for `?subleak=applied|divergent` — i.e. only
* when the LAST APPLY saw the state — because that is exactly what separates it
* from the panel's prediction about a draft.
*/
const SUB_FETCH_LEAK: StatusWarning = {
severity: 'critical',
section: 'subscription',
name: 'primary',
message:
'`fetch_via` is "proxy" and no detour is set, which reads as configured and is not: an empty detour resolves to the `direct` outbound. So this subscription\'s feed is fetched over your ordinary WAN: the provider that serves it sees your router\'s real IP address, and everyone on the path there — your ISP included — sees that this router talks to them. The fetch itself gives no sign of it: it SUCCEEDS, the node list updates, and it repeats on every scheduled refresh.',
}
/**
* The SAME finding for a subscription that is switched OFF — the daemon now
* files it (apply/warnings.go stopped skipping disabled subscriptions, because
* the premise that they are never fetched was false) and varies its own WHEN.
*
* It exists here so `?subleak=paused` can be applied as well as drafted: an
* `enabled=0` row wearing the enabled sentence would send the reader looking for
* a scheduled refresh that is not running, which is the same lie inverted.
*/
const SUB_FETCH_LEAK_OFF: StatusWarning = {
...SUB_FETCH_LEAK,
message:
'`fetch_via` is "proxy" and no detour is set, which reads as configured and is not: an empty detour resolves to the `direct` outbound. So this subscription\'s feed is fetched over your ordinary WAN: the provider that serves it sees your router\'s real IP address, and everyone on the path there — your ISP included — sees that this router talks to them. The fetch itself gives no sign of it: it SUCCEEDS and the node list updates. This subscription is switched OFF, so no scheduled refresh touches it — but `enabled=0` does NOT block a fetch you start yourself: the Fetch-now button on this row and `shaterd sub update <name>` both go through, over the plain WAN, exactly as described here.',
}
function mockWarnings(killSwitch: string): StatusWarning[] {
const q = typeof location === 'undefined' ? '' : location.search
const params = new URLSearchParams(q)
const mode = (CONFIG.Globals as { Untunnelable?: string }).Untunnelable ?? 'block'
const notes = untunnelableNote(mode, killSwitch)
const leak = mockSubLeak()
// Published unconditionally for those two knobs — a finding that only appears
// with `?warn` could not be paired with the config state it is about.
if (leak === 'applied' || leak === 'divergent' || leak === 'pausedapplied') {
// A switched-off subscription gets the daemon's OWN disabled sentence, not
// the enabled one — the two differ in WHEN, and the badge above them differs
// the same way.
const finding = leak === 'pausedapplied' ? SUB_FETCH_LEAK_OFF : SUB_FETCH_LEAK
return [{ ...finding }, ...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes]
}
// `?trunc` adds the daemon's "the published list is capped" disclosure, which
// it appends IN PLACE OF the last entry it had room for.
const trunc = params.has('trunc') ? [{ ...MOCK_TRUNCATION }] : []
// A degraded plane always comes with the findings that explain it.
if (params.has('warn') || params.get('plane')) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes]
if (params.has('warn') || params.get('plane') || trunc.length > 0) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes, ...trunc]
}
return notes
}
@@ -625,6 +877,33 @@ export async function getStatus(): Promise<Status> {
await wait(120)
const enabled = (CONFIG.Globals as { Enabled: boolean }).Enabled
const { plane, engine, killSwitch } = mockPlane()
// ?mock&cfg=unreadable — the daemon could not READ the configuration (full
// /overlay, or a `uci commit` caught half-written). It is not a hypothetical: it
// is the situation the fail-closed plane exists for, so it ships with the plane
// HOLDING and the whole LAN cut off deliberately — while `enabled`, `kill_switch`
// and `panel_port` are placeholders that mean nothing. Reproducing it here is how
// the "Turned off" misreading stays fixed: the panel must alarm, not reassure.
if (mockParam('cfg') === 'unreadable') {
return {
running: true,
enabled: false, // a placeholder, NOT "the owner switched it off"
active: false,
table: true,
hash,
version: '1.11.0-shater',
kill_switch: '', // placeholder likewise
panel_port: 0, // placeholder likewise
config_readable: false,
config_error: 'uci show shater: exit status 1',
can_rollback: armed || hasLastGood,
engine_running: false,
plane: 'hold',
traffic: mockTraffic('hold'),
warnings: [CONFIG_UNREADABLE_WARNING, ...mockWarnings(killSwitch)],
started_unix: MOCK_STARTED_UNIX,
uptime_seconds: Math.floor(Date.now() / 1000) - MOCK_STARTED_UNIX,
}
}
return {
running: true,
enabled,
@@ -634,6 +913,9 @@ export async function getStatus(): Promise<Status> {
version: '1.11.0-shater',
kill_switch: killSwitch,
panel_port: 8088,
// The daemon read the config fine in every other mock state. Sent explicitly
// rather than left off: absent means "no reading", which is a different claim.
config_readable: true,
can_rollback: armed || hasLastGood,
engine_running: engine,
plane,
@@ -643,11 +925,6 @@ export async function getStatus(): Promise<Status> {
// the reading ticks forward across polls exactly like the real daemon's does.
started_unix: MOCK_STARTED_UNIX,
uptime_seconds: Math.floor(Date.now() / 1000) - MOCK_STARTED_UNIX,
// ?nobyedpi flips the ciadpi-missing state so the gated Targets editor is
// exercisable in mock mode.
byedpi_installed: !new URLSearchParams(
typeof location === 'undefined' ? '' : location.search,
).has('nobyedpi'),
}
}
@@ -726,14 +1003,47 @@ export async function rollback(): Promise<ApplyResult> {
}
// A small rotating fixture query log so `?mock` renders a live-looking stream.
const MOCK_DOMAINS: Array<{ domain: string; action: string; server: string; device: string }> = [
{ domain: 'graph.facebook.com', action: 'proxy', server: 'cloudflare-doh', device: '192.168.1.77' },
{ domain: 'blocked-ad.example', action: 'block', server: 'router-local', device: 'anna-laptop' },
{ domain: 'www.gstatic.com', action: 'pass', server: 'router-local', device: '192.168.1.42' },
{ domain: 'telemetry.example', action: 'block', server: 'cloudflare-doh', device: 'iphone-anna' },
{ domain: 'github.com', action: 'proxy', server: 'cloudflare-doh', device: '192.168.1.77' },
// router's own resolutions (urltest probe / sub fetch) carry the literal "router"
{ domain: 'ntp.openwrt.org', action: 'pass', server: 'router-local', device: 'router' },
//
// TWO closed vocabularies are covered here on purpose, because they are what the
// page has to draw apart and a fixture set that only ever showed one value would
// let a collapse ship unnoticed:
//
// `outbound_kind` — all four: detour / default / local / '' (not recorded).
// `status` — all three: answered / failed / '' (not recorded), and both
// shapes of failure. The marquee row is `failed` on the
// `proxy` path: the lookup left through the tunnel and died
// there, which is the row that used to render as the
// healthiest-looking line in the whole log. There is also a
// failure with NO cause text, because "failed, cause not
// recorded" is a state the daemon can really produce.
const MOCK_DOMAINS: Array<{
domain: string
action: string
server: string
device: string
outbound_kind: string
outbound: string
status: string
error?: string
/** Overrides the derived rcode. -1 = the server never answered at all. */
rcode?: number
}> = [
{ domain: 'graph.facebook.com', action: 'proxy', server: 'cloudflare-doh', device: '192.168.1.77', outbound_kind: 'detour', outbound: 'nl-reality-1', status: 'answered' },
{ domain: 'blocked-ad.example', action: 'block', server: 'router-local', device: 'anna-laptop', outbound_kind: 'local', outbound: '', status: 'answered' },
{ domain: 'www.gstatic.com', action: 'pass', server: 'router-local', device: '192.168.1.42', outbound_kind: 'default', outbound: '', status: 'answered' },
// The tunnel is what broke: action says HOW it went, status says it never came
// back. Drawn from `action` alone this row was an accent-coloured "proxy".
{ domain: 'api.telegram.org', action: 'proxy', server: 'cloudflare-doh', device: '192.168.1.77', outbound_kind: 'detour', outbound: 'de-hysteria', status: 'failed', error: 'dial udp 1.1.1.1:53: i/o timeout', rcode: -1 },
{ domain: 'telemetry.example', action: 'block', server: 'cloudflare-doh', device: 'iphone-anna', outbound_kind: 'local', outbound: '', status: 'answered' },
// The same failure on the DIRECT path, and the resolver DID answer — with a
// refusal. A different report from the one above, and it has to read as one.
{ domain: 'broken.example', action: 'pass', server: 'router-local', device: '192.168.1.42', outbound_kind: 'default', outbound: '', status: 'failed', error: 'rejected', rcode: 2 },
{ domain: 'github.com', action: 'proxy', server: 'cloudflare-doh', device: '192.168.1.77', outbound_kind: 'detour', outbound: 'de-hysteria', status: 'answered' },
// A failure the producer recorded WITHOUT a cause — the panel must say so.
{ domain: 'nocause.example', action: 'pass', server: 'router-local', device: 'anna-laptop', outbound_kind: 'local', outbound: '', status: 'failed', error: '', rcode: -1 },
// router's own resolutions (urltest probe / sub fetch) carry the literal "router".
// Left UNRECORDED on BOTH axes so the not-recorded state is on screen too.
{ domain: 'ntp.openwrt.org', action: 'pass', server: 'router-local', device: 'router', outbound_kind: '', outbound: '', status: '' },
]
// --- seeded, seq-ordered fixture stores ------------------------------------
@@ -750,12 +1060,20 @@ function makeLogRow(seq: number, unix: number): QueryLogEntry {
unix,
domain: src.domain,
qtype: seq % 3 === 0 ? 'AAAA' : 'A',
rcode: src.action === 'block' ? 3 : 0,
rcode: src.rcode ?? (src.action === 'block' ? 3 : 0),
// A failure is never a filter block — the two are different facts and the
// daemon never sets both (stats.go: Blocked is the filter's own verdict).
blocked: src.action === 'block',
server: src.server,
action: src.action,
device: src.device,
seq,
status: src.status,
// Absent, not empty, when there is nothing to say — the wire field is
// `omitempty`, so a row with no cause must not carry an empty string either.
...(src.error ? { error: src.error } : {}),
outbound_kind: src.outbound_kind,
outbound: src.outbound,
}
}
@@ -790,32 +1108,69 @@ function mockLog(n: number): QueryLogEntry[] {
* - neither ⇒ the newest `limit` rows (the head).
*
* Only `after=` reports a backlog; the other two always report 0 / false.
*
* # The text filter and its scan budget
*
* `q=` is matched over the SAME closed field list the daemon searches
* (logRoute.logSearchFields / connSearchFields, mirroring stats/filter.go), so a
* search in `?mock` finds and misses exactly what a search on hardware does —
* including the deliberate misses: no ports, no `rule_kind`, no `outbound_kind`.
*
* A filtered walk is BUDGETED, as on the daemon, and reports `truncated` +
* `cursor` when the budget rather than the data ended it. The mock budget is
* tiny (MOCK_FILTER_SCAN) where the daemon's is 20000: the fixture store is a few
* hundred rows, and a state the fixtures cannot reach is a state nobody can look
* at before it ships.
*/
const MOCK_FILTER_SCAN = 60
function queryStorePage<T extends { seq: number }>(
store: T[],
q: StatsLogQuery,
grow: () => void,
): { rows: T[]; pending: number; more: boolean } {
fields?: (r: T) => string[],
): { rows: T[]; pending: number; more: boolean; truncated: boolean; cursor: number } {
const limit = Math.min(Math.max(1, q.limit ?? 50), 5000)
const clone = (rows: T[]) => rows.map((r) => ({ ...r }))
const needle = q.q?.trim() ?? ''
const filtered = needle !== '' && fields != null
const hit = (r: T) => !filtered || rowMatches(fields!(r), needle)
/** Walk `candidates` in selection order, spending the budget on rows EXAMINED. */
const walk = (candidates: T[]) => {
const out: T[] = []
let budget = MOCK_FILTER_SCAN
let cursor = 0
let stopped = false
for (const r of candidates) {
if (out.length >= limit) break
if (filtered) {
if (budget <= 0) {
stopped = true
break
}
budget--
}
cursor = r.seq
if (hit(r)) out.push(r)
}
return { out, cursor, stopped }
}
if (q.after != null) {
grow()
const after = q.after
// Ascending by seq, so we take the OLDEST rows above the cursor first.
const above = store.filter((r) => r.seq > after).sort((a, b) => a.seq - b.seq)
const chunk = above.slice(0, limit)
const pending = above.length - chunk.length
const { out, cursor, stopped } = walk(above)
const pending = above.filter((r) => r.seq > (out[out.length - 1]?.seq ?? after) && hit(r)).length
// Hand it back newest-first, exactly as the wire format does.
return { rows: clone(chunk).reverse(), pending, more: pending > 0 }
return { rows: clone(out).reverse(), pending, more: pending > 0, truncated: stopped, cursor }
}
if (q.before != null) {
const before = q.before
return { rows: clone(store.filter((r) => r.seq < before).slice(0, limit)), pending: 0, more: false }
}
return { rows: clone(store.slice(0, limit)), pending: 0, more: false }
const candidates = q.before != null ? store.filter((r) => r.seq < q.before!) : store
const { out, cursor, stopped } = walk(candidates)
return { rows: clone(out), pending: 0, more: false, truncated: stopped, cursor }
}
/** Rows-only view, for the callers that don't tail the stream. */
@@ -859,7 +1214,11 @@ export async function getStats(): Promise<Stats> {
engine_up: true,
updated_at: new Date().toISOString(),
backend,
totals: { queries: 1842, blocked: 317, allowed: 1525 },
// The three outcome counters are mutually exclusive and sum to `queries`.
// `failed` is non-zero on purpose: the fixture log carries failed rows, and an
// aggregate that said "0 failed" over a log full of failures would be exactly
// the disagreement between row and total that the status field exists to end.
totals: { queries: 1842, blocked: 317, failed: 46, allowed: 1479 },
top_domains: [
{ domain: 'graph.facebook.com', count: 210, blocked: 0 },
{ domain: 'blocked-ad.example', count: 143, blocked: 143 },
@@ -975,20 +1334,26 @@ export async function getStatsLog(q: StatsLogQuery = {}): Promise<QueryLogEntry[
export async function getStatsLogPage(q: StatsLogQuery = {}): Promise<StatsLogPage<QueryLogEntry>> {
await wait(60)
return queryStorePage(LOG_STORE, q, growLog)
return queryStorePage(LOG_STORE, q, growLog, logSearchFields)
}
// Connection events (feedback #2a "which device, where"): each row is a LAN
// client → destination flow the query log can't show. Mix of named + unnamed
// devices, tcp/tls + udp/quic, and a raw-IP/udp flow with no sniffed proto.
//
// `rule_kind` covers all THREE states — matched / default / '' (not recorded) —
// for the same reason MOCK_DOMAINS covers four: the two rule-less states must be
// visibly different, and only a fixture that carries both can show it.
const MOCK_CONNS: Array<Omit<ConnLogEntry, 'unix' | 'seq'>> = [
{ src_ip: '192.168.1.50', src_name: 'laptop', dest: 'graph.facebook.com', dest_ip: '157.240.1.35', port: 443, network: 'tcp', proto: 'tls', outbound: 'nl-reality-1' },
{ src_ip: '192.168.1.51', src_name: 'phone', dest: 'instagram.com', dest_ip: '157.240.1.174', port: 443, network: 'tcp', proto: 'tls', outbound: 'de-hysteria' },
{ src_ip: '192.168.1.50', src_name: 'laptop', dest: 'discord-media.example', dest_ip: '162.159.130.234', port: 443, network: 'udp', proto: 'quic', outbound: 'nl-reality-2' },
{ src_ip: '192.168.1.77', src_name: '192.168.1.77', dest: '5.9.100.200', dest_ip: '5.9.100.200', port: 51820, network: 'udp', proto: '', outbound: 'direct' },
{ src_ip: '192.168.1.51', src_name: 'phone', dest: 'gateway.icloud.com', dest_ip: '17.253.55.201', port: 443, network: 'tcp', proto: 'tls', outbound: 'direct' },
{ src_ip: '192.168.1.50', src_name: 'laptop', dest: 'github.com', dest_ip: '140.82.112.3', port: 443, network: 'tcp', proto: 'tls', outbound: 'nl-reality-1' },
{ src_ip: '192.168.1.42', src_name: 'ipad-kids', dest: 'blocked-ad.example', dest_ip: '203.0.113.77', port: 80, network: 'tcp', proto: 'http', outbound: 'block' },
{ src_ip: '192.168.1.50', src_name: 'laptop', dest: 'graph.facebook.com', dest_ip: '157.240.1.35', port: 443, network: 'tcp', proto: 'tls', outbound: 'nl-reality-1', rule_kind: 'matched', rule: 'protocol=tls rule_set=rs-social', chain: ['nl-reality-1', 'auto'] },
{ src_ip: '192.168.1.51', src_name: 'phone', dest: 'instagram.com', dest_ip: '157.240.1.174', port: 443, network: 'tcp', proto: 'tls', outbound: 'de-hysteria', rule_kind: 'matched', rule: 'domain_suffix=instagram.com', chain: ['de-hysteria', 'auto'] },
{ src_ip: '192.168.1.50', src_name: 'laptop', dest: 'discord-media.example', dest_ip: '162.159.130.234', port: 443, network: 'udp', proto: 'quic', outbound: 'nl-reality-2', rule_kind: 'matched', rule: 'network=udp rule_set=rs-voice' },
{ src_ip: '192.168.1.77', src_name: '192.168.1.77', dest: '5.9.100.200', dest_ip: '5.9.100.200', port: 51820, network: 'udp', proto: '', outbound: 'direct', rule_kind: 'default', rule: '' },
{ src_ip: '192.168.1.51', src_name: 'phone', dest: 'gateway.icloud.com', dest_ip: '17.253.55.201', port: 443, network: 'tcp', proto: 'tls', outbound: 'direct', rule_kind: 'default', rule: '' },
// A row from before rule capture: NOT RECORDED, which is a different fact from
// the two `default` rows above and has to look different on screen.
{ src_ip: '192.168.1.50', src_name: 'laptop', dest: 'github.com', dest_ip: '140.82.112.3', port: 443, network: 'tcp', proto: 'tls', outbound: 'nl-reality-1', rule_kind: '', rule: '' },
{ src_ip: '192.168.1.42', src_name: 'ipad-kids', dest: 'blocked-ad.example', dest_ip: '203.0.113.77', port: 80, network: 'tcp', proto: 'http', outbound: 'block', rule_kind: 'matched', rule: 'rule_set=rs-ads', chain: ['block'] },
]
function makeConnRow(seq: number, unix: number): ConnLogEntry {
@@ -1015,7 +1380,7 @@ export async function getStatsConns(q: StatsLogQuery = {}): Promise<ConnLogEntry
export async function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<ConnLogEntry>> {
await wait(60)
return queryStorePage(CONN_STORE, q, growConns)
return queryStorePage(CONN_STORE, q, growConns, connSearchFields)
}
// LAN discovery. `configured` / `name` / `blockCount` are DERIVED from CONFIG.Devices
@@ -1335,9 +1700,16 @@ class ApiErrorLike extends Error {
// The last three are absence of measurement, not a broken target, and the copy
// has to keep them apart. Results land one per GET poll, so the running/progress
// state is visible too.
// Every row carries `kind` and `source`, because the panel classifies on
// `source` now and a fixture that omitted it would exercise only the fallback
// path — the one that reads the error prose. Note which rows carry
// `source: ''`: every "no measurement exists" state, including the blocked
// chain, exactly as engine/grouptest.go files them.
const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_unix'>> = {
auto: {
selected: 'nl-reality-2',
kind: 'group',
source: 'observatory',
delay_ms: 42,
exit_ip: '185.12.34.56',
exit_country: 'NL',
@@ -1346,6 +1718,8 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
},
stealth: {
selected: 'nl-reality-1',
kind: 'group',
source: 'observatory',
delay_ms: 137,
exit_ip: '',
exit_country: '',
@@ -1361,6 +1735,18 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
// has a separate message for it.
'ewan-wg-subs': {
selected: '',
kind: 'chain',
// No end-to-end measurement of the chain exists — the walk stopped at hop 3
// — so the daemon files source:''. It is the one `source:''` that stays LOUD
// (testResult.blockedHop): a probe did run, at the hop, and failed.
source: '',
// THE SAME FACT AS THE SENTENCE BELOW, as a number. The panel reads this and
// nothing else: it used to keep the row red by matching a fragment of that
// sentence, which made a reworded daemon message a silent downgrade from red
// to grey. Note it sits beside `source:''` on purpose — nothing measured this
// chain's own exit, and stamping an instrument on it would be the lie the
// `source` field exists to prevent.
blocked_by: 3,
delay_ms: 0,
exit_ip: '',
exit_country: '',
@@ -1372,6 +1758,10 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
// this path and did not come back.
'via-tunnel': {
selected: '',
kind: 'group',
// A measurement was taken and it failed — the only shape in this fixture
// that is a health verdict rather than an absence of one.
source: 'observatory',
delay_ms: 0,
exit_ip: '',
exit_country: '',
@@ -1381,6 +1771,8 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
// Not a health verdict — nothing routes here, so no measurement of it exists.
fallback: {
selected: '',
kind: 'group',
source: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
@@ -1390,6 +1782,8 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
},
relay: {
selected: '',
kind: 'chain',
source: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
@@ -1400,6 +1794,8 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
// Routed, materialised, simply not reached yet. Untested is not dead.
'sub-fresh': {
selected: '',
kind: 'chain',
source: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
@@ -1419,13 +1815,119 @@ function shapeFor(group: string, i: number): GroupTestResult {
return { ...base, group, tested_unix: Math.floor(Date.now() / 1000) }
}
/** engine.groupTestCarryMax — the cap on rows an earlier run may leave behind. */
const CARRY_MAX = 64
/**
* engine.carryForward, in the fixture: the rows this run will NOT produce an
* answer for, newest first, capped.
*
* Identity is (name, KIND) and not name alone, exactly as the daemon has it —
* testing the node `nl-reality-1` must not drop a group of the same name, which
* is the whole reason `kind` is on the result. An empty kind on either side
* matches everything with that name (a name the box resolved to nothing could
* have been any of them).
*
* The fixture used to skip this entirely: every run set `results: []` and a
* comment claimed the daemon did the same. It no longer does, and a mock that
* lags the daemon here would hide the exact thing the panel now has to draw —
* a carried reading sitting beside a fresh one.
*/
function carryForward(
prev: GroupTestResult[],
covered: { group: string; kind: string }[],
): GroupTestResult[] {
const supersedes = (r: GroupTestResult) =>
covered.some((c) => c.group === r.group && (!c.kind || !r.kind || c.kind === r.kind))
const kept = prev.filter((r) => r.group && !supersedes(r))
if (kept.length <= CARRY_MAX) return kept
return [...kept].sort((a, b) => b.tested_unix - a.tested_unix).slice(0, CARRY_MAX)
}
let groupTest: GroupTestStatus = { running: false, done: 0, total: 0, scope: [], results: [] }
let groupTestQueue: string[] = []
/** A URL-seeded run stays running instead of draining (see the seed block below). */
let groupTestPinned = false
/** The in-flight run is a NODE run, so the row it lands is a node's. One board
* carries both kinds and only `kind` tells them apart — see GroupTestResult. */
let groupTestNodeRun = false
export async function postGroupsTest(name = ''): Promise<GroupTestStart> {
/**
* The single-NODE branch, `?mock&nodetest=<mode>`.
*
* All THREE refusals are reachable, because they are three different facts and
* a panel that draws one "error" for all of them is untestable in the way that
* matters:
*
* nodetest=400 → no name was sent (there is no "test every node")
* nodetest=404 → the daemon's configuration has no node by that name
* nodetest=503 → the configuration could not be read: an UNKNOWN
*
* 404 gets a knob rather than being left to "type a name that does not exist",
* because there is no way to type one: every button on the page carries a name
* the fixture's own config holds. On a real router it is the row that was
* renamed or removed under the open page — rare, and precisely the reason it
* needs to be exercisable at all.
*
* And all three READINGS:
*
* (default) → measured on demand, and it answered
* nodetest=dead → measured on demand, and it did not. A health verdict.
* nodetest=unmeasured → nothing measured it (source:''), which sits next to
* ok:false and is NOT a death. This is the pair the UI
* has to draw differently or the instrument is useless.
*/
const NODE_TEST_SHAPES: Record<string, Omit<GroupTestResult, 'group' | 'tested_unix'>> = {
ok: {
selected: '',
kind: 'node',
source: 'on-demand',
delay_ms: 61,
exit_ip: '185.12.34.56',
exit_country: 'NL',
ok: true,
error: '',
},
dead: {
selected: '',
kind: 'node',
source: 'on-demand',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: 'checked now over this node’s own configured path: the probe did not get through',
},
unmeasured: {
selected: '',
kind: 'node',
source: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error:
'no such node in the running engine — it is not in the applied configuration (not applied yet, or dropped as unusable)',
},
}
/** The node-test outcome this session is pinned to. Unrecognised ⇒ the measured
* success, never a silent failure mode. */
function nodeTestMode(): string {
const want = (mockParam('nodetest') ?? '').trim()
const refusal = want === '400' || want === '404' || want === '503'
return want in NODE_TEST_SHAPES || refusal ? want : 'ok'
}
function nodeShapeFor(name: string): GroupTestResult {
const mode = nodeTestMode()
const base = NODE_TEST_SHAPES[mode] ?? NODE_TEST_SHAPES.ok
return { ...base, group: name, tested_unix: Math.floor(Date.now() / 1000) }
}
export async function postGroupsTest(name = '', kind: TestKind = ''): Promise<GroupTestStart> {
await wait(60)
if (kind === 'node') return postNodeTest(name)
if (groupTest.running) return { started: false, reason: 'already running' }
// Empty name = every group AND every chain, exactly like the daemon.
const all = [
@@ -1436,16 +1938,67 @@ export async function postGroupsTest(name = ''): Promise<GroupTestStart> {
if (targets.length === 0) return { started: false, reason: `no group or chain named “${name}”` }
groupTestQueue = [...targets]
groupTestPinned = false
groupTestNodeRun = false
groupTest = {
running: true,
done: 0,
total: targets.length,
// The names this run covers — exactly what the panel tests card membership
// against. One name for a per-group run, every name for a run-all.
// against, both for the in-progress badge and for telling a CARRIED row from
// one this run produced. One name for a per-group run, every name for a
// run-all.
scope: [...targets],
// A re-test of ONE group replaces just that group's row and keeps the rest,
// exactly as a per-group daemon run would.
results: groupTest.results.filter((r) => !targets.includes(r.group)),
// CARRIED FORWARD, because engine.startTestRun carries: the board is no
// longer wiped per run, so a reading this run is not about survives with its
// own tested_unix. A group/chain run marks its targets with an EMPTY kind —
// the recoverable direction, matching any row of that name — because the
// fixture does not resolve which of the two each name is.
results: carryForward(
groupTest.results,
targets.map((group) => ({ group, kind: '' })),
),
}
return { started: true }
}
/**
* POST {"name":"<node>","kind":"node"} — the same singleton, the same run, the
* same GET below. The three refusals are thrown as ApiErrorLike so the page's
* error path sees the status codes the daemon really sends.
*/
async function postNodeTest(name: string): Promise<GroupTestStart> {
const mode = nodeTestMode()
if (mode === '400' || !name) {
throw new ApiErrorLike(
400,
'a node test needs a name: POST {"name":"<node>","kind":"node"}. There is deliberately no "test every node" — that sweep was removed',
)
}
if (mode === '503') {
throw new ApiErrorLike(
503,
'could not read the configuration to check that this node exists, so the test was not started — this is an unknown, not a verdict about the node: uci show shater: exit status 1',
)
}
if (mode === '404' || !(CONFIG.Nodes ?? []).some((n) => n.Name === name)) {
throw new ApiErrorLike(404, 'no node with that name in the configuration')
}
if (groupTest.running) return { started: false, reason: 'already running' }
// The case the carry-forward was built for: testing ONE node used to blank
// every group and chain card on the Targets screen, and nothing on screen said
// why. Now the other rows survive — identified by (name, kind), so a group
// that happens to share this node's name is NOT dropped — and they arrive on
// the board as carried readings the panel has to date rather than pass off as
// this run's.
groupTestQueue = [name]
groupTestPinned = false
groupTestNodeRun = true
groupTest = {
running: true,
done: 0,
total: 1,
scope: [name],
results: carryForward(groupTest.results, [{ group: name, kind: 'node' }]),
}
return { started: true }
}
@@ -1458,7 +2011,10 @@ export async function getGroupsTest(): Promise<GroupTestStatus> {
groupTest = {
...groupTest,
done: groupTest.done + 1,
results: [...groupTest.results, shapeFor(next, groupTest.done)],
results: [
...groupTest.results,
groupTestNodeRun ? nodeShapeFor(next) : shapeFor(next, groupTest.done),
],
}
}
if (groupTestQueue.length === 0) groupTest = { ...groupTest, running: false }
@@ -1495,5 +2051,41 @@ if (typeof location !== 'undefined') {
}
}
/**
* Land on a FINISHED board that holds both kinds of row, `?mock&board=carried`.
*
* The carried row is reachable by hand — test one target, then test another —
* but only after two runs and a wait, which makes the one rendering that must
* never be got wrong the hardest one to look at. This seeds it directly: a
* single group in scope with a reading taken just now, and every other target
* carrying a reading from ELEVEN MINUTES AGO.
*
* Eleven minutes, not eleven seconds, because the point is a row that is still
* true and no longer current. A fixture where every row is seconds old would let
* "shows when it was taken" pass while showing the same time twice.
*/
if (typeof location !== 'undefined') {
if (new URLSearchParams(location.search).get('board') === 'carried') {
const now = Math.floor(Date.now() / 1000)
const all = [
...(CONFIG.Groups ?? []).map((g) => g.Name),
...(CONFIG.Chains ?? []).map((c) => c.Name),
]
const fresh = all.slice(0, 1)
groupTestPinned = false
groupTestQueue = []
groupTest = {
running: false,
done: fresh.length,
total: fresh.length,
scope: [...fresh],
results: all.map((group, i) => ({
...shapeFor(group, i),
tested_unix: fresh.includes(group) ? now : now - 11 * 60,
})),
}
}
}
/** Exposed for potential UI hints; not part of the wire contract. */
export const isArmed = () => armed
+498
View File
@@ -0,0 +1,498 @@
/* Alerts section (rendered on Settings) — inherits the Faceplate tokens and the
* shared page chrome from App.css (.toast, .mono). Every rule below is a
* one-to-one copy of the DNS.css rule the markup used before the section moved
* here, renamed `dns-*` → `alr-*` so nothing collides. Orange stays an accent. */
/* ---- section shell (matches the Settings group plates one-to-one) ---- */
.alr-section {
margin-top: calc(var(--u, 8px) * 3.5);
}
.alr-sec-hd {
display: flex;
align-items: baseline;
gap: 12px;
padding-bottom: 10px;
border-bottom: 1px solid var(--groove);
}
.alr-sec-title {
margin: 0;
font-family: var(--font-mono);
font-size: 13px;
font-weight: 700;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--dim);
}
.alr-sec-count {
font-size: 11px;
letter-spacing: 0.06em;
color: var(--faint);
}
.alr-sec-note {
margin: 10px 2px 0;
font-family: var(--font-sans);
font-size: 12.5px;
line-height: 1.55;
color: var(--dim);
max-width: 56ch;
}
/* ---- add form ---- */
.alr-add {
display: flex;
flex-direction: column;
gap: 10px;
margin-top: calc(var(--u, 8px) * 2);
}
.alr-add-top {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px;
}
.alr-input {
min-width: 0;
padding: 9px 12px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 12.5px;
letter-spacing: 0.02em;
box-shadow: 0 1px 2px var(--shadow) inset;
transition: border-color 0.15s, box-shadow 0.15s;
}
.alr-input::placeholder {
color: var(--faint);
}
.alr-input:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-input:disabled {
opacity: 0.55;
}
.alr-input--name {
flex: 0 1 14rem;
}
/* A caption under a control, in the panel's quiet voice — used for the rule that
* an empty token box keeps the stored token, and for what a type switch does with
* the other type's settings. Not --amber: neither is a warning, both are the
* plain behaviour of the form, and colouring them would train the eye to skip the
* amber that does mean caution (.alr-note). */
.alr-hint {
margin: -4px 2px 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--dim);
max-width: 56ch;
}
/* The same plate as the add form, inline in the list and accented so it reads as
* the row currently being edited (mirrors .rt-add.rt-edit-form on Routing). */
.alr-edit-item {
list-style: none;
}
.alr-add.alr-edit-form {
margin-top: 0;
padding: 12px 14px;
border: 1px solid var(--accent);
border-radius: 8px;
background: var(--raised);
box-shadow: 0 1px 0 var(--edge) inset;
}
.alr-cancel {
padding: 6px 10px;
border: 0;
background: none;
color: var(--dim);
font-family: var(--font-sans);
font-size: 12px;
text-decoration: underline;
cursor: pointer;
}
.alr-cancel:hover:not(:disabled) {
color: var(--ink);
}
.alr-cancel:disabled {
opacity: 0.55;
cursor: default;
}
/* segmented type picker */
.alr-seg {
display: inline-flex;
border: 1px solid var(--groove);
border-radius: 7px;
overflow: hidden;
background: var(--sink);
}
.alr-seg-btn {
padding: 8px 14px;
border: 0;
background: transparent;
color: var(--dim);
font-family: var(--font-mono);
font-size: 11px;
letter-spacing: 0.08em;
text-transform: uppercase;
cursor: pointer;
transition: background 0.15s, color 0.15s;
}
.alr-seg-btn + .alr-seg-btn {
border-left: 1px solid var(--groove);
}
.alr-seg-btn.on {
background: var(--accent);
color: #fff;
}
.alr-seg-btn:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
}
.alr-resp {
display: inline-flex;
align-items: center;
gap: 8px;
}
.alr-resp-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-select {
padding: 8px 10px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 11.5px;
letter-spacing: 0.04em;
cursor: pointer;
}
.alr-select:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-add-actions {
display: flex;
align-items: center;
justify-content: flex-end;
gap: 14px;
flex-wrap: wrap;
}
.alr-field-err {
flex: 1;
min-width: 0;
margin: 0;
font-family: var(--font-mono);
font-size: 11.5px;
line-height: 1.5;
color: var(--crit);
}
/* ---- rows ---- */
.alr-rows {
list-style: none;
margin: calc(var(--u, 8px) * 2) 0 0;
padding: 0;
display: flex;
flex-direction: column;
gap: 8px;
}
.alr-row {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
padding: 12px 14px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(
180deg,
var(--raised),
color-mix(in srgb, var(--raised) 82%, var(--panel))
);
box-shadow: 0 1px 0 var(--edge) inset;
}
.alr-row-main {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 4px;
}
.alr-row-l1 {
display: flex;
align-items: center;
gap: 8px;
flex-wrap: wrap;
}
.alr-row-name {
font-family: var(--font-mono);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
color: var(--ink);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 24ch;
}
.alr-row-l2 {
display: flex;
align-items: center;
gap: 10px;
flex-wrap: wrap;
font-size: 11.5px;
letter-spacing: 0.02em;
}
.alr-row-detail {
color: var(--dim);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 40ch;
}
/* badge — groove-bordered, not orange (accent stays reserved) */
.alr-badge {
display: inline-block;
padding: 2px 7px;
border: 1px solid var(--groove);
border-radius: 5px;
background: color-mix(in srgb, var(--sink) 60%, transparent);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 600;
letter-spacing: 0.1em;
text-transform: uppercase;
color: var(--dim);
white-space: nowrap;
}
.alr-badge--accent {
border-color: color-mix(in srgb, var(--accent) 55%, var(--groove));
color: var(--accent);
}
.alr-masked {
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: 0.08em;
color: var(--faint);
text-transform: uppercase;
cursor: help;
}
.alr-row-actions {
flex: none;
display: flex;
align-items: center;
gap: 8px;
}
.alr-row-actions button {
padding: 6px 12px;
font-size: 10.5px;
}
.alr-del {
flex: none;
padding: 6px 12px;
font-size: 10.5px;
}
/* ---- empty plate ---- */
.alr-empty {
margin-top: calc(var(--u, 8px) * 2);
padding: calc(var(--u, 8px) * 3);
border: 1px dashed var(--groove);
border-radius: 9px;
background: color-mix(in srgb, var(--raised) 55%, transparent);
text-align: center;
}
.alr-empty-title {
display: block;
font-size: 13px;
font-weight: 700;
letter-spacing: 0.06em;
color: var(--dim);
}
.alr-empty-body {
margin: 8px auto 0;
max-width: 48ch;
font-family: var(--font-sans);
font-size: 13px;
line-height: 1.55;
color: var(--dim);
}
/* ---- loading skeleton ---- */
.alr-skel {
height: 62px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(90deg, var(--raised), var(--sink), var(--raised));
background-size: 200% 100%;
animation: alr-skel-shift 1.4s ease-in-out infinite;
}
@keyframes alr-skel-shift {
from {
background-position: 200% 0;
}
to {
background-position: -200% 0;
}
}
/* the per-row delivery picker sits inline in the row */
.alr-detour {
flex: none;
display: flex;
flex-direction: column;
gap: 5px;
min-width: 0;
}
.alr-detour-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-detour-select {
max-width: 22rem;
}
/* current delivery-path readout on the row */
.alr-path {
color: var(--faint);
white-space: nowrap;
}
.alr-path[data-active='on'] {
color: var(--dim);
}
.alr-path-name {
color: var(--led-on);
font-weight: 600;
}
.alr-path[data-missing='y'] .alr-path-name {
color: var(--amber);
}
.alr-path-flag {
color: var(--amber);
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.alr-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.alr-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.alr-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-row delivery controls sit inline in the row (like .alr-detour) */
.alr-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.alr-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.alr-note--row {
margin-top: 2px;
}
/* alert event checkboxes */
.alr-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.alr-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.alr-event input {
accent-color: var(--fp-accent, currentColor);
}
/* ---- responsive ---- */
@media (max-width: 640px) {
.alr-row {
flex-wrap: wrap;
}
.alr-row-main {
flex-basis: calc(100% - 90px);
}
.alr-del {
margin-left: auto;
}
.alr-input--name {
flex-basis: 100%;
}
.alr-detour {
flex-basis: 100%;
order: 3;
flex-wrap: wrap;
}
/* A <select> won't shrink below its widest option unless it's allowed to:
without min-width:0 the long detour labels push the page into a horizontal
scroll at 390px. Let them fill the row and clip instead. */
.alr-detour-select,
.alr-resp .alr-select {
max-width: 100%;
width: 100%;
min-width: 0;
}
.alr-resp {
display: flex;
flex-wrap: wrap;
max-width: 100%;
}
.alr-ctl {
flex-basis: 100%;
order: 3;
}
}
@media (prefers-reduced-motion: reduce) {
.alr-skel {
animation: none;
}
.alr-input,
.alr-seg-btn {
transition: none;
}
}
+790
View File
@@ -0,0 +1,790 @@
import './Alerts.css'
import { useCallback, useMemo, useState } from 'react'
import { Button, Toggle, useConfirm } from '../components'
import type { Alert, Model } from '../api'
import {
buildAlert,
draftFromAlert,
hasStoredToken,
validateAlertDraft,
TOKEN_KEEP_HINT,
TOKEN_NEW_HINT,
TYPE_SWITCH_NOTE,
} from '../alertEdit'
import type { AlertDraft, AlertType } from '../alertEdit'
// The Alerts section — out-of-band notifications (Telegram bot / webhook) for
// kill-switch trips, apply failures, new devices and subscription expiry. It
// lived at the bottom of the DNS page, which is the last place an operator
// looking for "tell me when the tunnel dies" would think to look; it now renders
// as a group on Settings. The component owns no I/O: every mutation goes through
// the `onSave` prop so Settings keeps a single dirty banner and a single toast.
//
// NOTE on duplication: the detour helpers below (DetourCatalog, canonDetour,
// detourValues, describeDetour, DetourSelect) plus asArray / uniqueName /
// maskUrl / EmptyPlate are deliberate copies of the ones in DNS.tsx. DNS keeps
// its own for resolvers and DNS rules; extracting a shared module would couple
// two pages that otherwise share nothing, and that refactor is out of scope
// here. If a third consumer ever appears, promote them then.
// ---- local Model extension --------------------------------------------------
/** The Model with the Alerts slice surfaced (index-signature passthrough). */
type AlertsModel = Model & { Alerts?: Alert[] | null }
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers (copies of DNS.tsx — see the header note) ----------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
/** A remote URL often carries a token in its query/path — show host only. */
function maskUrl(url: string): { host: string; masked: boolean } {
try {
const u = new URL(url)
return { host: u.host, masked: u.search !== '' || u.pathname.replace(/\/+$/, '') !== '' }
} catch {
return { host: url || '—', masked: false }
}
}
function uniqueName(base: string, taken: Set<string>): string {
const seed = base.trim() || 'alert'
if (!taken.has(seed)) return seed
let i = 2
while (taken.has(`${seed}-${i}`)) i++
return `${seed}-${i}`
}
/** The live targets an alert's delivery can be pinned to (the picker). */
interface DetourCatalog {
groups: string[]
chains: string[]
egresses: { name: string; type: string }[]
nodes: string[]
}
/**
* Normalise a stored `Via` to a picker option value. Empty/`direct` ⇒
* `direct`; already-prefixed values (`group:`/`chain:`/`egress:`/`node:`) pass
* through; a bare legacy name is resolved against the catalog so a still-valid
* setup isn't mislabelled; anything unresolved is kept verbatim (shown stale).
*/
function canonDetour(raw: string | undefined, cat: DetourCatalog): string {
const d = (raw ?? '').trim()
if (!d || d.toLowerCase() === 'direct') return 'direct'
if (/^(node|group|chain|egress):/i.test(d)) return d
if (cat.egresses.some((e) => e.name === d)) return `egress:${d}`
if (cat.groups.includes(d)) return `group:${d}`
if (cat.chains.includes(d)) return `chain:${d}`
if (cat.nodes.includes(d)) return `node:${d}`
return d
}
/** Every valid option value for a catalog, including `direct`. */
function detourValues(cat: DetourCatalog): Set<string> {
const s = new Set<string>(['direct'])
for (const g of cat.groups) s.add(`group:${g}`)
for (const c of cat.chains) s.add(`chain:${c}`)
for (const e of cat.egresses) s.add(`egress:${e.name}`)
for (const n of cat.nodes) s.add(`node:${n}`)
return s
}
/** Describe a canonical detour value for the row readout. */
function describeDetour(
canon: string,
cat: DetourCatalog,
valid: Set<string>,
): { direct: boolean; prefix: string; name: string; missing: boolean } {
if (canon === 'direct') return { direct: true, prefix: '', name: '', missing: false }
const i = canon.indexOf(':')
const kind = i === -1 ? '' : canon.slice(0, i)
const name = i === -1 ? canon : canon.slice(i + 1)
const missing = !valid.has(canon)
let prefix = 'via'
if (kind === 'group') prefix = 'via group'
else if (kind === 'chain') prefix = 'via chain'
else if (kind === 'node') prefix = 'via node'
else if (kind === 'egress') {
const eg = cat.egresses.find((e) => e.name === name)
prefix = eg?.type === 'interface' ? 'via interface' : 'via egress'
}
return { direct: false, prefix, name, missing }
}
// ---- section ----------------------------------------------------------------
export function AlertsSection({
config,
busy,
loading,
onSave,
}: {
/** Full desired-state model; null until it has loaded. */
config: Model | null
/** A save/apply is in flight — controls lock. */
busy: boolean
/** The config is still loading — show a skeleton row. */
loading: boolean
/** Persist the whole next model; resolves true on success (Settings' `save`). */
onSave: (next: Model, okMsg: string) => Promise<boolean>
}): JSX.Element {
const confirm = useConfirm()
const model = config as AlertsModel | null
const alerts = useMemo<Alert[]>(() => asArray(model?.Alerts), [model])
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Built straight from the Model with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({
groups: asArray(config?.Groups).map((g) => g.Name),
chains: [],
egresses: asArray(config?.Egresses).map((e) => ({ name: e.Name, type: e.Type })),
nodes: asArray(config?.Nodes).map((n) => n.Name),
}),
[config],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
// Which row is open for editing — one at a time, like the rule bus. Keyed by
// INDEX and not by name: names are the thing an edit can change, and a config
// is free to carry two channels with the same one.
const [editing, setEditing] = useState<number | null>(null)
// ---- mutations — all writes go through onSave -----------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!model) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return onSave({ ...model, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[model, alerts, onSave],
)
/**
* Save an edited channel in place.
*
* The next Alert is built by alertEdit.buildAlert, which is where the rule about
* unshown fields lives — an empty token box keeps the stored token, and the
* other type's settings survive a type switch. Nothing about that decision is
* repeated here, so there is one place it can be got wrong.
*/
const editAlert = useCallback(
async (idx: number, next: Alert): Promise<boolean> => {
if (!model) return false
const ok = await onSave(
{ ...model, Alerts: alerts.map((a, i) => (i === idx ? next : a)) },
`Updated ${next.Name}`,
)
if (ok) setEditing(null)
return ok
},
[model, alerts, onSave],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void onSave({ ...model, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[model, alerts, onSave],
)
const removeAlert = useCallback(
async (idx: number) => {
if (!model) return
const target = alerts[idx]
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void onSave({ ...model, Alerts: next }, `Deleted ${target.Name}`)
},
[model, alerts, onSave, confirm],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!model) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void onSave(
{ ...model, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[model, alerts, onSave],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void onSave(
{ ...model, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[model, alerts, onSave],
)
return (
<div className="alr-section" aria-label="Alerts">
<header className="alr-sec-hd">
<h2 className="alr-sec-title">Alerts</h2>
<span className="alr-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="alr-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AlertForm
stored={null}
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onSubmit={addAlert}
/>
{loading ? (
<ul className="alr-rows" aria-hidden="true">
<li className="alr-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="alr-rows">
{alerts.map((a, i) =>
editing === i ? (
<li className="alr-edit-item" key={`edit-${i}`}>
<AlertForm
stored={a}
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onSubmit={(next) => editAlert(i, next)}
onCancel={() => setEditing(null)}
/>
</li>
) : (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
editingOther={editing !== null}
catalog={alertCatalog}
valid={alertValid}
onEdit={() => setEditing(i)}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
),
)}
</ul>
)}
</div>
)
}
// ---- alert add form + row ----------------------------------------------------
/**
* ONE form for adding a channel and for editing one, chosen by `stored`.
*
* They were never going to be two forms for long. Everything the add form asks —
* type, token, chat ID, URL, events, delivery — is a thing an existing channel
* must be able to change, and the reason it could not was simply that no editor
* existed: the row offered Enabled/Via/Fallback and nothing else, so a typo in a
* chat ID meant deleting the channel and going back to BotFather for a token you
* already owned. Sharing the component is what keeps the two paths from drifting
* into different validation rules, which is exactly how the first one grew a
* "Telegram needs a bot token" check that an edit could not satisfy.
*
* The one asymmetry is the secret, and it is stated in the interface rather than
* implied: the token box starts empty when editing and the hint under it says
* that empty means keep. The rule itself lives in alertEdit.buildAlert.
*/
function AlertForm({
stored,
busy,
disabled,
taken,
catalog,
valid,
onSubmit,
onCancel,
}: {
/** The channel being edited, or null to add a new one. */
stored: Alert | null
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onSubmit: (a: Alert) => Promise<boolean>
onCancel?: () => void
}) {
const editing = stored !== null
const initial = useMemo<AlertDraft>(
() =>
stored
? draftFromAlert(stored, canonDetour(stored.Via, catalog))
: {
name: '',
type: 'telegram',
token: '',
chatId: '',
url: '',
events: ['killswitch'],
via: 'direct',
fallback: false,
},
// Built once per mounted form: re-deriving it as the catalog polls would
// throw away half-typed input. The list keys each form by row, so opening a
// different one mounts a fresh component with a fresh prefill.
// eslint-disable-next-line react-hooks/exhaustive-deps
[],
)
const [name, setName] = useState(initial.name)
const [type, setType] = useState<AlertType>(initial.type)
const [token, setToken] = useState('')
const [chatId, setChatId] = useState(initial.chatId)
const [url, setUrl] = useState(initial.url)
const [events, setEvents] = useState<string[]>(initial.events)
const [via, setVia] = useState(initial.via)
const [fallback, setFallback] = useState(initial.fallback)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName(initial.name)
setType(initial.type)
setToken('')
setChatId(initial.chatId)
setUrl(initial.url)
setEvents(initial.events)
setVia(initial.via)
setFallback(initial.fallback)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const draft: AlertDraft = { name, type, token, chatId, url, events, via, fallback }
const problem = validateAlertDraft(draft, taken, stored)
if (problem) {
setErr(problem)
return
}
setErr(null)
const ok = await onSubmit(buildAlert(stored, draft))
if (ok && !editing) reset()
}
const routed = via !== 'direct'
const keepsToken = editing && hasStoredToken(stored)
const typeSwitched = editing && stored!.Type !== type
return (
<form
className={editing ? 'alr-add alr-edit-form' : 'alr-add'}
aria-label={editing ? `Edit alert ${stored!.Name}` : 'Add an alert channel'}
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="alr-add-top">
<input
className="alr-input alr-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="alr-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{typeSwitched && <p className="alr-hint">{TYPE_SWITCH_NOTE}</p>}
{type === 'telegram' ? (
<>
<input
className="alr-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder={keepsToken ? 'Bot token — leave empty to keep' : 'Bot token (kept secret)'}
aria-label={keepsToken ? 'Telegram bot token — leave empty to keep the stored one' : 'Telegram bot token'}
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
{/* The rule, in the interface rather than in someone's head. A masked
field with no caption means "type it again"; this one does not. */}
<p className="alr-hint">{keepsToken ? TOKEN_KEEP_HINT : TOKEN_NEW_HINT}</p>
<input
className="alr-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="alr-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="alr-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="alr-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="alr-delivery">
<label className="alr-resp">
<span className="alr-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="alr-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="alr-add-actions">
{err && (
<p className="alr-field-err" role="alert">
{err}
</p>
)}
{editing && (
<button type="button" className="alr-cancel" onClick={onCancel} disabled={busy}>
Cancel
</button>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : editing ? 'Save alert' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
editingOther,
catalog,
valid,
onEdit,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
/** Another row is open for editing — freeze this one's controls. */
editingOther: boolean
catalog: DetourCatalog
valid: Set<string>
onEdit: () => void
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
// While a row is open for editing, every other row's writes are frozen: an
// Alerts save PUTs the whole list, so a toggle landing under an open form would
// save one edit over the other. Edit itself stays live — it just moves the form.
const frozen = busy || editingOther
return (
<li className="alr-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={frozen}
/>
<div className="alr-row-main">
<div className="alr-row-l1">
<span className="alr-row-name">{alert.Name}</span>
<span className="alr-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="alr-badge alr-badge--accent">
{e}
</span>
))}
</div>
<div className="alr-row-l2 mono">
<span className="alr-row-detail">{detail.text}</span>
{detail.masked && (
<span className="alr-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="alr-path">direct</span>
) : (
<span className="alr-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="alr-path-name">{route.name}</strong>
{route.missing && <span className="alr-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="alr-note alr-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="alr-ctl">
<label className="alr-detour">
<span className="alr-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={frozen || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
<div className="alr-row-actions">
<Button onClick={onEdit} disabled={busy} aria-label={`Edit alert ${alert.Name}`}>
Edit
</Button>
<Button
className="alr-del"
onClick={onDelete}
disabled={frozen}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</div>
</li>
)
}
/** The live delivery picker: option list built from the Model's targets. */
function DetourSelect({
value,
catalog,
valid,
busy,
disabled,
ariaLabel,
onChange,
directLabel = 'Direct (no proxy)',
}: {
value: string // canonical value
catalog: DetourCatalog
valid: Set<string>
busy: boolean
disabled: boolean
ariaLabel: string
onChange: (v: string) => void
directLabel?: string
}) {
const missing = value !== 'direct' && !valid.has(value)
return (
<select
className="alr-select alr-detour-select"
value={value}
onChange={(e) => onChange(e.target.value)}
disabled={busy || disabled}
aria-label={ariaLabel}
>
<option value="direct">{directLabel}</option>
{catalog.groups.length > 0 && (
<optgroup label="Groups">
{catalog.groups.map((g) => (
<option key={g} value={`group:${g}`}>
Group {g} (balancer)
</option>
))}
</optgroup>
)}
{catalog.chains.length > 0 && (
<optgroup label="Chains">
{catalog.chains.map((c) => (
<option key={c} value={`chain:${c}`}>
Chain {c}
</option>
))}
</optgroup>
)}
{catalog.egresses.length > 0 && (
<optgroup label="Interfaces / egresses">
{catalog.egresses.map((e) => (
<option key={e.name} value={`egress:${e.name}`}>
Interface/egress {e.name}
{e.type ? ` (${e.type})` : ''}
</option>
))}
</optgroup>
)}
{catalog.nodes.length > 0 && (
<optgroup label="Nodes">
{catalog.nodes.map((n) => (
<option key={n} value={`node:${n}`}>
Node {n}
</option>
))}
</optgroup>
)}
{missing && <option value={value}>{value} (missing)</option>}
</select>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="alr-empty">
<span className="alr-empty-title mono">{title}</span>
<p className="alr-empty-body">{body}</p>
</div>
)
}
+180 -52
View File
@@ -10,8 +10,9 @@ import {
getStatus,
ApiError,
} from '../api'
import type { Globals, Status } from '../api'
import { engineReadout } from '../planeState'
import type { Model, Status } from '../api'
import { applyRisk, engineReadout, killSwitchReadout, serviceIntent } from '../planeState'
import type { ApplyRisk } from '../planeState'
import { onPendingConfirmExpire, usePendingConfirm } from '../pendingConfirm'
// Short, readable config hash — drops the "sha256:" prefix like the footer does.
@@ -48,7 +49,11 @@ interface ActionResult {
export default function Apply() {
const [status, setStatus] = useState<Status | null>(null)
const [statusError, setStatusError] = useState<string | null>(null)
const [globals, setGlobals] = useState<Globals | null>(null)
// The WHOLE model, not just Globals: the pre-apply warning below is decided
// from the rules, the inbounds and the resolvers as well as the globals, and
// splitting the read would have let those two drift a poll apart.
const [config, setConfig] = useState<Model | null>(null)
const globals = config?.Globals ?? null
const [configError, setConfigError] = useState<string | null>(null)
const [busy, setBusy] = useState<Busy>(null)
@@ -93,8 +98,7 @@ export default function Apply() {
const loadConfig = useCallback(async () => {
try {
setConfigError(null)
const cfg = await getConfig()
setGlobals(cfg.Globals)
setConfig(await getConfig())
} catch (e) {
setConfigError(msg(e))
}
@@ -103,30 +107,64 @@ export default function Apply() {
void loadConfig()
}, [loadConfig])
// The window running out is the daemon reverting on its own — observe it and
// say so. The countdown itself ticks inside usePendingConfirm; this only reacts
// to the end of it, and the store makes sure that fires exactly once even with
// the app-wide band mounted alongside.
// The window running out does NOT mean the daemon rolled back.
//
// apply.ArmRollback captures the data-plane generation when it arms, and on
// expiry it compares. If anything re-applied the plane in between — another
// panel apply, SIGHUP, a hotplug or the once-a-minute cron reconcile, the WAN
// profile auto-switch — it disarms and KEEPS the running config, logging "NOT
// rolling back" and nothing else. That is the common case on a production
// router, and this page used to print "daemon auto-rolled back to last-good
// config" for it: a confident report of an event that did not happen, with a
// hash pair underneath that quietly said "unchanged".
//
// The panel cannot see which branch ran — the daemon says so only in its log.
// So it reports the one thing it CAN observe, the live config hash, and waits
// for the revert to land before reading it (a rollback is a full re-apply and
// does not complete the instant the timer fires).
const liveHashRef = useRef('')
liveHashRef.current = status?.hash ?? ''
useEffect(
() =>
onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Auto-rolled back')
void (async () => {
const after = (await refreshStatus())?.hash ?? ''
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed — daemon auto-rolled back to last-good config.',
before,
after,
})
})()
}),
[flash, refreshStatus],
)
useEffect(() => {
let cancelled = false
const off = onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Confirm window elapsed')
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed. Reading what the daemon did…',
before,
after: before,
})
void (async () => {
let after = before
for (let i = 0; i < 4 && !cancelled; i++) {
await new Promise((r) => window.setTimeout(r, 1500))
if (cancelled) return
after = (await refreshStatus())?.hash ?? after
if (after !== before) break
}
if (cancelled) return
setResult({
kind: 'expire',
tone: 'warn',
text:
after !== before
? 'Confirm window elapsed and the live config changed — the daemon reverted to its last-good config.'
: 'Confirm window elapsed and the live config has not changed, so this config is still running. ' +
'The daemon only reverts if nothing else re-applied the data plane while the window was open; ' +
'otherwise it stands down and keeps what is live. Which one happened is in the daemon log — ' +
'download it from Settings, or run `logread -e shater`.',
before,
after,
})
})()
})
return () => {
cancelled = true
off()
}
}, [flash, refreshStatus])
const confirmWindow = globals?.ConfirmTimeout ?? 0
@@ -232,15 +270,46 @@ export default function Apply() {
flash('Rollback failed')
} finally {
setBusy(null)
// A rollback restores the PREVIOUS config, so the desired state this page
// reads — and the pre-apply warning derived from it — is now stale.
void loadConfig()
}
}, [status, flash, refreshStatus])
}, [status, flash, loadConfig, refreshStatus])
// ---- derived display state (mirrors Overview's LED semantics) ----
//
// The LIVE kill-switch wins over the saved one, exactly as on Overview: this row
// is a status readout, and the config on disk can already differ from what is
// installed. Falls back to the config only while /api/status is unread.
const killArmed = (status?.kill_switch ?? globals?.KillSwitch ?? 'closed') === 'closed'
// Whether that setting is actually installed — same three-plus-unknown reading
// as Overview, so the two pages cannot disagree about the same router. There is
// no separate `killArmed` here any more: it compared the raw string (so "Closed"
// read as fail-OPEN) and, being a boolean, could not express "the configuration
// could not be read". Both facts come off this one readout now.
// THE ONE FACT THAT RE-READS THE OTHER THREE PIPS. With the service switched
// off, "engine stopped" / "no table" / "kill-switch not in effect" are all true
// and none of them is a fault — this row showed three crit lamps on a router
// that had installed correctly and been touched by nobody. Under `on`, and
// under an unreadable config (`unknown`), every one of them keeps its crit.
const serviceOff = serviceIntent(status) === 'off'
const kill = killSwitchReadout(status, globals?.KillSwitch)
const killWord =
kill.state === 'open'
? 'open'
: kill.state === 'armed'
? 'fail-closed'
: kill.state === 'standby'
? // Configured, nothing to guard yet. Reads as pending, not as failed —
// which is the whole difference between this and `inert` below.
'closed · standing by'
: kill.state === 'inert'
? 'closed · not in effect'
: // The unknown branch splits: "closed · not reported" asserts the policy
// and doubts only the install, which is wrong when the policy itself is
// a placeholder from a configuration nothing could read.
kill.setting === 'not known'
? 'not known'
: 'closed · not reported'
// Every engine mark on this page comes from ONE reading, and that reading is
// able to say "stopped" — see planeState.engineState for why `status.running`
// could not. This page is where someone lands when the network is down; three
@@ -251,8 +320,26 @@ export default function Apply() {
// that is a leak (crit); under an open one it is the documented choice (amber).
// It used to go amber whenever `running` was true — i.e. always — and unlit
// otherwise, so the one state worth shouting about had no colour of its own.
const dataVariant: LedVariant = status?.table ? 'on' : !status ? 'off' : killArmed ? 'crit' : 'amber'
const configVariant: LedVariant = status?.enabled ? 'on' : 'amber'
// The service being off is checked BEFORE the kill-switch, because the
// fail-closed reading is the one that turned "there is no plane, and none was
// asked for" into a leak alarm.
const dataVariant: LedVariant =
status?.table
? 'on'
: !status || serviceOff
? 'off'
: kill.state !== 'open'
? 'crit'
: 'amber'
// "no table" is the installed truth; on a switched-off service it needs the
// reason attached or it reads as the thing that failed.
const dataWord = status?.table ? 'nft installed' : serviceOff ? 'none — service off' : 'no table'
// `enabled` is sourced from the configuration, so it means nothing when that
// could not be read (Status.config_readable): unlit, not amber, and the pip
// beside it says so rather than printing "disabled".
const configUnreadable = status?.config_readable === false
const configVariant: LedVariant = configUnreadable ? 'off' : status?.enabled ? 'on' : 'amber'
const configWord = configUnreadable ? 'unreadable' : status?.enabled ? 'enabled' : 'disabled'
const liveHash = short(status?.hash ?? '')
const pct = armed ? Math.max(0, Math.round((armed.remaining / armed.pending.total) * 100)) : 0
@@ -262,6 +349,11 @@ export default function Apply() {
// is nothing to roll back to, so the idle control is hidden entirely (no dead-end).
const canRollback = status?.can_rollback ?? false
// Predicted from the config this apply would install — never from `status`,
// which describes the config already running. Null on every config that does
// not have this outcome, which is nearly all of them.
const risk = applyRisk(config)
return (
<section className="page apply-page" aria-label="Apply and rollback">
{/* What's live now — engine + data-plane state in LED form. */}
@@ -275,18 +367,10 @@ export default function Apply() {
<StatusPip
label="Config"
variant={configVariant}
value={status?.enabled ? 'enabled' : 'disabled'}
/>
<StatusPip
label="Data plane"
variant={dataVariant}
value={status?.table ? 'nft installed' : 'no table'}
/>
<StatusPip
label="Kill-switch"
variant={killArmed ? 'on' : 'amber'}
value={killArmed ? 'fail-closed' : 'open'}
value={configWord}
/>
<StatusPip label="Data plane" variant={dataVariant} value={dataWord} />
<StatusPip label="Kill-switch" variant={kill.variant} value={killWord} />
</div>
{statusError && (
@@ -315,12 +399,8 @@ export default function Apply() {
led={{ variant: configVariant }}
rows={[
{ k: 'engine', v: engine.word, hot: engineVariant === 'crit' },
{
k: 'data plane',
v: status?.table ? 'nft installed' : 'no table',
hot: dataVariant === 'crit',
},
{ k: 'kill-switch', v: killArmed ? 'fail-closed' : 'open', hot: !killArmed },
{ k: 'data plane', v: dataWord, hot: dataVariant === 'crit' },
{ k: 'kill-switch', v: killWord, hot: kill.variant === 'crit' || kill.settingHot },
]}
/>
<Module
@@ -334,7 +414,10 @@ export default function Apply() {
led={{ variant: engineVariant }}
rows={[
{ k: 'state', v: engine.word, hot: engineVariant === 'crit' },
{ k: 'config', v: status?.enabled ? 'enabled' : 'disabled' },
// Same reading as the pip above — `status.enabled` is a placeholder
// when the configuration could not be read, and "disabled" is the one
// word that must not be printed for it.
{ k: 'config', v: configWord, hot: configUnreadable },
{ k: 'schema', v: globals ? `v${globals.SchemaVersion}` : '—' },
]}
/>
@@ -351,6 +434,13 @@ export default function Apply() {
/>
</div>
{/* Predicted outcome of pressing the button below. Sits directly above the
control room rather than at the top of the page: this is a statement
about the action, and it belongs where the action is. Suppressed while a
confirm window is armed — that config is already live, so the thing to
read there is the countdown, not a forecast of an apply that happened. */}
{risk && !armed && <ApplyRiskBand risk={risk} />}
{/* Control room — the signature: idle actions, or the armed countdown. */}
{armed ? (
<div className="cc" role="group" aria-label="Commit-confirm window">
@@ -374,8 +464,9 @@ export default function Apply() {
is what "live but not kept" means — so the readout survives a
reload instead of depending on what this tab remembers. */}
Applied config <span className="mono">{liveHash}</span> is live but not yet kept.
Confirm to keep it — otherwise the daemon rolls back to the last-good config when
the timer hits zero.
Confirm to keep it. At zero the daemon rolls back to the last-good config — unless
something else re-applies the data plane first, in which case it stands down and
keeps whatever is live.
</p>
<div className="cc-bar" aria-hidden="true">
<span className="cc-bar-fill" style={{ width: `${pct}%` }} />
@@ -483,6 +574,41 @@ export default function Apply() {
)
}
/**
* The one thing said BEFORE the most dangerous apply this router can do.
*
* Shared by Apply and Overview because both carry an "Apply config" button, and a
* warning that appears beside one of them is a warning half the operators never
* see. The wording is entirely {@link applyRisk}'s, for the same reason
* protectionState owns the wording of the observed states: two screens describing
* one router in two ways is how they come to disagree.
*
* It uses the crit vocabulary — the outcome is a network with no way out, which
* is exactly crit-severity — but it must not read as a fault that has already
* happened, so it leads with the tense ("BEFORE YOU APPLY") and ends in the
* accent colour: crit says how bad this is, the accent says which control fixes
* it. Nothing here disables the Apply button. It is the operator's router, the
* config is legal, and a warning that blocks the action just gets worked around.
*/
export function ApplyRiskBand({ risk }: { risk: ApplyRisk }) {
return (
<section className="risk-band" role="alert" aria-label="Before you apply">
<div className="risk-band-hd">
<Led variant="crit" />
<span className="risk-eyebrow">Before you apply</span>
</div>
<p className="risk-headline">{risk.headline}</p>
<p className="risk-detail">{risk.detail}</p>
<p className={risk.noAutoRollback ? 'risk-undo hot' : 'risk-undo'}>{risk.undo}</p>
<ul className="risk-steps">
{risk.steps.map((s) => (
<li key={s}>{s}</li>
))}
</ul>
</section>
)
}
function labelFor(kind: ActionKind): string {
switch (kind) {
case 'apply':
@@ -492,7 +618,9 @@ function labelFor(kind: ActionKind): string {
case 'rollback':
return 'Rollback'
case 'expire':
return 'Auto-rollback'
// NOT "Auto-rollback": on expiry the daemon either reverts or stands down,
// and this page cannot tell which. Name the event it did observe.
return 'Window elapsed'
}
}
+31 -67
View File
@@ -104,6 +104,16 @@
font-size: 11.5px;
color: var(--faint);
}
/* Nested inside .dns-filter-copy the note is an ordinary paragraph, but the
endpoint-resolver footnote sits as a DIRECT child of the card — which makes it
a grid item. Without a span it auto-placed into the toggle's `auto` column and
sized that column to its own max-content (322px on desktop, 237px at 390px),
which starved the `1fr` copy column down to 0px: the heading then laid out one
word per line and spilled 2px past the viewport, scrolling the whole page
sideways. It is a full-width footnote under the readout — say so. */
.dns-filter-card > .dns-filter-note {
grid-column: 1 / -1;
}
.dns-readout {
display: flex;
flex-direction: column;
@@ -217,6 +227,12 @@
.dns-input--name {
flex: 0 1 14rem;
}
/* Refresh cadence — a duration, so it needs room for `24h`, not a URL. */
.dns-input--interval {
flex: 0 0 5.5rem;
width: 5.5rem;
text-align: center;
}
.dns-textarea {
width: 100%;
resize: vertical;
@@ -648,73 +664,6 @@
}
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.dns-alert-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.dns-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.dns-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-alert-row delivery controls sit inline in the row (like .dns-detour) */
.dns-alert-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.dns-alert-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.dns-alert-note--row {
margin-top: 2px;
}
@media (max-width: 640px) {
.dns-alert-ctl {
flex-basis: 100%;
order: 3;
}
}
/* alert event checkboxes */
.dns-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.dns-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.dns-event input {
accent-color: var(--fp-accent, currentColor);
}
@media (prefers-reduced-motion: reduce) {
.dns-skel {
animation: none;
@@ -789,12 +738,27 @@
.dns-rule-arrow {
color: var(--groove);
}
/* The per-row editor panel. Same column stack as .dns-add so the shared field
sets (ListFields / ResolverFields / DNSRuleFields) lay out identically whether
they are mounted in the add form or inside a row. */
.dns-rule-edit {
flex-basis: 100%;
width: 100%;
margin-top: 12px;
padding-top: 12px;
border-top: 1px solid var(--groove);
display: flex;
flex-direction: column;
gap: 10px;
}
/* The row's Edit / Delete pair. Pushed to the end so it lines up with the same
pair on a DNS-rule row above it. */
.dns-row-actions {
display: flex;
align-items: center;
gap: 8px;
margin-left: auto;
}
/* A caveat that changes behaviour, not just a note — carries the amber LED. */
+846 -940
View File
File diff suppressed because it is too large Load Diff
+146 -98
View File
@@ -98,7 +98,10 @@
color-mix(in srgb, var(--raised) 82%, var(--panel))
);
box-shadow: 0 1px 0 var(--edge) inset;
overflow: hidden;
/* NOT `overflow: hidden`. The card used to clip its children to the rounded
* corners, which also clipped the policy picker's popover to the bottom of the
* card — the list you were choosing from was cut off mid-row. The two children
* that carry a background round their own corners instead. */
}
.dev-card[data-configured='yes'] {
border-color: color-mix(in srgb, var(--accent) 22%, var(--groove));
@@ -240,73 +243,152 @@
gap: calc(var(--u, 8px) * 2);
padding: calc(var(--u, 8px) * 2) 14px 14px;
border-top: 1px solid var(--groove);
border-radius: 0 0 8px 8px;
background: color-mix(in srgb, var(--sink) 24%, transparent);
}
.dev-controls[data-paused='yes'] {
opacity: 0.72;
}
/* ---- domain (block / allow) editor ---- */
.dev-domains {
border: 1px solid var(--groove);
border-radius: 8px;
padding: 11px 12px;
background: color-mix(in srgb, var(--raised) 50%, transparent);
}
.dev-domains-hd {
/* ---- the order rail ----
*
* Five stops, drawn as a machined strip: this is the order the ENGINE decides a
* name in, and it is the one place on this page where numbering carries real
* information rather than decorating a list. It exists because the two controls
* below it CANNOT show that order by position — the typed lane and the attached
* lane of one control are two steps apart, with the other control's lane in
* between — so the numbers on the rail and the numbers stamped on each lane are
* what tie them together.
*
* A stop is lit only when this device actually has something at that step, which
* makes the rail the card's summary as well as its legend. */
.dev-order {
list-style: none;
margin: 0;
padding: 0;
display: flex;
align-items: baseline;
justify-content: space-between;
flex-wrap: wrap;
align-items: center;
gap: 4px 0;
}
.dev-order-step {
display: inline-flex;
align-items: center;
gap: 6px;
padding: 3px 9px 3px 5px;
border: 1px solid var(--groove);
border-right-width: 0;
background: color-mix(in srgb, var(--sink) 55%, transparent);
}
.dev-order-step:first-child {
border-radius: 6px 0 0 6px;
}
.dev-order-step:last-child {
border-right-width: 1px;
border-radius: 0 6px 6px 0;
}
/* A lit stop: this device has something at that step. Accent is the panel's one
* emphasis colour; it says "in use", not "good". */
.dev-order-step[data-on='yes'] {
border-color: color-mix(in srgb, var(--accent) 45%, var(--groove));
background: color-mix(in srgb, var(--accent) 11%, var(--raised));
}
.dev-order-step[data-on='yes'] + .dev-order-step {
border-left-color: color-mix(in srgb, var(--accent) 45%, var(--groove));
}
.dev-order-n {
display: inline-flex;
align-items: center;
justify-content: center;
min-width: 15px;
height: 15px;
border: 1px solid var(--groove);
border-radius: 3px;
background: var(--sink);
box-shadow: 0 1px 1px var(--shadow) inset;
color: var(--faint);
font-size: 9.5px;
font-weight: 700;
line-height: 1;
}
.dev-order-step[data-on='yes'] .dev-order-n {
border-color: color-mix(in srgb, var(--accent) 55%, var(--groove));
color: var(--accent);
}
.dev-order-lab {
font-size: 10px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--faint);
white-space: nowrap;
}
.dev-order-step[data-on='yes'] .dev-order-lab {
color: var(--ink);
}
.dev-order-cap {
margin: 8px 0 0;
max-width: 72ch;
font-family: var(--font-sans);
font-size: 12px;
line-height: 1.5;
color: var(--dim);
}
.dev-controls[data-paused='yes'] .dev-order-step[data-on='yes'] {
border-color: var(--groove);
background: color-mix(in srgb, var(--sink) 55%, transparent);
}
.dev-controls[data-paused='yes'] .dev-order-step[data-on='yes'] .dev-order-n {
border-color: var(--groove);
color: var(--faint);
}
.dev-controls[data-paused='yes'] .dev-order-step[data-on='yes'] .dev-order-lab {
color: var(--faint);
}
/* ---- the two policy controls ---- */
.dev-pol {
display: flex;
flex-direction: column;
gap: 8px;
}
.dev-pol-row {
display: grid;
grid-template-columns: 58px minmax(0, 1fr);
align-items: start;
gap: 10px;
}
.dev-domains-title {
font-size: 11px;
.dev-pol-lab {
padding-top: 10px;
font-size: 10.5px;
font-weight: 700;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--dim);
}
.dev-domains-count {
font-size: 11px;
color: var(--faint);
.dev-pol-row[data-kind='allow'] .dev-pol-lab {
color: color-mix(in srgb, var(--led-on) 70%, var(--dim));
}
.dev-domains-hint {
margin: 6px 0 10px;
font-family: var(--font-sans);
font-size: 12px;
line-height: 1.45;
color: var(--dim);
max-width: 56ch;
.dev-pol-row[data-kind='block'] .dev-pol-lab {
color: color-mix(in srgb, var(--crit) 60%, var(--dim));
}
.dev-add {
/* Said on the card, not only in the picker: attaching a broad allow list is how
* a device quietly loses the network's filtering, and the person who did it
* should not have to reopen a popover to be told. */
.dev-pol-warn {
display: flex;
align-items: flex-start;
gap: 8px;
}
.dev-input {
flex: 1;
min-width: 0;
padding: 8px 11px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
margin: 0;
max-width: 78ch;
font-family: var(--font-sans);
font-size: 12px;
letter-spacing: 0.02em;
box-shadow: 0 1px 2px var(--shadow) inset;
transition: border-color 0.15s, box-shadow 0.15s;
line-height: 1.5;
color: var(--dim);
}
.dev-input::placeholder {
color: var(--faint);
}
.dev-input:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.dev-input:disabled {
opacity: 0.55;
.dev-pol-warn > .led {
margin-top: 4px;
flex: 0 0 auto;
}
.dev-field-err {
@@ -317,57 +399,6 @@
color: var(--crit);
}
/* chips */
.dev-chips {
list-style: none;
margin: 10px 0 0;
padding: 0;
display: flex;
flex-wrap: wrap;
gap: 7px;
}
.dev-chip {
display: inline-flex;
align-items: center;
gap: 6px;
padding: 3px 4px 3px 9px;
border: 1px solid var(--groove);
border-radius: 6px;
background: var(--sink);
}
.dev-domains[data-kind='allow'] .dev-chip {
border-color: color-mix(in srgb, var(--led-on) 40%, var(--groove));
}
.dev-chip-dom {
font-size: 11.5px;
letter-spacing: 0.01em;
color: var(--ink);
}
.dev-chip-x {
display: inline-flex;
align-items: center;
justify-content: center;
width: 18px;
height: 18px;
padding: 0;
border: 0;
border-radius: 4px;
background: none;
color: var(--faint);
font-size: 11px;
line-height: 1;
cursor: pointer;
transition: color 0.15s, background 0.15s;
}
.dev-chip-x:hover:not(:disabled) {
color: var(--crit);
background: color-mix(in srgb, var(--crit) 14%, transparent);
}
.dev-chip-x:disabled {
opacity: 0.5;
cursor: default;
}
/* ---- card footer ---- */
.dev-card-foot {
display: flex;
@@ -432,6 +463,23 @@
.dev-card-hd {
flex-wrap: wrap;
}
/* The rail wraps rather than scrolls: five stops do not fit a phone, and a
horizontally scrolling legend is a legend nobody reads to the end. Each stop
keeps its own rounded shoulders once the strip is broken up. */
.dev-order-step {
border-right-width: 1px;
border-radius: 6px;
}
.dev-order {
gap: 4px;
}
.dev-pol-row {
grid-template-columns: minmax(0, 1fr);
gap: 4px;
}
.dev-pol-lab {
padding-top: 0;
}
.dev-id {
flex-basis: calc(100% - 100px);
}
+250 -158
View File
@@ -1,9 +1,17 @@
import './Devices.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Module, Toggle, useConfirm } from '../components'
import type { LedVariant } from '../components'
import { apply as apiApply, getConfig, getDevices, putConfig, ApiError } from '../api'
import type { Device, DiscoveredDevice, Model } from '../api'
import { Button, Led, ListPicker, Module, Toggle, useConfirm } from '../components'
import type { LedVariant, ListOption } from '../components'
import {
apply as apiApply,
getConfig,
getDevices,
getRulesetStatus,
putConfig,
ApiError,
} from '../api'
import type { Allowlist, Blocklist, Device, DiscoveredDevice, Model, RulesetStatus } from '../api'
import { listLoad } from '../deviceLists'
// The Devices page is a thin editor over the desired-state Model — exactly like
// Nodes.tsx / DNS.tsx. Liveness (who's online right now) is polled separately
@@ -39,17 +47,67 @@ function matchDevice(devs: Device[], mac: string, ip: string): number {
return -1
}
/** Normalise a typed domain for a per-device block/allow list; null if invalid. */
function cleanDomain(raw: string): string | null {
const d = raw
.trim()
.toLowerCase()
.replace(/^\*\./, '')
.replace(/^\.+/, '')
.replace(/\.+$/, '')
if (!d) return null
if (!/^[a-z0-9.-]+$/.test(d)) return null
return d
// The typed-entry parser lives in ../deviceLists (with its tests): it is the one
// piece of this page that decides whether a parent's Block entry can ever match,
// and `node --test` cannot load a .tsx file.
/** Where a list's content comes from, for the picker row's detail column. */
function listDetail(l: Blocklist | Allowlist): string {
if (l.Source === 'url') {
try {
return `url · ${new URL(l.URL ?? '').host}`
} catch {
return 'url'
}
}
if (l.Source === 'file') return `file · ${l.Path || '—'}`
if (l.Source === 'geosite') {
const cats = (l.Categories ?? []).filter(Boolean)
if (cats.length === 0) return 'geosite'
return cats.length === 1 ? `geosite · ${cats[0]}` : `geosite · ${cats[0]} +${cats.length - 1}`
}
const n = l.Entries?.length ?? 0
return `${n} domain${n === 1 ? '' : 's'}`
}
/** The four list slots on a `config device`, in the order the engine reads them. */
type ListField = 'Allow' | 'Block' | 'Allowlists' | 'Blocklists'
const LIST_VERB: Record<ListField, { add: string; del: string }> = {
Allow: { add: 'Allowed', del: 'Removed allow' },
Block: { add: 'Blocked', del: 'Unblocked' },
Allowlists: { add: 'Attached allowlist', del: 'Detached allowlist' },
Blocklists: { add: 'Attached blocklist', del: 'Detached blocklist' },
}
/** Read one slot. Closed switch — a slot added to the model has to be decided
* about here rather than silently reading `undefined`. */
function readList(d: Device | undefined, field: ListField): string[] {
if (!d) return []
switch (field) {
case 'Allow':
return asArray(d.Allow)
case 'Block':
return asArray(d.Block)
case 'Allowlists':
return asArray(d.Allowlists)
case 'Blocklists':
return asArray(d.Blocklists)
}
}
/** Write one slot, leaving the other three exactly as they were. */
function withList(d: Device, field: ListField, next: string[]): Device {
switch (field) {
case 'Allow':
return { ...d, Allow: next }
case 'Block':
return { ...d, Block: next }
case 'Allowlists':
return { ...d, Allowlists: next }
case 'Blocklists':
return { ...d, Blocklists: next }
}
}
const STATE_LED: Record<string, LedVariant> = {
@@ -118,6 +176,40 @@ export default function Devices() {
return () => window.clearInterval(id)
}, [loadDevices])
// ---- did the attached lists actually LOAD? --------------------------------
//
// The same endpoint and the same tags the DNS page reads (`bl-<name>` /
// `al-<name>`), for the same reason it reads them: attaching a list is not
// evidence that anything is filtered. A url list whose fetch never succeeded, a
// file that is not there, a geosite category that resolved to nothing — each is
// a parental control that does NOT work, and it has to be visible as one on the
// card where the parent attached it, not only on the DNS page. Grouped by NAME
// because a geo list with N categories reports N records. Slow poll: lists
// refresh on a ~24h cadence, so 20s only has to catch a manual Update-now.
const [listStatus, setListStatus] = useState<Map<string, RulesetStatus[]>>(new Map())
const loadListStatus = useCallback(async () => {
try {
const all = await getRulesetStatus()
const m = new Map<string, RulesetStatus[]>()
for (const s of all) {
if (s.kind !== 'blocklist' && s.kind !== 'allowlist') continue
const key = `${s.kind}:${s.name}`
const arr = m.get(key)
if (arr) arr.push(s)
else m.set(key, [s])
}
setListStatus(m)
} catch {
// Engine stopped or an older daemon — keep the last reading. Chips fall back
// to "load unknown", which claims nothing either way.
}
}, [])
useEffect(() => {
void loadListStatus()
const id = window.setInterval(() => void loadListStatus(), 20000)
return () => window.clearInterval(id)
}, [loadListStatus])
// ---- toast + persistent apply banner --------------------------------------
const [toast, setToast] = useState<string | null>(null)
const toastTimer = useRef<number | undefined>(undefined)
@@ -218,6 +310,35 @@ export default function Devices() {
return out
}, [devices, config])
// ---- what a device can attach, and what the engine says about it ----------
const blockOptions = useMemo<ListOption[]>(
() =>
asArray(config?.Blocklists).map((b) => ({
name: b.Name,
detail: listDetail(b),
networkEnabled: b.Enabled,
response: b.Response,
load: listLoad(listStatus.get(`blocklist:${b.Name}`), true),
})),
[config, listStatus],
)
const allowOptions = useMemo<ListOption[]>(
() =>
asArray(config?.Allowlists).map((a) => ({
name: a.Name,
detail: listDetail(a),
networkEnabled: a.Enabled,
load: listLoad(listStatus.get(`allowlist:${a.Name}`), true),
})),
[config, listStatus],
)
// Step 5 of the ladder. "The network filter is in force" is the master switch
// AND at least one blocklist switched on — the switch alone blocks nothing, and
// drawing the last rung lit with no list behind it would be the same lie the
// rest of this page exists to avoid.
const networkFilterOn =
(config?.Globals?.DNSFilter ?? false) && asArray(config?.Blocklists).some((b) => b.Enabled)
const onlineCount = rows.filter((r) => r.state === 'online').length
const idleCount = rows.filter((r) => r.state === 'idle').length
const offlineCount = rows.filter((r) => r.state === 'offline').length
@@ -251,6 +372,27 @@ export default function Devices() {
const nameOf = (row: DeviceRow) => row.cfg?.Name || row.hostname || row.ip || 'device'
/**
* Write one of the four list slots back whole.
*
* The picker emits the complete next list rather than one add/remove, so the
* toast is derived by diffing — which also keeps it truthful when a change
* touches more than one entry. Slots are a CLOSED switch, not a computed key:
* `{ ...d, [field]: next }` would compile against a string index signature and
* happily write a slot that does not exist.
*/
const setDeviceList = useCallback(
(row: DeviceRow, field: ListField, next: string[]) => {
const cur = readList(row.cfg, field)
const added = next.filter((x) => !cur.includes(x))
const removed = cur.filter((x) => !next.includes(x))
const what = added[0] ?? removed[0] ?? ''
const verb = added.length > 0 ? LIST_VERB[field].add : LIST_VERB[field].del
editDevice(row, (d) => withList(d, field, next), `${verb} ${what} — ${nameOf(row)}`)
},
[editDevice],
)
const removeControl = useCallback(
async (row: DeviceRow) => {
if (!config) return
@@ -310,9 +452,10 @@ export default function Devices() {
]}
/>
<p className="dev-summary-note">
Clients appear here as they join the LAN. Manage a device to block or always-allow domains
just for it, or pause its policy. To route a device through a specific exit, add a Routing
rule with this device as the source.
Clients appear here as they join the LAN. Manage a device to give it its own domain
policy — type entries by hand, attach a named list from the DNS page, or pause the lot.
To route a device through a specific exit, add a Routing rule with this device as the
source.
</p>
</div>
@@ -375,18 +518,10 @@ export default function Devices() {
onToggleEnabled={(on) =>
editDevice(row, (d) => ({ ...d, Enabled: on }), `${nameOf(row)} policy ${on ? 'active' : 'paused'}`)
}
onAddBlock={(dom) =>
editDevice(row, (d) => ({ ...d, Block: [...asArray(d.Block), dom] }), `Blocked ${dom} for ${nameOf(row)}`)
}
onRemoveBlock={(dom) =>
editDevice(row, (d) => ({ ...d, Block: asArray(d.Block).filter((x) => x !== dom) }), `Unblocked ${dom}`)
}
onAddAllow={(dom) =>
editDevice(row, (d) => ({ ...d, Allow: [...asArray(d.Allow), dom] }), `Allowed ${dom} for ${nameOf(row)}`)
}
onRemoveAllow={(dom) =>
editDevice(row, (d) => ({ ...d, Allow: asArray(d.Allow).filter((x) => x !== dom) }), `Removed allow ${dom}`)
}
blockOptions={blockOptions}
allowOptions={allowOptions}
networkFilterOn={networkFilterOn}
onSetList={(field, next) => setDeviceList(row, field, next)}
/>
))}
</ul>
@@ -408,28 +543,30 @@ interface DeviceCardProps {
row: DeviceRow
name: string
busy: boolean
/** Every `config blocklist` / `config allowlist`, with the engine's load reading. */
blockOptions: ListOption[]
allowOptions: ListOption[]
/** Step 5: is the network-wide DNS filter actually blocking anything? */
networkFilterOn: boolean
onAddControl: () => void
onRemoveControl: () => void
onRename: (name: string) => void
onToggleEnabled: (on: boolean) => void
onAddBlock: (d: string) => void
onRemoveBlock: (d: string) => void
onAddAllow: (d: string) => void
onRemoveAllow: (d: string) => void
onSetList: (field: ListField, next: string[]) => void
}
function DeviceCard({
row,
name,
busy,
blockOptions,
allowOptions,
networkFilterOn,
onAddControl,
onRemoveControl,
onRename,
onToggleEnabled,
onAddBlock,
onRemoveBlock,
onAddAllow,
onRemoveAllow,
onSetList,
}: DeviceCardProps) {
const cfg = row.cfg
const configured = !!cfg
@@ -437,6 +574,11 @@ function DeviceCard({
const paused = configured && cfg?.Enabled === false
const netLabel = networkLabel(row)
const allowDomains = asArray(cfg?.Allow)
const blockDomains = asArray(cfg?.Block)
const allowLists = asArray(cfg?.Allowlists)
const blockLists = asArray(cfg?.Blocklists)
// ---- inline rename ---------------------------------------------------------
// Names default to the DHCP hostname; this lets the operator pin a friendly one.
// Renaming an unmanaged device upserts it into the managed set (editDevice does
@@ -561,27 +703,77 @@ function DeviceCard({
<p className="dev-unmanaged">Using network defaults</p>
) : (
<div className="dev-controls" data-paused={paused ? 'yes' : 'no'}>
{/* per-device block list */}
<DomainEditor
kind="block"
title="Blocked for this device"
hint="These domains are refused for this device only, on top of the network blocklists."
values={asArray(cfg?.Block)}
busy={busy}
onAdd={onAddBlock}
onRemove={onRemoveBlock}
/>
{/* The whole point of the card: the order the engine decides a name in.
It is drawn, not described, because the two controls under it cannot
show it by position — the typed lane and the attached lane of ONE
control are two steps apart, with the other control's lane between
them. Each rung lights only when this device actually has something
at that step, so the rail doubles as the card's summary. */}
<ol className="dev-order" aria-label={`How a domain is decided for ${name} — first match wins`}>
{[
{ n: 1, label: 'allow · typed', on: allowDomains.length > 0 },
{ n: 2, label: 'block · typed', on: blockDomains.length > 0 },
{ n: 3, label: 'allow · lists', on: allowLists.length > 0 },
{ n: 4, label: 'block · lists', on: blockLists.length > 0 },
{ n: 5, label: 'network filter', on: networkFilterOn },
].map((s) => (
<li key={s.n} className="dev-order-step" data-on={s.on ? 'yes' : 'no'}>
<span className="dev-order-n mono">{s.n}</span>
<span className="dev-order-lab mono">{s.label}</span>
</li>
))}
</ol>
<p className="dev-order-cap">
{paused
? 'Policy is paused, so steps 1–4 are not applied at all — this device falls straight through to the network filter.'
: 'First match wins. Typed entries beat attached lists, and an allow beats a block at the same step.'}
</p>
{/* per-device allow list */}
<DomainEditor
kind="allow"
title="Always allowed for this device"
hint="These domains always resolve for this device, even if a blocklist would block them."
values={asArray(cfg?.Allow)}
busy={busy}
onAdd={onAddAllow}
onRemove={onRemoveAllow}
/>
<div className="dev-pol">
<div className="dev-pol-row" data-kind="allow">
<span className="dev-pol-lab mono">Allow</span>
<ListPicker
kind="allow"
lists={allowLists}
domains={allowDomains}
options={allowOptions}
typedStep={1}
listStep={3}
disabled={busy}
ariaLabel={`Always allowed for ${name}`}
onListsChange={(v) => onSetList('Allowlists', v)}
onDomainsChange={(v) => onSetList('Allow', v)}
/>
</div>
<div className="dev-pol-row" data-kind="block">
<span className="dev-pol-lab mono">Block</span>
<ListPicker
kind="block"
lists={blockLists}
domains={blockDomains}
options={blockOptions}
typedStep={2}
listStep={4}
disabled={busy}
ariaLabel={`Blocked for ${name}`}
onListsChange={(v) => onSetList('Blocklists', v)}
onDomainsChange={(v) => onSetList('Block', v)}
/>
</div>
</div>
{/* Not shown while the policy is paused: with steps 1-4 out of force the
sentence would describe filtering this device is not losing. */}
{allowLists.length > 0 && !paused && (
<p className="dev-pol-warn" role="note">
<Led variant="amber" />
<span>
An attached allow list is terminal: everything{' '}
{allowLists.length === 1 ? allowLists[0] : `these ${allowLists.length} lists`} covers
also stops being filtered by the network blocklists for {name}.
</span>
</p>
)}
<div className="dev-card-foot">
<Button className="dev-unmanage" onClick={onRemoveControl} disabled={busy}>
@@ -593,103 +785,3 @@ function DeviceCard({
</li>
)
}
// ---- domain (block / allow) editor -----------------------------------------
function DomainEditor({
kind,
title,
hint,
values,
busy,
onAdd,
onRemove,
}: {
kind: 'block' | 'allow'
title: string
hint: string
values: string[]
busy: boolean
onAdd: (d: string) => void
onRemove: (d: string) => void
}) {
const [text, setText] = useState('')
const [err, setErr] = useState<string | null>(null)
const submit = () => {
const dom = cleanDomain(text)
if (!dom) {
setErr('Enter a domain like example.com.')
return
}
if (values.some((v) => v.toLowerCase() === dom)) {
setErr(`${dom} is already ${kind === 'block' ? 'blocked' : 'allowed'}.`)
return
}
setErr(null)
setText('')
onAdd(dom)
}
return (
<div className="dev-domains" data-kind={kind}>
<div className="dev-domains-hd">
<span className="dev-domains-title mono">{title}</span>
<span className="dev-domains-count mono">{values.length}</span>
</div>
<p className="dev-domains-hint">{hint}</p>
<form
className="dev-add"
onSubmit={(e) => {
e.preventDefault()
submit()
}}
>
<input
className="dev-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder={kind === 'block' ? 'example.com' : 'school.example.edu'}
aria-label={`Domain to ${kind === 'block' ? 'block' : 'allow'}`}
value={text}
onChange={(e) => {
setText(e.target.value)
if (err) setErr(null)
}}
disabled={busy}
/>
<Button type="submit" disabled={busy}>
{kind === 'block' ? 'Block' : 'Allow'}
</Button>
</form>
{err && (
<p className="dev-field-err" role="alert">
{err}
</p>
)}
{values.length > 0 && (
<ul className="dev-chips">
{values.map((d) => (
<li key={d} className="dev-chip">
<span className="dev-chip-dom mono">{d}</span>
<button
type="button"
className="dev-chip-x"
onClick={() => onRemove(d)}
disabled={busy}
aria-label={`${kind === 'block' ? 'Unblock' : 'Remove allow for'} ${d}`}
title="Remove"
>
✕
</button>
</li>
))}
</ul>
)}
</div>
)
}
+399 -2
View File
@@ -797,9 +797,186 @@
.ins-conn--dns {
grid-template-columns: 62px minmax(60px, 116px) 12px minmax(0, 1fr) minmax(72px, 160px) auto;
}
.ins-conn--dns .tag {
/* The verdict cell: PATH then OUTCOME, in that order, so the eye lands last on
* what came of the lookup. Wraps rather than squeezes — on a narrow screen the
* outcome drops under the path instead of pushing the domain to nothing. */
.ins-dns-v {
display: flex;
align-items: center;
justify-content: flex-end;
flex-wrap: wrap;
gap: 4px 6px;
justify-self: end;
}
/* ---- the four situations one DNS row can be in ------------------------------
* answered · blocked · failed · not recorded. No two of them may look alike, and
* the unrecorded one may not look like any of the three answers.
*
* The OUTCOME is what marks the row (a rail down its start edge) because the
* outcome is what the reader is sorting rows into. The PATH keeps its own colour
* beside it — a failure in the tunnel and a failure on the direct path are
* different reports — but it is muted below whenever the outcome is not a plain
* answer, so no failed row can ever again be the healthiest-looking line here. */
.ins-dns-st {
flex: none;
padding: 1px 5px;
border: 1px solid var(--groove);
border-radius: 4px;
font-family: var(--font-mono);
font-size: 9.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
white-space: nowrap;
color: var(--dim);
cursor: help;
}
/* an answer came back — stated, quietly. It is the common case and must not shout */
.ins-dns-st.s-answered {
color: var(--led-on);
border-color: color-mix(in srgb, var(--led-on) 40%, var(--groove));
}
/* NO usable answer. Amber, filled, and it takes the row's rail with it — crit is
* already spoken for by the filter's deliberate block, and a fault the box did
* not intend is a different thing from a kill it did. */
.ins-dns-st.s-failed {
color: var(--amber);
border-color: color-mix(in srgb, var(--amber) 60%, var(--groove));
background: color-mix(in srgb, var(--amber) 16%, transparent);
font-weight: 700;
}
/* NOT RECORDED — dashed and faint, the same language the routing chips use for
* the state that is not an answer at all. */
.ins-dns-st.s-unrecorded {
color: var(--faint);
border-style: dashed;
background: none;
}
/* The row rail. `answered` gets none — a clean row is the baseline everything
* else is read against. */
.ins-conn--dns.st-blocked,
.ins-conn--dns.st-failed,
.ins-conn--dns.st-unrecorded {
border-inline-start: 2px solid transparent;
padding-inline-start: 10px;
}
.ins-conn--dns.st-blocked {
border-inline-start-color: color-mix(in srgb, var(--crit) 70%, transparent);
}
.ins-conn--dns.st-failed {
border-inline-start-color: var(--amber);
background: color-mix(in srgb, var(--amber) 7%, transparent);
}
.ins-conn--dns.st-unrecorded {
border-inline-start-style: dashed;
border-inline-start-color: color-mix(in srgb, var(--faint) 60%, transparent);
}
/* ---- the same axis on a CONNECTION row --------------------------------------
* killed · carried · no exit recorded. A connection sent to the engine's `block`
* outbound was not carried at all — a rule with target Block, or the kill-switch
* default — and it has to be as findable here as the DNS block is above it. Same
* rail, same crit, because it is the same statement about the same box.
*
* `carried` takes no rail: it is the baseline the other two are read against, and
* it is deliberately NOT green — the log records the routing decision, not
* whether the flow then worked, so there is no health to claim. */
.ins-conn.st-killed,
.ins-conn.st-unrecorded {
border-inline-start: 2px solid transparent;
padding-inline-start: 10px;
}
.ins-conn.st-killed {
border-inline-start-color: color-mix(in srgb, var(--crit) 70%, transparent);
background: color-mix(in srgb, var(--crit) 6%, transparent);
}
.ins-conn.st-unrecorded {
border-inline-start-style: dashed;
border-inline-start-color: color-mix(in srgb, var(--faint) 60%, transparent);
}
/* The fate mark. Same chrome as the DNS outcome chip (.ins-dns-st) so the two
* logs answer "what came of it" in one visual language. */
.ins-fate {
flex: none;
padding: 1px 5px;
border: 1px solid var(--groove);
border-radius: 4px;
font-family: var(--font-mono);
font-size: 9.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
white-space: nowrap;
color: var(--dim);
cursor: help;
}
/* nothing left the router for this connection */
.ins-fate.f-killed {
color: var(--crit);
border-color: color-mix(in srgb, var(--crit) 60%, var(--groove));
background: color-mix(in srgb, var(--crit) 15%, transparent);
font-weight: 700;
}
/* it left through a named outbound — stated, quietly. The common case, and it
* must not shout and must not read as "it worked". */
.ins-fate.f-carried {
color: var(--dim);
}
/* nothing was recorded about the exit. Dashed and faint — the same language the
* routing chips use for the state that is not an answer at all. */
.ins-fate.f-unrecorded {
color: var(--faint);
border-style: dashed;
background: none;
}
/* The exit tag itself, in the column that a narrow screen drops. It carries the
* kill's colour so the two readings agree wherever both are visible; it is never
* the only place the kill is said. */
.ins-conn-out.e-killed {
color: var(--crit);
}
.ins-conn-out.e-unrecorded {
color: var(--faint);
}
.ins-conn--dns .tag {
flex: none;
}
/* The path is still the path — but it stops carrying a health signal the moment
* the outcome says the lookup did not produce one. This is the whole defect:
* `action=proxy` + `status=failed` used to render as a bright, healthy tunnel. */
.ins-conn--dns.st-failed .tag,
.ins-conn--dns.st-unrecorded .tag {
color: var(--faint);
border-color: color-mix(in srgb, var(--faint) 45%, transparent);
background: none;
}
/* A path nobody recorded, in either outcome. Dashed, so it cannot pass for one
* of the three real paths. */
.ins-conn--dns .tag.unknown {
color: var(--faint);
border-style: dashed;
border-color: color-mix(in srgb, var(--faint) 55%, transparent);
background: none;
}
/* the failure's own cause, verbatim from the resolver, and what the rcode adds */
.ins-why-cause {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
color: var(--amber);
cursor: help;
}
.ins-why-rcode {
flex: none;
color: var(--faint);
cursor: help;
}
@media (max-width: 560px) {
/* resolver (rendered via .ins-conn-out) is hidden by the .ins-conn-out rule
* above; collapse to time · device · → · domain · tag so the source device, the
@@ -809,10 +986,30 @@
/* trimmed device/domain mins so time + device + domain + tag fit the narrow
* scroll box (no inner h-scroll); domain keeps a small hard min so the queried
* name never fully collapses. */
grid-template-columns: 48px minmax(34px, 64px) 8px minmax(30px, 1fr) auto;
grid-template-columns: 48px minmax(34px, 64px) 8px minmax(30px, 1fr);
column-gap: 6px;
padding-inline: 8px;
}
/* The verdict is TWO chips now, and a ~86px column for them left the queried
* domain pinned at its 30px minimum — the one field nobody can do without,
* truncated to three letters. So on a narrow screen the verdict takes its own
* full-width line under the query instead of competing with it: the domain gets
* the whole remainder of line one, and outcome + path stay side by side and
* right-aligned, in the same reading order as on a wide screen. */
.ins-dns-v {
grid-column: 1 / -1;
justify-content: flex-end;
margin-top: 2px;
}
/* the outcome rail eats 2px of the 8px inset, so the text still lines up with
* the rows that have no rail */
.ins-conn--dns.st-blocked,
.ins-conn--dns.st-failed,
.ins-conn--dns.st-unrecorded,
.ins-conn.st-killed,
.ins-conn.st-unrecorded {
padding-inline-start: 6px;
}
}
/* ---- logging-off state ---- *
@@ -905,3 +1102,203 @@
animation: none;
}
}
/* ---- why it went there / where it went out ----------------------------------
* A second line under each log row carrying the ROUTING RECORD: for a connection
* the rule that matched, for a DNS query the channel the lookup left through.
* It spans the whole row grid so the columns above stay unchanged (and keep
* working when the narrow breakpoint drops one of them). */
.ins-why {
grid-column: 1 / -1;
display: flex;
align-items: baseline;
flex-wrap: wrap;
gap: 4px 7px;
margin-top: 3px;
min-width: 0;
font-size: 10.5px;
line-height: 1.35;
}
.ins-why-kind {
flex: none;
padding: 1px 5px;
border: 1px solid var(--groove);
border-radius: 4px;
font-family: var(--font-mono);
font-size: 9.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
color: var(--dim);
background: color-mix(in srgb, var(--raised) 60%, transparent);
cursor: help;
}
.ins-why-txt,
.ins-why-path {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
color: var(--faint);
}
.ins-why-path::before {
content: 'path ';
color: var(--groove);
}
/* The three/four vocabularies. Each state gets its OWN look — that is the whole
* point of the fields (stats.go: "" is reserved for NOT RECORDED, so it can never
* be read as "no rule matched" / "nothing egressed"). Do not merge any two. */
/* a rule matched — a plain recorded answer */
.ins-why-kind.k-matched {
color: var(--ink);
border-color: color-mix(in srgb, var(--accent) 40%, var(--groove));
}
/* nothing matched, so the default route carried it — an ANSWER, stated quietly */
.ins-why-kind.k-default {
color: var(--dim);
border-style: solid;
}
/* the resolver names no detour: this lookup left over the plain WAN, past the
* tunnel. Marked, not alarmed — a box with no anti-leak intent reads this all day */
.ins-why-kind.o-default {
color: var(--amber);
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
background: color-mix(in srgb, var(--amber) 12%, transparent);
}
/* answered on the box — nothing egressed at all */
.ins-why-kind.o-local {
color: var(--dim);
}
/* the lookup left through a named detour */
.ins-why-kind.o-detour {
color: var(--ink);
border-color: color-mix(in srgb, var(--accent) 40%, var(--groove));
}
/* NOT RECORDED — the only state that is not a fact about routing. Dashed and
* faint so it can never be mistaken for one of the answers above. */
.ins-why-kind.k-unrecorded,
.ins-why-kind.o-unrecorded {
color: var(--faint);
border-style: dashed;
background: none;
}
/* ---- the scan stopped before the log did ------------------------------------
* A filtered walk is bounded by the daemon; a page that ended on that bound is
* short for a reason that has nothing to do with how much data exists. Naming it
* is the whole job — with the resume control right next to the sentence. */
.ins-scan {
display: flex;
align-items: baseline;
flex-wrap: wrap;
gap: 6px 9px;
margin: 0 0 10px;
padding: 9px 11px;
border: 1px solid color-mix(in srgb, var(--amber) 50%, var(--groove));
border-radius: 8px;
background: color-mix(in srgb, var(--amber) 10%, transparent);
font-size: 11.5px;
line-height: 1.45;
color: var(--dim);
}
.ins-scan-lab {
flex: none;
font-family: var(--font-mono);
font-size: 9.5px;
font-weight: 700;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--amber);
}
.ins-scan-txt {
flex: 1 1 220px;
min-width: 0;
}
.ins-scan-go {
flex: none;
padding: 4px 9px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 6px;
background: var(--raised);
color: var(--ink);
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.04em;
cursor: pointer;
}
.ins-scan-go:hover:not(:disabled) {
border-color: var(--amber);
}
.ins-scan-go:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.ins-scan-go:disabled {
opacity: 0.55;
cursor: default;
}
/* ---- log search -------------------------------------------------------------
* The daemon filters, not the panel. The hint under the box names the fields that
* ARE searched, so an empty result reads as "not in these fields" rather than
* "did not happen". */
.ins-search {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 3px 6px;
min-width: 0;
max-width: 100%;
}
.ins-search-in {
flex: 1 1 150px;
min-width: 0;
max-width: 100%;
padding: 5px 8px;
border: 1px solid var(--groove);
border-radius: 6px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 11px;
box-shadow: 0 1px 2px var(--shadow) inset;
}
.ins-search-in::placeholder {
color: var(--faint);
}
.ins-search-in:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.ins-search-clr {
flex: none;
padding: 3px 6px;
border: 1px solid var(--groove);
border-radius: 6px;
background: var(--raised);
color: var(--dim);
font-size: 10px;
line-height: 1;
cursor: pointer;
}
.ins-search-clr:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.ins-search-fields {
flex: 1 0 100%;
font-size: 9.5px;
letter-spacing: 0.03em;
color: var(--faint);
text-align: right;
}
@media (max-width: 560px) {
.ins-search {
flex: 1 1 100%;
}
.ins-search-fields {
text-align: left;
}
}
+438 -78
View File
@@ -1,11 +1,25 @@
import './Insights.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, SegMeter } from '../components'
import type { LedVariant, QueryTag } from '../components'
import type { LedVariant } from '../components'
import { usePrefersReducedMotion } from '../components/usePrefersReducedMotion'
import { fmtBytes, fmtClock } from '../format'
import { getStats, getStatsConnsPage, getStatsLogPage } from '../api'
import type { StatsLogPage, StatsLogQuery } from '../api'
import { serviceIntent } from '../planeState'
import {
CONN_SEARCH_HINT,
LOG_SEARCH_HINT,
chainPath,
connFate,
connRule,
dnsOutbound,
dnsRowMark,
historyExhausted,
logCountLabel,
logEndNote,
resumeCursor,
} from '../logRoute'
import type {
ConnLogEntry,
DeviceDomains,
@@ -16,6 +30,7 @@ import type {
RuleStat,
ServerStat,
Stats,
Status,
TimelineBucket,
} from '../api'
@@ -32,12 +47,6 @@ function fmtNum(n: number): string {
return (n || 0).toLocaleString('en-US')
}
function actionTag(action?: string): QueryTag {
if (action === 'block') return 'block'
if (action === 'proxy') return 'proxy'
return 'pass'
}
// ---- small presentational pieces -------------------------------------------
type Tone = 'accent' | 'flow' | 'crit' | 'good'
@@ -148,12 +157,31 @@ function HostRow({ h, max }: { h: HostStat; max: number }) {
* One connection event: which LAN device (name, else IP) reached which
* destination (:port), with the transport/proto chip and the exit it took. This
* is the "from which device and where" the DNS query log structurally can't show.
*
* TWO AXES, AS ON THE DNS ROW ABOVE. `rule_kind` answers WHY the connection went
* where it went; it does not answer whether it went anywhere. A row whose exit is
* the engine's `block` outbound was NOT carried — that is a rule with target
* Block, or the kill-switch default — and it used to be drawn as plain mono text,
* indistinguishable from `nl-reality-1`, while a DNS row about the same host on
* this very page got a crit rail and a BLOCK mark. This log is where a "the site
* does not open" report is read first, so the row that IS the answer now carries
* its own mark (logRoute.connFate) and its own rail.
*/
function ConnRow({ c, fresh }: { c: ConnLogEntry; fresh: boolean }) {
const dev = c.src_name || c.src_ip
const dest = c.port ? `${c.dest}:${c.port}` : c.dest
// WHY it went there. The three states of `rule_kind` are decided once, in
// logRoute.connRule, and they must stay visibly different here: "no rule
// matched" is an answer, "not recorded" is the absence of one.
const why = connRule(c)
// WHETHER it went at all. Decided once, in logRoute.connFate — the page derives
// nothing about `block` on its own.
const fate = connFate(c)
const path = chainPath(c.chain)
const cls = ['ins-conn', `st-${fate.fate}`]
if (fresh) cls.push('new')
return (
<li className={fresh ? 'ins-conn new' : 'ins-conn'}>
<li className={cls.join(' ')}>
<span className="ins-conn-t mono">{fmtClock(c.unix)}</span>
<span className="ins-conn-dev mono" title={c.src_ip}>
{dev}
@@ -165,27 +193,64 @@ function ConnRow({ c, fresh }: { c: ConnLogEntry; fresh: boolean }) {
{dest}
</span>
<NetBadge network={c.network} proto={c.proto} />
<span className="ins-conn-out mono" title={`exit · ${c.outbound}`}>
{/* The exit column is the first thing dropped on a narrow screen, so it may
never be the only place the kill is stated — the mark and the rail below
survive at every width. */}
<span className={`ins-conn-out mono e-${fate.fate}`} title={fate.title}>
{c.outbound || '—'}
</span>
<span className="ins-why">
<span className={`ins-fate f-${fate.fate}`} title={fate.title}>
{fate.label}
</span>
<span className={`ins-why-kind k-${why.kind}`} title={why.title}>
{why.label}
</span>
{why.text ? (
<span className="ins-why-txt mono" title={why.title}>
{why.text}
</span>
) : null}
{path ? (
<span className="ins-why-path mono" title={`outbound path · ${path}`}>
{path}
</span>
) : null}
</span>
</li>
)
}
/**
* One DNS decision: time · device → domain · resolver · action tag
* (block/proxy/pass). Mirrors ConnRow's `device → dest` reading (same `.ins-conn`
* chrome + accent arrow) so the DNS log and the Connections log read identically:
* the "who asked" is the LAN device, the resolver that answered is the trailing
* secondary column. The backend drops the router's own resolutions (urltest probes /
* sub fetches) at ingestion, so every row here is a real client; an empty device
* (older data) falls back to a dash.
* One DNS decision: time · device → domain · resolver · path tag · OUTCOME.
* Mirrors ConnRow's `device → dest` reading (same `.ins-conn` chrome + accent
* arrow) so the DNS log and the Connections log read identically: the "who asked"
* is the LAN device, the resolver that answered is the trailing secondary column.
* The backend drops the router's own resolutions (urltest probes / sub fetches) at
* ingestion, so every row here is a real client; an empty device (older data)
* falls back to a dash.
*
* TWO AXES, AND THE OUTCOME LEADS. The row used to be coloured by `action` alone,
* which answers "which way did it go" — so a lookup that left through a detour and
* then TIMED OUT came out as an accent-coloured `proxy`, the single healthiest-
* looking row in the log, describing the exact moment the tunnel failed. The
* outcome now marks the whole row (`st-<state>`, four states, four looks) and the
* path stays right beside it, because "it failed" and "it failed in the tunnel"
* are different reports and the second one is the useful one.
*/
function DnsRow({ r, fresh }: { r: QueryLogEntry; fresh: boolean }) {
const tag = actionTag(r.action)
// Outcome + path in one reading (logRoute.dnsRowMark) — the page derives
// neither on its own, so the four states cannot collapse into three here.
const mark = dnsRowMark(r)
const dev = r.device.trim()
// WHERE the lookup left the box. Four states, four readings (logRoute
// .dnsOutbound) — `default` is the one worth spotting: the resolver names no
// detour, so this query went out past the tunnel.
const out = dnsOutbound(r)
const cls = ['ins-conn', 'ins-conn--dns', `st-${mark.state}`]
if (fresh) cls.push('new')
return (
<li className={fresh ? 'ins-conn ins-conn--dns new' : 'ins-conn ins-conn--dns'}>
<li className={cls.join(' ')}>
<span className="ins-conn-t mono">{fmtClock(r.unix)}</span>
<span className="ins-conn-dev mono" title={dev || 'source device unknown'}>
{dev || '—'}
@@ -199,7 +264,43 @@ function DnsRow({ r, fresh }: { r: QueryLogEntry; fresh: boolean }) {
<span className="ins-conn-out mono" title={`resolver · ${r.server}`}>
{r.server || '—'}
</span>
<span className={`tag ${tag}`}>{tag}</span>
<span className="ins-dns-v">
<span className={`tag ${mark.path}`} title={mark.pathTitle}>
{mark.pathLabel}
</span>
<span className={`ins-dns-st s-${mark.outcome}`} title={mark.title}>
{mark.label}
</span>
</span>
<span className="ins-why">
<span className={`ins-why-kind o-${out.kind}`} title={out.title}>
{out.label}
</span>
{out.tag ? (
<span className="ins-why-txt mono" title={out.title}>
{out.tag}
</span>
) : null}
{/* Only a failure has a cause, and a failure ALWAYS shows one — "cause not
recorded" included, because an empty cell reads as nothing wrong. */}
{mark.cause ? (
<span className="ins-why-cause mono" title={mark.title}>
{mark.cause}
</span>
) : null}
{mark.rcodeNote ? (
<span
className="ins-why-rcode mono"
title={
r.rcode === -1
? 'The server never answered at all — a timeout, a network error, or a lookup that never left.'
: `The server answered with response code ${r.rcode} (2 SERVFAIL, 5 REFUSED).`
}
>
{mark.rcodeNote}
</span>
) : null}
</span>
</li>
)
}
@@ -415,6 +516,7 @@ function LogShell({
pending,
onTogglePause,
onFlush,
truncated,
}: {
ariaLabel: string
children: React.ReactNode
@@ -427,9 +529,12 @@ function LogShell({
pending: number
onTogglePause: () => void
onFlush: () => void
/** The daemon stopped scanning before the data ran out — see ScanNotice. */
truncated: boolean
}) {
return (
<>
{truncated ? <ScanNotice onContinue={onMore} busy={busy} /> : null}
{/* aria-live off while paused so a screen reader isn't nagged by rows the
* user deliberately froze; polite while live so new rows are announced. */}
<div
@@ -481,6 +586,106 @@ function LogShell({
)
}
/**
* THE SEARCH STOPPED BEFORE THE LOG DID.
*
* A filtered walk has to examine rows it will not return, so the daemon bounds
* it and reports X-Stats-Log-Truncated when that bound — not the data — ended
* the walk. Drawing such a page as the end of the log is the exact lie this
* notice exists to prevent: the operator would read "nothing found" and stop,
* with the evidence still sitting further back in the ring.
*
* So the state is NAMED, and it comes with the only thing that can resolve it:
* the daemon's resume cursor, spent by the same "load older" path the button
* below drives. It is a note, not an alarm — nothing is broken.
*/
function ScanNotice({ onContinue, busy }: { onContinue: () => void; busy: boolean }) {
return (
<p className="ins-scan" role="status">
<span className="ins-scan-lab">Scan stopped</span>
<span className="ins-scan-txt">
The daemon limits how many rows one search may scan, and this search hit that limit.
These are the matches found so far — not the end of the log.
</span>
<button type="button" className="ins-scan-go" onClick={onContinue} disabled={busy}>
{busy ? 'Searching…' : 'Keep searching further back'}
</button>
</p>
)
}
/**
* The log search box.
*
* The daemon filters, not the panel: `q=` is applied inside the store on the
* same walk as the seq cursor, so `limit` counts MATCHING rows and paging stays
* correct. It matters because the ring is 200 rows by default — by the time
* anyone opens the panel the evidence is usually gone, and search is what makes
* the log worth keeping.
*
* `fields` names what is searched, and by omission what is not. That is not
* decoration: a person who types a port number and gets nothing deserves to see
* that ports were never in the list, rather than conclude the traffic did not
* happen.
*/
function LogSearch({
id,
value,
onChange,
fields,
placeholder,
}: {
id: string
value: string
onChange: (v: string) => void
fields: string
placeholder: string
}) {
return (
<div className="ins-search">
<input
id={id}
type="search"
className="ins-search-in mono"
value={value}
placeholder={placeholder}
aria-label={`Search this log by ${fields}`}
aria-describedby={`${id}-fields`}
onChange={(e) => onChange(e.target.value)}
/>
{value ? (
<button
type="button"
className="ins-search-clr"
onClick={() => onChange('')}
aria-label="Clear the search"
title="Clear the search"
>
✕
</button>
) : null}
<span id={`${id}-fields`} className="ins-search-fields mono">
{fields}
</span>
</div>
)
}
/**
* Hold a value still for `ms` after the last change. The search text drives a
* server round-trip AND discards the accumulated rows, so sending one per
* keystroke would both hammer the daemon's scan budget and make the list flicker
* through the prefixes of what is being typed.
*/
function useDebounced<T>(value: T, ms: number): T {
const [held, setHeld] = useState(value)
useEffect(() => {
const id = window.setTimeout(() => setHeld(value), ms)
return () => window.clearTimeout(id)
}, [value, ms])
return held
}
// ---- live, cursor-paginated log model --------------------------------------
type HasSeq = { seq: number }
@@ -519,6 +724,13 @@ interface LiveLog<T> {
paused: boolean // live stream paused — new rows buffer instead of prepend
pending: number // rows waiting in the paused buffer ("N new")
freshSeqs: Set<number> // seqs from the most recent live prepend (animate these)
/**
* The last HISTORY page (head or `before=`) ended on the daemon's scan budget
* rather than on the data. Never true without a filter — an unfiltered walk has
* no budget — and it is the one thing that must not be folded into "no more
* rows": see ScanNotice, and logRoute.historyExhausted for the decision it feeds.
*/
truncated: boolean
reload: () => void // re-fetch the newest page (reset to the live head)
loadMore: () => void // append the next older page to the bottom
togglePause: () => void // pause ⇄ resume (resume flushes the buffer)
@@ -543,6 +755,10 @@ function useLiveLog<T extends HasSeq>(
/** Changing this discards the accumulated rows and re-reads from the head.
* Driven by the stats backend id — see the reset effect below (F6). */
resetKey?: string,
/** Server-side substring filter (`q=`). Changing it invalidates every row we
* hold — they were selected under the old needle — so it is folded into the
* reset key below rather than applied to the next page only. */
filter = '',
opts?: { page?: number; livePage?: number; cap?: number; interval?: number },
): LiveLog<T> {
const page = opts?.page ?? 100
@@ -558,6 +774,18 @@ function useLiveLog<T extends HasSeq>(
const [paused, setPaused] = useState(false)
const [buffer, setBuffer] = useState<T[]>([])
const [freshSeqs, setFreshSeqs] = useState<Set<number>>(() => new Set())
const [truncated, setTruncated] = useState(false)
// The daemon's resume point for the HISTORY direction. It is authoritative:
// on a complete page it equals the oldest row's seq, and on a truncated page
// with no matches it is the ONLY way back — there is no row to page from.
const scanCursorRef = useRef(0)
// The FORWARD (live tail) cursor. It used to be derived from rows[0].seq, which
// breaks the moment a filter is on: a filtered `after=` walk that matches
// nothing returns zero rows, rows[0] never moves, and the poller rescans the
// same window every tick forever. The daemon reports the last row it examined,
// so the tail advances on rows examined rather than rows matched.
const liveCursorRef = useRef(0)
// Live values for the stable poll effect, so it never needs to re-subscribe
// (and reset its interval) when rows/paused/buffer change.
@@ -573,12 +801,16 @@ function useLiveLog<T extends HasSeq>(
const reload = useCallback(async () => {
setBusy(true)
try {
const { rows: res } = await fetcher({ limit: page })
const sorted = [...res].sort((a, b) => b.seq - a.seq)
const p = await fetcher({ limit: page, q: filter })
const sorted = [...p.rows].sort((a, b) => b.seq - a.seq)
setRows(sorted)
setBuffer([])
setFreshSeqs(new Set())
setExhausted(res.length < page)
setTruncated(p.truncated)
scanCursorRef.current = p.cursor
liveCursorRef.current = sorted[0]?.seq ?? 0
// A short page is the end of the log ONLY if the scan reached it.
setExhausted(historyExhausted({ returned: p.rows.length, page, truncated: p.truncated }))
setErr(false)
setLoaded(true)
} catch {
@@ -586,18 +818,24 @@ function useLiveLog<T extends HasSeq>(
} finally {
setBusy(false)
}
}, [fetcher, page])
}, [fetcher, page, filter])
const loadMore = useCallback(async () => {
const cur = rowsRef.current
if (busyRef.current || cur.length === 0) return
if (busyRef.current) return
// Not `rows[last].seq`: a truncated page can be EMPTY and still have more log
// behind it, and only the daemon's cursor knows where that is.
const from = resumeCursor({ cursor: scanCursorRef.current, rows: rowsRef.current })
if (from == null) return
setBusy(true)
try {
const { rows: res } = await fetcher({ before: cur[cur.length - 1].seq, limit: page })
if (res.length < page) setExhausted(true)
if (res.length > 0) {
const p = await fetcher({ before: from, limit: page, q: filter })
setTruncated(p.truncated)
if (p.cursor > 0) scanCursorRef.current = p.cursor
else if (p.rows.length > 0) scanCursorRef.current = p.rows[p.rows.length - 1].seq
setExhausted(historyExhausted({ returned: p.rows.length, page, truncated: p.truncated }))
if (p.rows.length > 0) {
setFreshSeqs(new Set()) // an older page never animates
setRows((prev) => mergeSeqDesc(prev, res, cap))
setRows((prev) => mergeSeqDesc(prev, p.rows, cap))
}
setErr(false)
} catch {
@@ -605,7 +843,7 @@ function useLiveLog<T extends HasSeq>(
} finally {
setBusy(false)
}
}, [fetcher, page, cap])
}, [fetcher, page, cap, filter])
const flushPending = useCallback(() => {
const buf = bufferRef.current
@@ -663,48 +901,60 @@ function useLiveLog<T extends HasSeq>(
const poll = async () => {
if (cancelled || document.hidden || busyRef.current) return
const cold = rowsRef.current.length === 0
const cold = rowsRef.current.length === 0 && liveCursorRef.current === 0
try {
if (cold) {
const { rows: res } = await fetcher({ limit: page })
const p = await fetcher({ limit: page, q: filter })
if (cancelled) return
setExhausted(res.length < page)
if (res.length === 0) return
setTruncated(p.truncated)
scanCursorRef.current = p.cursor
setExhausted(historyExhausted({ returned: p.rows.length, page, truncated: p.truncated }))
if (p.rows.length === 0) return
liveCursorRef.current = p.rows[0].seq
// A cold fill is a whole first page, not an arrival — land it quietly,
// exactly as reload() does, instead of animating 100 rows at once.
setFreshSeqs(new Set())
setRows((prev) => mergeSeqDesc(prev, res, cap))
setRows((prev) => mergeSeqDesc(prev, p.rows, cap))
return
}
for (let i = 0; i < MAX_CATCHUP_PAGES; i++) {
const cursor = rowsRef.current[0]?.seq
if (cursor == null) return
const { rows: res, pending, more } = await fetcher({ after: cursor, limit: livePage })
const cursor = liveCursorRef.current || rowsRef.current[0]?.seq
if (!cursor) return
const p = await fetcher({ after: cursor, limit: livePage, q: filter })
if (cancelled) return
if (res.length === 0) return
// Advance on rows EXAMINED, not rows matched: with a filter on, a walk
// that matched nothing still consumed the window, and a cursor that only
// moved on matches would rescan it every tick for as long as the box is up.
liveCursorRef.current = Math.max(cursor, p.cursor, p.rows[0]?.seq ?? 0)
if (p.rows.length === 0) {
if (!p.more) return
continue
}
// Backlog beyond what we can even hold: stop walking it and re-head.
if (pending > cap) {
const { rows: head } = await fetcher({ limit: page })
if (p.pending > cap) {
const head = await fetcher({ limit: page, q: filter })
if (cancelled) return
setBuffer([])
setFreshSeqs(new Set())
setRows(head)
setExhausted(head.length < page)
setRows(head.rows)
setTruncated(head.truncated)
scanCursorRef.current = head.cursor
liveCursorRef.current = head.rows[0]?.seq ?? liveCursorRef.current
setExhausted(
historyExhausted({ returned: head.rows.length, page, truncated: head.truncated }),
)
return
}
if (pausedRef.current) {
setBuffer((prev) => mergeSeqDesc(prev, res, cap))
setBuffer((prev) => mergeSeqDesc(prev, p.rows, cap))
} else {
setFreshSeqs(new Set(res.map((r) => r.seq)))
setRows((prev) => mergeSeqDesc(prev, res, cap))
setFreshSeqs(new Set(p.rows.map((r) => r.seq)))
setRows((prev) => mergeSeqDesc(prev, p.rows, cap))
}
if (!more) return
// Paused: rows land in the buffer, so rowsRef never advances and the
// cursor would repeat. Let the next tick continue instead.
if (pausedRef.current) return
if (!p.more) return
}
} catch {
// transient — keep the last good list and retry on the next tick
@@ -716,7 +966,7 @@ function useLiveLog<T extends HasSeq>(
cancelled = true
window.clearInterval(id)
}
}, [enabled, fetcher, page, livePage, cap, interval])
}, [enabled, fetcher, page, livePage, cap, interval, filter])
// F6 — a stats-backend switch invalidates every cursor we hold. `seq` is only
// monotonic across a sqlite→sqlite restart: memory→* restarts the counter at 1,
@@ -725,25 +975,43 @@ function useLiveLog<T extends HasSeq>(
// `after=<unreachable>` forever and the stream would sit silently frozen. Drop
// everything and re-read from the head instead — one visible reload beats a dead
// feed. Skipped on the first observation (there is nothing to invalidate yet).
const backendRef = useRef<string | undefined>(resetKey)
//
// The FILTER is folded into the same key for the same reason: every row we hold
// was selected under the old needle, so keeping them while the new one is in
// force would show matches for a search nobody ran.
//
// The separator is NUL because no reset key and no needle can contain one, so
// ("a","b") and ("ab","") cannot collide. It is written as the ESCAPE and never
// as a raw byte: one literal NUL in this source makes `file` call it binary, and
// ripgrep, `git grep` and every tree-wide search then SKIP the file in silence.
// On a codebase reviewed by grepping the tree, a file that answers no search is
// worse than the collision this separator guards against.
const gen = `${resetKey ?? ''}\u0000${filter}`
const backendRef = useRef<string>(gen)
useEffect(() => {
if (backendRef.current === resetKey) return
backendRef.current = resetKey
if (backendRef.current === gen) return
backendRef.current = gen
setRows([])
setBuffer([])
setFreshSeqs(new Set())
setExhausted(false)
setTruncated(false)
scanCursorRef.current = 0
liveCursorRef.current = 0
setLoaded(false) // re-arms the initial load effect above
}, [resetKey])
}, [gen])
return {
rows,
busy,
err,
showMore: rows.length > 0 && !exhausted,
// A truncated page has more log behind it even with zero rows on screen, so
// the way onward must stay reachable — that is the whole recovery path.
showMore: (rows.length > 0 || truncated) && !exhausted,
paused,
pending: buffer.length,
freshSeqs,
truncated,
reload: () => void reload(),
loadMore: () => void loadMore(),
togglePause,
@@ -753,7 +1021,7 @@ function useLiveLog<T extends HasSeq>(
// ---- page ------------------------------------------------------------------
export default function Insights() {
export default function Insights({ status }: { status: Status | null }) {
// Live aggregate snapshot — poll on Overview's 3s cadence, degrade to an honest
// "unavailable" state rather than freezing on stale numbers.
const [stats, setStats] = useState<Stats | null>(null)
@@ -803,13 +1071,36 @@ export default function Insights() {
// so no fetch fires in the off state. Cursors are derived from the rows, so
// Load more appends the next OLDER page and the live poll prepends NEW rows.
const loggingOff = stats?.backend === 'off'
// THE SERVICE ITSELF IS OFF — the reason this page is empty on a router nobody
// has switched on yet, and the reason it said nothing about it. The `ins-off`
// short-circuit below was written for one cause (logging switched off) and the
// shipped default is `memory`, not `off`, so it never fired: ten sections drew
// ten well-mannered "nothing yet" states and not one of them named the switch.
//
// Read through serviceIntent so `off` means the operator's switch and never an
// unreadable configuration — that one keeps the alarms it already has, and an
// unreachable stats endpoint is still reported by `statsOk` below.
const serviceOff = serviceIntent(status) === 'off'
// Wait for the first snapshot before touching the log endpoints: firing a fetch
// while `stats` is still null would hit a backend that may turn out to be "off".
const logsEnabled = stats !== null && !loggingOff
// Nothing is recorded with the service down either, so the pollers stay parked.
const logsEnabled = stats !== null && !loggingOff && !serviceOff
// The backend id doubles as the cursor-generation key: when it changes, `seq`
// may have restarted or jumped, so both logs reset and re-read from the head.
const conns = useLiveLog<ConnLogEntry>(getStatsConnsPage, logsEnabled, stats?.backend)
const dns = useLiveLog<QueryLogEntry>(getStatsLogPage, logsEnabled, stats?.backend)
// Search text: what the box holds vs what the daemon is actually filtering on.
// The debounced value is the one that goes on the wire, so a half-typed domain
// does not spend a scan budget or discard the rows already on screen.
const [connQ, setConnQ] = useState('')
const [dnsQ, setDnsQ] = useState('')
const connFilter = useDebounced(connQ, 350)
const dnsFilter = useDebounced(dnsQ, 350)
const conns = useLiveLog<ConnLogEntry>(
getStatsConnsPage,
logsEnabled,
stats?.backend,
connFilter,
)
const dns = useLiveLog<QueryLogEntry>(getStatsLogPage, logsEnabled, stats?.backend, dnsFilter)
// True from mount until the first log page has actually been fetched, so the
// panels read "loading…" instead of flashing "nothing logged yet".
const logsPending = !logsEnabled && statsOk
@@ -914,6 +1205,29 @@ export default function Insights() {
// effective backend is "off", nothing is recorded, so every list is empty — show
// an honest off state instead of a wall of "no data yet" cards. Absent backend on
// older daemons ⇒ treated as collecting, so this never trips on legacy snapshots.
// Said ONCE, before the grid, and it comes first: with the service off there is
// no engine to resolve a name or carry a packet, so "logging is on" is true and
// irrelevant and every section below is empty for this reason and no other.
// Ten honest empty states are still ten guesses to make about which one broke.
if (serviceOff) {
return (
<section className="page insights" aria-label="Insights">
<section className="ins-off" aria-label="The service is off">
<Led variant="off" />
<h2 className="ins-off-title">The service is off</h2>
<p className="ins-off-body">
Nothing is being routed, filtered, or recorded, so there is nothing to show here yet.
Turn the service on and apply the config — traffic starts appearing on this page within
seconds of the first lookup.
</p>
<a className="ins-off-cta" href="#/settings">
Turn it on in Settings
</a>
</section>
</section>
)
}
if (loggingOff) {
return (
<section className="page insights" aria-label="Insights">
@@ -1194,18 +1508,36 @@ export default function Insights() {
? 'log unavailable'
: (conns.busy || logsPending) && conns.rows.length === 0
? 'loading…'
: `${fmtNum(conns.rows.length)} shown · newest first`
: connFilter
? `${fmtNum(conns.rows.length)} ${conns.rows.length === 1 ? 'match' : 'matches'}`
: `${fmtNum(conns.rows.length)} shown · newest first`
}
right={
<LogSearch
id="ins-conn-q"
value={connQ}
onChange={setConnQ}
fields={CONN_SEARCH_HINT}
placeholder="find a device, host or rule"
/>
}
wide
>
{conns.rows.length === 0 ? (
<Empty>
{conns.err
? 'Connections log unreachable — the endpoint retries on the next refresh.'
: conns.busy || logsPending
? 'Reading the connection log…'
: 'No connections logged yet — once LAN clients open flows, device → destination events stream in here.'}
</Empty>
<>
{/* Named BEFORE the empty state, because the empty state is only true
when the scan reached the end — see ScanNotice. */}
{conns.truncated ? <ScanNotice onContinue={conns.loadMore} busy={conns.busy} /> : null}
<Empty>
{conns.err
? 'Connections log unreachable — the endpoint retries on the next refresh.'
: conns.busy || logsPending
? 'Reading the connection log…'
: connFilter || conns.truncated
? logEndNote({ filter: connFilter, truncated: conns.truncated, matches: 0 }).text
: 'No connections logged yet — once LAN clients open flows, device → destination events stream in here.'}
</Empty>
</>
) : (
<LogShell
ariaLabel="Connection events, newest first"
@@ -1213,7 +1545,13 @@ export default function Insights() {
onMore={conns.loadMore}
busy={conns.busy}
showMore={conns.showMore}
count={conns.showMore ? `${fmtNum(conns.rows.length)} loaded` : `${fmtNum(conns.rows.length)} · all loaded`}
truncated={conns.truncated}
count={logCountLabel({
rows: conns.rows.length,
showMore: conns.showMore,
truncated: conns.truncated,
filtered: !!connFilter,
})}
paused={conns.paused}
pending={conns.pending}
onTogglePause={conns.togglePause}
@@ -1238,18 +1576,34 @@ export default function Insights() {
? 'log unavailable'
: (dns.busy || logsPending) && dns.rows.length === 0
? 'loading…'
: `${fmtNum(dns.rows.length)} shown · newest first`
: dnsFilter
? `${fmtNum(dns.rows.length)} ${dns.rows.length === 1 ? 'match' : 'matches'}`
: `${fmtNum(dns.rows.length)} shown · newest first`
}
right={
<LogSearch
id="ins-dns-q"
value={dnsQ}
onChange={setDnsQ}
fields={LOG_SEARCH_HINT}
placeholder="find a domain, device or exit"
/>
}
wide
>
{dns.rows.length === 0 ? (
<Empty>
{dns.err
? 'DNS log unreachable — the endpoint retries on the next refresh.'
: dns.busy || logsPending
? 'Reading the DNS log…'
: 'No DNS decisions logged yet — resolve some DNS and they stream in here.'}
</Empty>
<>
{dns.truncated ? <ScanNotice onContinue={dns.loadMore} busy={dns.busy} /> : null}
<Empty>
{dns.err
? 'DNS log unreachable — the endpoint retries on the next refresh.'
: dns.busy || logsPending
? 'Reading the DNS log…'
: dnsFilter || dns.truncated
? logEndNote({ filter: dnsFilter, truncated: dns.truncated, matches: 0 }).text
: 'No DNS decisions logged yet — resolve some DNS and they stream in here.'}
</Empty>
</>
) : (
<LogShell
ariaLabel="DNS decisions, newest first"
@@ -1257,7 +1611,13 @@ export default function Insights() {
onMore={dns.loadMore}
busy={dns.busy}
showMore={dns.showMore}
count={dns.showMore ? `${fmtNum(dns.rows.length)} loaded` : `${fmtNum(dns.rows.length)} · all loaded`}
truncated={dns.truncated}
count={logCountLabel({
rows: dns.rows.length,
showMore: dns.showMore,
truncated: dns.truncated,
filtered: !!dnsFilter,
})}
paused={dns.paused}
pending={dns.pending}
onTogglePause={dns.togglePause}
+16
View File
@@ -495,6 +495,22 @@
border-left: 0;
border-top: 1px solid var(--groove);
}
/* Collapsing to one column was not enough on a phone. A grid column is sized by
its widest item's MIN-CONTENT, and a <select> reports the width of its longest
option ("Allow everything — most compatible", in the mono face) — so the
column stayed ~20px wider than the plate and the COPY beside it was clipped
mid-word at the right edge, which is how a sentence about what leaks loses its
second half. The select is allowed to shrink and ellipsise its own label
instead; the chosen option is still fully readable once opened, and no text
that states a consequence is cut. */
.nw-policy-ctl {
min-width: 0;
}
.nw-policy-ctl .fp-select {
max-width: 100%;
min-width: 0;
text-overflow: ellipsis;
}
}
/* A daemon info note about the current policy — neutral by design: it states a
+374 -41
View File
@@ -2,9 +2,11 @@ import './Networks.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
import type { Inbound, Interface, Model, Status } from '../api'
import type { Complete, Inbound, Interface, Model, Status } from '../api'
import { isLanNetwork, isWanNetwork, useInterfaces } from '../srcOptions'
import { sectionNotes } from '../findings'
import { killSwitchClosed } from '../planeState'
import { interceptLive } from '../intercept'
// The Networks page is the INGRESS editor — a thin editor over Model.Inbounds,
// following the same save-then-Apply contract as Nodes/DNS/Routing: every edit
@@ -98,6 +100,36 @@ const DEFAULT_TPROXY_PORT = 12345
*
* `icmp` still carries that destination-dependence for its NON-ping half, so its
* cost line says so rather than claiming "nothing else gets out".
*
* WHY THIS COPY IS NOW A FUNCTION AND NOT A TABLE. Two audit findings, one cause:
* a constant string cannot be true about a router whose behaviour three OTHER
* settings can override.
*
* 1. MULTICAST IPTV WAS PROMISED, AND NEVER WORKS. `direct` said "Ping, multicast
* IPTV, and connecting to a VPN ... all work" — an INSTRUCTION, and the worst
* kind of wrong: someone who wants IPTV reads it, moves to the most open rung
* on the ladder (which also permits a client's ESP/GRE straight past the
* proxy), and still has no IPTV. The stream is UDP; every rule this policy
* emits carries `meta l4proto != { tcp, udp }` so UDP never reaches one, and
* the fail-closed forward chain accepts only the RFC1918/link-local daddr
* sets — 224.0.0.0/4 is not there, and unconditional drops follow. The daemon
* says exactly this in the note rendered a few pixels below on this same page.
* IPTV is now stated ONCE, as its own line, and it says it does not work.
* 2. THE `block` COST LINE WAS UNCONDITIONAL. Three settings contradict it:
* - an OPEN kill-switch — the forward chain emits no drops at all;
* - Globals.L3Tunnel — prerouting marks ICMP echo into the engine's TUN
* BEFORE the forward chain, so ping keeps working, through the tunnel;
* - Globals.UntunnelableEgress — ESP/AH/GRE/IGMP/SCTP are marked and routed
* out a named interface, so the forward chain never rules on them.
* Neither of the last two existed in the panel's `Globals` type, so the page
* could not have told the truth about them even in principle; they were added
* (api.ts) rather than papered over with a vaguer sentence.
*
* The copy therefore describes only what the POLICY still decides, and a separate
* line names whatever another setting has taken off it. Detail beyond that belongs
* to the daemon's own note for this section (`policyNotes`), which is computed
* from the running plane and rendered right underneath — this copy's job is to not
* contradict it.
*/
type Untunnelable = 'block' | 'icmp' | 'direct'
@@ -122,31 +154,222 @@ interface PolicyCopy {
works: string
cost: string | null
tone: 'good' | 'warn'
/** What some OTHER setting decides instead of this one. `null` ⇒ nothing; this policy owns it all. */
claimed: string | null
}
const UNTUNNELABLE_COPY: Record<Untunnelable, PolicyCopy> = {
block: {
// Scoped to "this traffic" on purpose. The old line — "Nothing leaves except
// through the tunnel" — was doubly loose: it was false (see the note above),
// and even read charitably it collides with directly-routed TCP, which does
// leave outside the tunnel by design.
works:
'None of this traffic leaves the router — it’s dropped, whatever your routing rules say. It’s the only setting whose promise doesn’t depend on how the rules are written.',
cost: 'Ping and traceroute stop working from your devices. So do IPsec and PPTP VPN connections made from a device on your network, multicast IPTV, and SCTP. VPNs that run over UDP — WireGuard, OpenVPN-UDP, and IPsec through NAT (IKEv2/NAT-T) — are unaffected: they go through the tunnel like everything else.',
tone: 'good',
},
icmp: {
works: 'Ping and traceroute work everywhere, so you can check whether something is reachable.',
cost: 'Whatever you ping sees your real IP address instead of the tunnel’s. IPsec, PPTP and IPTV also get out — but only toward addresses your routing rules already send direct, so a VPN app on a device can still open its own connection beside this one if its server is one of those.',
tone: 'warn',
},
direct: {
works: 'Ping, multicast IPTV, and connecting to a VPN from a device on your network all work.',
cost: 'All of it goes out with your real IP, around the tunnel. A VPN app left running on a device keeps its own connection open beside this one — traffic through it isn’t proxied or filtered.',
tone: 'warn',
},
/**
* How much of this traffic the policy still decides.
*
* Only three combinations are reachable, which is why this is an enum and not two
* booleans: UntunnelableEgress claims EVERY untunnelable protocol (the kernel
* routes them out its device before the forward chain runs), so once it is set
* there is nothing left for L3Tunnel to change about the policy's scope.
*
* all — neither override is on. The policy decides everything.
* exceptPing — L3Tunnel only. Ping rides the tunnel; ESP/AH/GRE/SCTP are the
* policy's.
* none — UntunnelableEgress is set. Routing settles all of it first; the
* policy answers only for the case where that route fails to come up.
*/
type PolicyScope = 'all' | 'exceptPing' | 'none'
interface PolicyContext {
/** Globals.L3Tunnel. */
l3: boolean
/** Globals.UntunnelableEgress, trimmed. */
egress: string
/** The LIVE kill-switch, normalised the daemon's way. Open ⇒ the chain has no drops. */
killSwitchOpen: boolean
}
function policyScope(ctx: PolicyContext): PolicyScope {
if (ctx.egress) return 'none'
return ctx.l3 ? 'exceptPing' : 'all'
}
/** The sentence that stops an operator "fixing" a UDP VPN that was never broken. */
const UDP_VPNS_FINE =
'VPNs that run over UDP — WireGuard, OpenVPN-UDP, and IPsec through NAT (IKEv2/NAT-T) — are unaffected either way: they go through the tunnel like everything else.'
/** Which setting took this traffic off the policy, and what it does with it. */
function claimedCopy(ctx: PolicyContext): string | null {
// Each of these says only WHY the copy above has the shape it has — which other
// setting took the traffic, and therefore why the familiar promise is missing.
// What that setting then DOES with it is the daemon's note, published for this
// same section and rendered immediately below from the RUNNING plane. Saying it
// twice would make the shorter, staler one look like a second opinion.
if (ctx.egress && ctx.l3) {
return `Two other settings decide this before the one above is asked: ping goes through the tunnel (l3_tunnel), and everything else the tunnel can’t carry is routed out “${ctx.egress}” (untunnelable_egress).`
}
if (ctx.egress) {
return `Another setting decides this before the one above is asked: everything the tunnel can’t carry is routed out “${ctx.egress}” (untunnelable_egress).`
}
if (ctx.l3) {
// Deliberately shorter than the two above: when only the L3 ingress is on, the
// daemon publishes its own note for this section directly underneath and says
// the rest (which addresses can be pinged, and why the others cannot). This
// line exists to explain the SHAPE of the copy above it — why ping is missing
// from a policy that used to decide it — not to restate the daemon.
return 'Ping and Windows tracert are taken through the tunnel before the setting above is asked (l3_tunnel), so it no longer decides them.'
}
return null
}
/**
* The copy for the policy as it is actually behaving right now.
*
* Read it as: an open kill-switch beats everything (no drops are emitted at all,
* so no rung promises anything), then the scope decides how much of the ladder's
* usual story is still this setting's to tell.
*/
function policyCopy(policy: Untunnelable, ctx: PolicyContext): PolicyCopy {
const scope = policyScope(ctx)
const claimed = claimedCopy(ctx)
// Fail-open: the forward chain emits no drops, so every rung is inert. Saying
// what IS happening beats repeating a promise nothing is keeping.
if (ctx.killSwitchOpen) {
if (scope === 'none') {
return {
works:
'Nothing is being dropped, and nothing is left for this setting to decide: the kill-switch is open, and another setting has already taken this traffic.',
cost: 'If that route ever fails to come up, the traffic leaves through your normal connection with your real IP address, quietly, instead of failing.',
tone: 'warn',
claimed,
}
}
return {
works:
scope === 'exceptPing'
? 'Nothing is being dropped: with the kill-switch open the forward chain has no drops at all, so a device’s own IPsec or PPTP connection works too.'
: 'Nothing is being dropped: with the kill-switch open the forward chain has no drops at all, so ping and a device’s own IPsec or PPTP connection both work.',
// No "set the kill-switch to fail-closed" here: the moot note below owns
// that instruction, and printing it twice in one section is how the second
// copy stops being read.
cost: `It reaches the internet with your real IP address, around the tunnel. ${UDP_VPNS_FINE}`,
tone: 'warn',
claimed,
}
}
switch (policy) {
case 'block':
if (scope === 'none') {
return {
works:
'Where this setting still applies, the packet is dropped rather than let out — so a route that fails to come up fails honestly instead of leaking.',
cost: null,
tone: 'good',
claimed,
}
}
return {
// Scoped to "this traffic" on purpose. The old line — "Nothing leaves
// except through the tunnel" — was doubly loose: it was false (see the
// note above), and even read charitably it collides with directly-routed
// TCP, which does leave outside the tunnel by design.
works:
scope === 'exceptPing'
? 'Everything this setting still decides is dropped, whatever your routing rules say. If the ping route ever fails to come up, ping fails outright rather than leaking.'
: 'None of this traffic leaves the router — it’s dropped, whatever your routing rules say. It’s the only setting whose promise doesn’t depend on how the rules are written.',
cost:
scope === 'exceptPing'
? `IPsec and PPTP VPN connections made from a device on your network stop working, and so does SCTP. ${UDP_VPNS_FINE}`
: `Ping stops working from your devices. So do IPsec and PPTP VPN connections made from a device on your network, and SCTP. ${UDP_VPNS_FINE}`,
tone: 'good',
claimed,
}
case 'icmp':
if (scope === 'none') {
return {
works:
'Where this setting still applies, it lets ping out directly and drops the rest.',
cost: 'So if a route ever fails to come up, ping quietly leaves with your real IP address instead of failing.',
tone: 'warn',
claimed,
}
}
if (scope === 'exceptPing') {
return {
works:
'Ping already travels through the tunnel, so this rung’s exception for it only matters if that route fails to come up.',
cost: 'IPsec, PPTP and SCTP get out toward addresses your routing rules already send direct, with your real IP address — so a VPN app on a device can still open its own connection beside this one if its server is one of those. And if the ping route fails, ping leaves with your real address rather than failing.',
tone: 'warn',
claimed,
}
}
return {
works:
'Ping works everywhere, so you can check whether something is reachable.',
cost: 'Whatever you ping sees your real IP address instead of the tunnel’s. IPsec and PPTP also get out — but only toward addresses your routing rules already send direct, so a VPN app on a device can still open its own connection beside this one if its server is one of those.',
tone: 'warn',
claimed,
}
case 'direct':
if (scope === 'none') {
return {
works:
'Nothing is left for this setting to decide: another setting has already taken this traffic.',
cost: 'If that route ever fails to come up, this setting lets the traffic leave through your normal connection with your real IP address, quietly, instead of failing.',
tone: 'warn',
claimed,
}
}
return {
works:
scope === 'exceptPing'
? 'A device on your network can make its own IPsec or PPTP VPN connection. Ping already travels through the tunnel.'
: 'Ping works, and a device on your network can make its own IPsec or PPTP VPN connection.',
cost:
scope === 'exceptPing'
? 'That traffic goes out with your real IP, around the tunnel. A VPN app left running on a device keeps its own connection open beside this one — traffic through it isn’t proxied or filtered. If the ping route ever fails to come up, ping does the same instead of failing.'
: 'All of it goes out with your real IP, around the tunnel. A VPN app left running on a device keeps its own connection open beside this one — traffic through it isn’t proxied or filtered.',
tone: 'warn',
claimed,
}
}
}
/**
* Multicast IPTV, said once and said straight.
*
* It is stated unconditionally because it is unconditionally true — the stream is
* UDP and no rule this policy emits can match UDP, on any of the three rungs — and
* it is stated at all because the page used to promise the opposite under `direct`
* and under `icmp`. Someone whose IPTV is broken arrives here looking for the
* setting that fixes it; the useful thing to tell them is that there isn't one.
*/
const IPTV_LINE =
'Multicast IPTV is not one of these things: it doesn’t pass this router on any of the three settings, and “Allow everything” won’t bring it back.'
/**
* Plain `traceroute`, said once and said the daemon's way.
*
* FOUR RUNGS OF THIS PAGE CLAIMED "ping and traceroute work". They do not. On the
* production router an ordinary `traceroute` from a LAN device prints `* * *` and
* nothing else, on every setting here, `direct` included — its UDP probes are
* diverted by tproxy and delivered LOCALLY to the engine, and local delivery is
* not forwarding, so nothing on the path is ever provoked into a `time-exceeded`.
* Nobody debugging a blank trace was going to find that by themselves, and the
* page was actively sending them to look for a fault at their own end.
*
* The wording is `apply/warnings.go udpTracerouteFacts`, which the daemon
* publishes for this very section and which is rendered from the running plane a
* few pixels below this line. Two versions of one fact on one screen is how the
* shorter one becomes a second opinion, so this is deliberately the same three
* claims in the same order: it prints no hops, why, and what to use instead.
* netplane/untunnelable.go states it too. Change one, change all three.
*
* It is unconditional for the same reason IPTV_LINE is: it is unconditionally
* true, and someone whose trace printed nothing arrives here looking for the
* setting that fixes it. There isn't one.
*/
const TRACEROUTE_LINE =
'Plain traceroute on Linux and macOS is a separate matter, and it reads the same under every setting here, “Allow everything” included: its UDP probes are tunnelled and do reach the target, but it prints no hops at all, only * * *. Each probe is delivered locally to the engine, and local delivery is not forwarding, so nothing on the path is ever asked for a time-exceeded — there is no fault at your end to go looking for. Use traceroute -I (ICMP probes, which is what Windows tracert already sends) for a trace that prints hops.'
/** The addr:port an inbound binds — the generator's clash key (listenKey). */
function listenKey(in_: Inbound): string {
if (effectiveType(in_) === 'tproxy') {
@@ -348,11 +571,24 @@ export default function Networks({ status }: { status?: Status | null }) {
// Untunnelable-traffic policy. Normalised the same way the daemon does, so an
// absent/unknown UCI value reads as `block` here too rather than as blank.
const untunnelable = normUntunnelable(config?.Globals?.Untunnelable)
const untunnelableCopy = UNTUNNELABLE_COPY[untunnelable]
// Prefer the LIVE kill-switch off /api/status; fall back to the saved config
// when the shell hasn't got a status yet.
const killSwitchOpen =
(status?.kill_switch ?? config?.Globals?.KillSwitch ?? 'closed').toLowerCase() === 'open'
// when the shell hasn't got a status yet. The comparison is the daemon's own
// (planeState.killSwitchClosed) — this page normalised and planeState.ts did
// not, so the same router read differently on two pages.
const killSwitchOpen = !killSwitchClosed(status?.kill_switch ?? config?.Globals?.KillSwitch)
// The two settings that decide part of this traffic BEFORE the policy is
// consulted. Both are read from the saved config, like `untunnelable` itself:
// /api/status reports neither, and this section describes the setting the
// operator is editing. The daemon's own note below is the live counterpart.
const untunnelableCopy = useMemo(
() =>
policyCopy(untunnelable, {
l3: config?.Globals?.L3Tunnel === true,
egress: (config?.Globals?.UntunnelableEgress ?? '').trim(),
killSwitchOpen,
}),
[untunnelable, config?.Globals?.L3Tunnel, config?.Globals?.UntunnelableEgress, killSwitchOpen],
)
// The daemon's info notes about this policy — shown beside the control they
// describe. The fail-open case has its own dedicated line below, so drop that
// one here to avoid saying the same thing twice.
@@ -374,6 +610,9 @@ export default function Networks({ status }: { status?: Status | null }) {
[lanIfaces, inbounds],
)
const coveredCount = coverage.filter((c) => c.by).length
// Whether the configured coverage is in effect right now. Read live off
// /api/status, never from the config that produced `coverage` — see interceptLive.
const live = interceptLive(status ?? null)
// ---- mutations -----------------------------------------------------------
const addInbound = useCallback(
@@ -458,16 +697,33 @@ export default function Networks({ status }: { status?: Status | null }) {
<div className="nw-section" aria-label="Interception coverage">
<header className="nw-sec-hd">
<h2 className="nw-sec-title">Interception</h2>
{/* "configured" is now part of the count, because the lamps below can no
longer be read as "and it is happening" — see interceptLive. */}
<span className="nw-sec-count mono">
{coveredCount} / {lanIfaces.length} networks
{coveredCount} / {lanIfaces.length} configured
</span>
</header>
<p className="nw-sec-note">
Only a transparent (tproxy) inbound that is switched on sends a network’s traffic through
the engine. SOCKS and HTTP listeners are ports the router offers to whoever asks for them —
they don’t capture anything by themselves.
the engine, and only while the engine is running. SOCKS and HTTP listeners are ports the
router offers to whoever asks for them — they don’t capture anything by themselves.
</p>
{/* Said once, above the board, so the lamps below don't each have to carry
it. `stopped` is amber and not crit on purpose: the engine being down is
already an alarm, and it is raised once by the shell's status banner —
repeating it per network would be three copies of one fault. */}
{live !== 'running' && (
<p className="nw-sec-note nw-policy-note" role="status">
<Led variant={live === 'stopped' ? 'amber' : 'off'} />
<span>
{live === 'stopped'
? 'Nothing is being intercepted right now: the engine is not running, or no data plane is installed. What the lamps below show is what this config asks for, not what the router is doing.'
: 'Whether any of this is in effect has not been reported yet. The lamps below show what this config asks for.'}
</span>
</p>
)}
{lanIfaces.length === 0 ? (
<div className="nw-plate">
<p className="nw-plate-title">No LAN networks found</p>
@@ -479,19 +735,40 @@ export default function Networks({ status }: { status?: Status | null }) {
) : (
<ul className="nw-cov" aria-label="LAN networks and their interception state">
{coverage.map(({ iface, by }) => (
<li key={iface.name} className={by ? 'nw-cov-item on' : 'nw-cov-item off'}>
// Green only for coverage that is CONFIGURED and RUNNING. Configured
// but not running is amber (wired, not connected); configured with no
// reading is an unlit socket. An uncovered network stays unlit: it is
// not a fault, it is a network nobody asked to intercept.
<li
key={iface.name}
className={by && live === 'running' ? 'nw-cov-item on' : 'nw-cov-item off'}
>
<div className="nw-cov-hd">
<Led variant={by ? 'on' : 'off'} />
<Led
variant={
!by ? 'off' : live === 'running' ? 'on' : live === 'stopped' ? 'amber' : 'off'
}
/>
<span className="nw-cov-name mono">{iface.name}</span>
</div>
<span className="nw-cov-cidr mono">{iface.subnet || '—'}</span>
<span className="nw-cov-state">
{by ? (
{!by ? (
'Goes straight out — not intercepted'
) : live === 'running' ? (
<>
Through the tunnel via <strong className="mono">{by.Name}</strong>
</>
) : live === 'stopped' ? (
<>
Set to go through <strong className="mono">{by.Name}</strong> — not while the
engine is stopped
</>
) : (
'Goes straight out — not intercepted'
<>
Set to go through <strong className="mono">{by.Name}</strong> — not reported
as running
</>
)}
</span>
</li>
@@ -519,8 +796,8 @@ export default function Networks({ status }: { status?: Status | null }) {
</header>
<p className="nw-sec-note">
The tunnel carries the traffic almost everything uses — web, video, games, email. A few
things can’t go through it no matter what: ping, and the protocols that carry IPTV or a VPN
connection. Choose what happens to those.
things can’t go through it no matter what: ping, and the protocols a device uses to make
its own VPN connection. Choose what happens to those.
</p>
<div className="nw-policy">
@@ -549,13 +826,32 @@ export default function Networks({ status }: { status?: Status | null }) {
</div>
</div>
{/* Fail-open makes the whole policy moot — say so instead of letting the
page imply something is being blocked when nothing is. */}
{/* What another setting decides instead of this one. Unlit lamp, like the
daemon's notes below: it reports a configuration, not a fault. */}
{untunnelableCopy.claimed && (
<p className="nw-sec-note nw-policy-note" role="status">
<Led variant="off" />
<span>{untunnelableCopy.claimed}</span>
</p>
)}
{/* Said once, on every setting, because it is true on every setting. */}
<p className="nw-sec-note nw-policy-note">
<Led variant="off" />
<span>{IPTV_LINE}</span>
</p>
<p className="nw-sec-note nw-policy-note">
<Led variant="off" />
<span>{TRACEROUTE_LINE}</span>
</p>
{/* Fail-open makes the whole policy moot. The copy above now says what IS
happening; this line stays because it is the one that says what to DO. */}
{killSwitchOpen && (
<p className="nw-sec-note nw-policy-moot" role="status">
<Led variant="amber" /> This setting isn’t doing anything right now: the kill-switch is
set to fail-open, so traffic keeps flowing directly whenever the tunnel is down. Set it
to fail-closed in Settings for this choice to take effect.
<Led variant="amber" /> The kill-switch is set to fail-open, so traffic keeps flowing
directly whenever the tunnel is down. Set it to fail-closed in Settings for this choice
to take effect.
</p>
)}
@@ -758,7 +1054,22 @@ function toDraft(in_: Inbound): Draft {
* and a dokodemo listener binds whatever `TargetNetwork` says. Writing anything
* else would make the daemon warn about a flag no one can see.
*/
function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound {
/**
* Build the Inbound this draft describes — REBUILT per type, never extended.
*
* Rebuilding is the point: an inbound switched from `socks` to `tproxy` binds
* 0.0.0.0:TproxyPort, so a surviving `Listen`/`Auth` from its previous life would
* be a setting the panel shows nobody and the generator ignores. `base` is
* accepted and deliberately unused for that reason.
*
* The return type is `Complete<Inbound>` so the rebuild cannot go stale in
* silence: adding a field to `Inbound` fails the BUILD in all three branches
* below until each says what it wants done with it. That is the only mechanism
* here that makes the omission loud — `Ruleset.Format` is what a missing one
* costs (see ruleset.ts). `undefined` is a positive statement of "this shape has
* no use for it", and JSON.stringify drops it, so the wire form is unchanged.
*/
function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Complete<Inbound> {
void base
const common = { Name: d.Name.trim(), Enabled: enabled }
const port = Number.parseInt(d.Port, 10)
@@ -771,6 +1082,16 @@ function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound
TproxyPort: Number.parseInt(d.TproxyPort, 10) || DEFAULT_TPROXY_PORT,
TCP: d.TCP,
UDP: d.UDP,
// Binds 0.0.0.0:TproxyPort and cannot authenticate or rewrite a
// destination, so none of the listener fields mean anything here.
Listen: undefined,
Port: undefined,
Auth: undefined,
User: undefined,
Pass: undefined,
TargetAddr: undefined,
TargetPort: undefined,
TargetNetwork: undefined,
}
case 'socks':
case 'http':
@@ -784,6 +1105,12 @@ function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound
Pass: d.Auth === 'password' ? d.Pass : '',
TCP: true,
UDP: true,
// A local listener diverts no network and has no fixed target.
Network: undefined,
TproxyPort: undefined,
TargetAddr: undefined,
TargetPort: undefined,
TargetNetwork: undefined,
}
case 'dokodemo':
return {
@@ -796,6 +1123,12 @@ function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound
TargetNetwork: d.TargetNetwork,
TCP: true,
UDP: true,
// Diverts no network, and forwards everything on without authenticating.
Network: undefined,
TproxyPort: undefined,
Auth: undefined,
User: undefined,
Pass: undefined,
}
}
}
+220
View File
@@ -196,6 +196,138 @@
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
color: var(--amber);
}
/* "not built" — the saved switch says on and the engine has no such outbound. */
.badge--crit {
border-color: color-mix(in srgb, var(--crit) 55%, var(--groove));
color: var(--crit);
}
/* ---- last-apply findings, attached to the row they are about ----
Sits under the row's own two lines, inside the row plate, so a node the
generator threw away cannot read as an ordinary enabled node. Severity carries
the colour; the accent stays reserved for controls. */
.row-findings {
margin: 6px 0 0;
padding: 0;
list-style: none;
display: flex;
flex-direction: column;
gap: 5px;
}
.row-finding {
display: flex;
align-items: flex-start;
gap: 8px;
padding: 7px 9px;
border: 1px solid color-mix(in srgb, var(--amber) 40%, var(--groove));
border-radius: 6px;
background: color-mix(in srgb, var(--sink) 35%, transparent);
}
.row-finding--critical {
border-color: color-mix(in srgb, var(--crit) 45%, var(--groove));
}
.row-finding-msg {
flex: 1;
min-width: 0;
font-size: 12px;
line-height: 1.5;
color: var(--ink);
max-width: 82ch;
overflow-wrap: anywhere;
}
/* ---- one node's test reading ----
* The whole design brief for this strip is one distinction: a MEASURED failure
* and a row nothing measured must not look alike. So the failure is crit-red
* text behind a red lamp, and the unmeasured row is the faintest text on the
* page behind an UNLIT socket — the same register the Targets card uses for its
* unmeasured state, so the two pages say "no reading" the same way. Orange stays
* out of it; good / warn / crit carry the meaning. */
.node-test {
display: flex;
align-items: center;
gap: 8px;
flex-wrap: wrap;
margin-top: 6px;
font-size: 11.5px;
letter-spacing: 0.02em;
color: var(--dim);
}
.node-test-msg {
flex: 1 1 14ch;
min-width: 0;
overflow-wrap: anywhere;
}
.node-test-delay {
font-weight: 700;
color: var(--ink);
}
.node-test-exit {
color: var(--dim);
}
.node-test-exit--unknown {
font-family: var(--font-sans);
font-style: italic;
color: var(--faint);
cursor: help;
}
/* Which instrument took the number, in the quietest voice available: it matters
when the reading is questioned and never before. */
.node-test-src {
font-family: var(--font-sans);
font-size: 11px;
color: var(--faint);
cursor: help;
}
.node-test-at {
margin-left: auto;
font-size: 10.5px;
color: var(--faint);
}
/* A CARRIED reading — one an earlier run measured, kept because a new run no
longer wipes the board. On an inventory of 300 nodes most rows are carried, so
this must be visibly a different statement from "just measured": a bracket and
a brighter ink, never amber, because age is a qualifier and not a fault. */
.node-test-at--carried {
padding-left: 6px;
border-left: 2px solid color-mix(in srgb, var(--dim) 45%, transparent);
color: var(--dim);
cursor: help;
}
/* The hop that stopped the walk, read off `blocked_by`. Crit, because it names a
probe that ran and failed — the one `source:''` row that is a finding. */
.node-test-hop {
font-family: var(--font-mono);
font-size: 10.5px;
font-weight: 700;
color: var(--crit);
white-space: nowrap;
cursor: help;
}
.node-test--wait .node-test-msg {
color: var(--amber);
}
.node-test--bad .node-test-msg,
.node-test--crit .node-test-msg {
color: var(--crit);
}
.node-test--warn .node-test-msg {
color: var(--amber);
}
/* "Nothing measured this" — and it must be impossible to mistake for the line
above it. */
.node-test--none .node-test-msg,
.node-test--unknown .node-test-msg {
font-family: var(--font-sans);
color: var(--faint);
}
/* Findings that belong to no single row (see Nodes.tsx globalFindings). */
.node-findings {
margin-bottom: calc(var(--u, 8px) * 2);
}
.node-findings .row-findings {
margin-top: 0;
}
/* masked-credential marker */
.masked {
@@ -382,6 +514,80 @@
line-height: 1.45;
color: var(--faint);
}
/* protocol filter — a row of checkboxes that wraps rather than stretching the
grid column it sits in */
.opt-checks {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 4px 12px;
padding: 7px 0 1px;
}
.opt-check {
display: inline-flex;
align-items: center;
gap: 5px;
font-family: var(--font-mono);
font-size: 11.5px;
color: var(--dim);
cursor: pointer;
}
.opt-check input {
accent-color: var(--accent);
cursor: pointer;
}
.opt-check input:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 2px;
}
/* a Toggle sitting inside an .opt-field, with its explanation beside it */
.opt-toggle-row {
display: flex;
align-items: center;
gap: 10px;
padding-top: 4px;
}
.opt-toggle-row .opt-hint {
min-width: 0;
}
/* A caveat inside an options fieldset that changes what the setting DOES —
carries the amber LED, same plate as .dns-sec-warn on the DNS page. */
.opt-warn {
display: flex;
align-items: flex-start;
gap: 8px;
margin: 2px 0 0;
padding: 9px 11px;
border: 1px solid color-mix(in srgb, var(--amber) 45%, var(--groove));
border-radius: 8px;
background: linear-gradient(
180deg,
color-mix(in srgb, var(--amber) 8%, var(--raised)),
var(--raised)
);
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--ink);
max-width: 72ch;
}
/* One honest next step under a row that has nothing in it yet. */
.row-hint {
margin: 4px 0 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--faint);
max-width: 64ch;
}
.row-hint strong {
color: var(--dim);
font-weight: 600;
}
/* selects inherit .fp-input; give them a little room for the native arrow */
select.fp-input {
appearance: none;
@@ -641,6 +847,20 @@ select.fp-input {
letter-spacing: 0.06em;
color: var(--faint);
}
/* A collapsed bucket has to carry its own bad news: a 300-node subscription is
closed by default, and the per-row findings inside it are otherwise unreachable
without knowing to look. */
.group-flagged {
flex: none;
display: inline-flex;
align-items: center;
gap: 6px;
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
color: var(--amber);
}
.group-rows {
margin-top: 8px;
}
+812 -99
View File
File diff suppressed because it is too large Load Diff
+172 -27
View File
@@ -1,7 +1,7 @@
import { useCallback, useEffect, useRef, useState } from 'react'
import { Button, Led, Module } from '../components'
import type { LedVariant } from '../components'
import { fmtDateTime, fmtDuration } from '../format'
import { fmtClock, fmtDateTime, fmtDuration } from '../format'
import {
apply as apiApply,
rollback as apiRollback,
@@ -14,18 +14,24 @@ import type { Model, Stats, Status, StatusWarning } from '../api'
import { confirmTimeout } from '../pendingConfirm'
import { navigate } from '../router'
import type { Route } from '../router'
import { attentionFindings } from '../findings'
import { engineReadout, protectionState } from '../planeState'
import { attentionFindings, truncationNote } from '../findings'
import { applyRisk, engineReadout, killSwitchReadout, protectionState } from '../planeState'
import { rejectedReading, shortHash } from '../appliedConfig'
import type { RejectedReading } from '../appliedConfig'
import { ApplyRiskBand } from './Apply'
// null-safe length for a Go slice that may arrive as null.
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
const enabledCount = <T extends { Enabled?: boolean }>(a: T[] | null | undefined): number =>
(a ?? []).filter((x) => x.Enabled).length
// The config hash, shortened. Delegates to appliedConfig.shortHash so this row
// and the not-applied band cannot print two different strings for one
// configuration — in that state telling two configurations apart is the entire
// point of the reading. The em dash for "no hash" stays here: the band has no
// empty case to render.
function short(hash: string): string {
if (!hash) return '—'
const h = hash.replace(/^sha256:/, '')
return h.length > 12 ? h.slice(0, 12) : h
return hash ? shortHash(hash) : '—'
}
// Confirm is no longer one of them — see the note beside the controls row.
@@ -226,11 +232,11 @@ export function Overview({
// ---- derived display state ----
const g = config?.Globals
// The LIVE kill-switch from /api/status wins over the saved config: this is a
// status readout, so it must describe what is actually installed. Reading the
// config here let the strip claim "fail-closed" while an apply-time finding
// said the running plane was fail-open — two truths on one screen.
const killArmed = (status?.kill_switch ?? g?.KillSwitch ?? 'closed') === 'closed'
// There is deliberately no local `killArmed` any more. The page asked the same
// question twice — once here and once inside killSwitchReadout — and the local
// copy was the poorer of the two: it compared the raw string (so "Closed" read
// as fail-OPEN) and it was a boolean, which cannot say "the configuration could
// not be read". Both answers now come from the readout below.
// Offer rollback only when the daemon has something to revert to (armed
// commit-confirm snapshot or an engine last-good); otherwise hide the control.
const canRollback = status?.can_rollback ?? false
@@ -309,15 +315,36 @@ export function Overview({
const engineVariant: LedVariant = engine.variant
const protection = protectionState(status)
// Configured fail-closed AND actually enforcing it. `none` means nothing is
// installed, so the setting is inert no matter what it says.
const killInEffect = killArmed && status?.plane !== 'none'
// Configured fail-closed, actually enforcing it, or not known — three answers,
// and the third is not folded into the first. See planeState.killSwitchReadout.
const kill = killSwitchReadout(status, g?.KillSwitch)
// Findings that need attention. `info` notes are statements about the config,
// not problems, so they live beside the setting they describe (see findings.ts)
// — keeping this list to things someone could actually act on.
const warnings = attentionFindings(status?.warnings)
const criticalCount = warnings.filter((w) => w.severity === 'critical').length
// The daemon caps the published list at 50 and says so in an `info` note — the
// one channel this page filters away. Carried separately so the list can admit
// it is not the whole list. See findings.ts truncationNote.
const truncated = truncationNote(status?.warnings)
// What pressing "Apply config" below would do, when what it would do is take
// the whole network off the internet. Null otherwise — which is nearly always.
const risk = applyRisk(config)
/**
* THE CONFIGURATION ON DISK WAS REFUSED, and something older is running.
*
* Null on a healthy box and on a daemon too old to say — the page then looks
* exactly as it did. When it is not null, every reading below is about a
* DIFFERENT configuration than the one the operator saved, and the ones that
* name it (the hash, the traffic default, the findings) say so. See
* appliedConfig; the router-clock instant is formatted here because that module
* has no locale, and it is never turned into a browser-computed elapsed time.
*/
const rejected = rejectedReading(status, fmtClock(status?.apply_failed_since_unix ?? 0))
const stale = rejected !== null
return (
<section className="page" aria-label="Overview">
@@ -340,7 +367,16 @@ export function Overview({
</p>
)}
<Findings warnings={warnings} criticalCount={criticalCount} />
{/* The configuration on disk is not the one running. Above the findings,
because it says what the findings are ABOUT. */}
{rejected && <NotAppliedBand r={rejected} />}
<Findings
warnings={warnings}
criticalCount={criticalCount}
truncated={truncated}
stale={stale}
/>
<div className="grid">
{/* Groups, not nodes: a group is where a dial path is defined, so it is the
@@ -407,7 +443,15 @@ export function Overview({
}}
rows={[
{ k: 'egresses', v: String(len(config?.Egresses)) },
{ k: 'default', v: defaultTarget(status, config), hot: true },
// Where traffic goes under the config that is RUNNING. When the config
// on disk was refused that is the older one, and the row has to say so
// rather than answer "where does my traffic go" about a config the
// operator replaced.
{
k: stale ? 'default · previous config' : 'default',
v: defaultTarget(status, config),
hot: true,
},
]}
/>
@@ -447,17 +491,21 @@ export function Overview({
/>
{/* A kill-switch set to fail-closed is only ARMED if something is actually
installed to enforce it. With no plane it is configured but inert, and
saying "ARMED" there would be a false reassurance next to a readout
that says nothing is protected. */}
installed to enforce it, and "we haven't been told" is neither. With no
plane it is configured but inert; with no reading the lamp stays unlit
rather than joining the healthy branch by default. */}
<Module
name="Kill-switch"
value={killInEffect ? 'ARMED' : killArmed ? 'NOT IN EFFECT' : 'OPEN'}
led={{ variant: killInEffect ? 'on' : killArmed ? 'crit' : 'amber' }}
value={kill.value}
led={{ variant: kill.variant }}
rows={[
{ k: 'setting', v: killArmed ? 'fail-closed' : 'fail-open', hot: !killArmed },
...(killArmed && !killInEffect
? [{ k: 'blocking now', v: 'no — nothing installed', hot: true }]
// From the readout, not from a second local comparison: the old
// `killArmed ? 'fail-closed' : 'fail-open'` had no third answer, so an
// unreadable configuration printed a confident "fail-closed" beneath a
// lamp that said NOT REPORTED. See planeState.killSwitchReadout.
{ k: 'setting', v: kill.setting, hot: kill.settingHot },
...(kill.blockingNow
? [{ k: 'blocking now', v: kill.blockingNow, hot: kill.hot }]
: [{ k: 'ipv6', v: g?.IPv6 ? 'covered' : 'off' }]),
{ k: 'confirm', v: g?.ConfirmTimeout ? `${g.ConfirmTimeout}s window` : 'no auto-rollback' },
]}
@@ -481,7 +529,15 @@ export function Overview({
led={{ variant: engineVariant }}
rows={[
{ k: 'process', v: engine.word, hot: engineVariant === 'crit' },
{ k: 'config hash', v: <span className="mono">{short(status?.hash ?? '')}</span> },
// The hash identifies the config the ENGINE is running, which is not
// always the one on disk. Unqualified it is the most misleading value
// on this page in the rejected state: a real hash of a real running
// configuration, just not the one that was saved.
{
k: stale ? 'config hash · not yours' : 'config hash',
v: <span className="mono">{short(status?.hash ?? '')}</span>,
hot: stale,
},
// Uptime of the daemon PROCESS. "started" is the moment it came up,
// by the router's clock — not the moment a config was applied.
{ k: 'running for', v: uptimeText || '—' },
@@ -492,6 +548,13 @@ export function Overview({
/>
</div>
{/* The same forecast the Apply page shows, because this page has the same
button. It is predicted from the CONFIG, so it is silent on the router
that has already been applied into that state — protectionState's
"Nothing is getting out" is the readout for that one, and it is at the
top of this page. See planeState.applyRisk. */}
{risk && <ApplyRiskBand risk={risk} />}
{/* Apply / Confirm / Rollback — active voice, honest results. */}
<div className="controls" role="group" aria-label="Config actions">
<div className="controls-btns">
@@ -536,11 +599,23 @@ const SECTION_ROUTE: Record<string, Route> = {
rule: 'routing',
ruleset: 'routing',
blocklist: 'dns',
allowlist: 'dns',
resolver: 'dns',
dns_rule: 'dns',
device: 'devices',
chain: 'targets',
group: 'targets',
// A node the generator dropped (unparseable share link, duplicate WireGuard
// key, name colliding with a reserved tag) is reported under `node` — and had
// nowhere to jump to, so the one page that could show it a green toggle was
// also the one page the finding could not reach.
node: 'nodes',
subscription: 'nodes',
egress: 'targets',
inbound: 'networks',
interface: 'networks',
profile: 'profiles',
alert: 'settings',
// The standing note about non-TCP/UDP traffic — its control lives on Networks.
untunnelable: 'networks',
}
@@ -558,11 +633,22 @@ const SECTION_ROUTE: Record<string, Route> = {
function Findings({
warnings,
criticalCount,
truncated,
stale,
}: {
warnings: StatusWarning[]
criticalCount: number
/** The daemon's "N further suppressed" note, when the list was capped. */
truncated: StatusWarning | null
/**
* The config on disk was refused. These findings come from the last SUCCESSFUL
* apply, so they are about the configuration that is running — not the one the
* operator saved. Under "Last apply" that reads as a report on their edit, and
* a clean list then reads as "your edit is fine".
*/
stale: boolean
}) {
if (warnings.length === 0) return null
if (warnings.length === 0 && !truncated) return null
const rank = { critical: 0, warning: 1, info: 2 } as const
const sorted = [...warnings].sort((a, b) => rank[a.severity] - rank[b.severity])
@@ -570,13 +656,26 @@ function Findings({
return (
<section className="findings" aria-label="Findings from the last apply">
<header className="findings-hd">
<h2 className="findings-title">Last apply</h2>
<h2 className="findings-title">{stale ? 'Last successful apply' : 'Last apply'}</h2>
<span className="findings-count mono">
{/* "at least" whenever the list was capped: the counts below it are a
floor, not a total, and the cap drops the least severe FIRST — so
on a config with fifty criticals the thing it drops is a critical. */}
{truncated ? 'at least ' : ''}
{criticalCount > 0
? `${criticalCount} critical · ${warnings.length} total`
: `${warnings.length} note${warnings.length === 1 ? '' : 's'}`}
</span>
</header>
{/* Its own line, not a third item in the header: at 390 px the header's
three cells shrank the sentence to four words a column. It is a sentence
about the whole list, so it sits above the list. */}
{stale && (
<p className="findings-stale">
These describe the configuration that is RUNNING — not the one you saved. The refusal
is in the band above.
</p>
)}
<ul className="findings-list">
{sorted.map((w, i) => {
const route = SECTION_ROUTE[w.section]
@@ -610,6 +709,20 @@ function Findings({
</li>
)
})}
{/* The list saying it is not the whole list. Last, because it is about
everything above it — and never filtered out with the other `info`
notes, which is where it used to disappear. */}
{truncated && (
<li className="finding finding--truncated">
<Led variant="amber" />
<div className="finding-copy">
<span className="finding-where mono">list truncated</span>
<span className="finding-msg">
Some findings are missing from this list. {truncated.message}
</span>
</div>
</li>
)}
</ul>
</section>
)
@@ -630,6 +743,38 @@ const NAV_LABEL: Record<Route, string> = {
apply: 'Apply',
}
/**
* THE CONFIGURATION ON DISK IS NOT THE ONE RUNNING.
*
* It borrows the .risk-band vocabulary because the severity matches, but the
* eyebrow names a different tense: the hazard band is a FORECAST of what a button
* would do, this one reports a state the box is already in.
*
* The last line is the one the defect was missing. Saying "refused" is not
* enough: `engine_running` is true, the LED is green, the hash looks normal, and
* the traffic verdict answers confidently — all about a configuration the
* operator replaced. So the band names what everything else on the page is about
* before the reader gets to any of it.
*/
function NotAppliedBand({ r }: { r: RejectedReading }) {
return (
<section className="stale-band" role="alert" aria-label="Configuration not applied">
<div className="risk-band-hd">
<Led variant="crit" />
<span className="risk-eyebrow">Not applied</span>
</div>
<p className="risk-headline">The configuration on disk is not the one running.</p>
<p className="risk-detail">
/etc/config/shater was refused at <strong className="stale-stage">{r.stageText}</strong>.
{r.stageHint ? ` ${r.stageHint}` : ''}
</p>
<p className="risk-detail stale-cause mono">{r.cause}</p>
<p className="risk-detail">{r.persistence}</p>
<p className="risk-undo hot">{r.scope}</p>
</section>
)
}
// (StatusPip lived here. It backed the five-pip status strip, which collapsed
// into the single protection-state readout above; Apply.tsx keeps its own copy
// for the apply/rollback flow, where the individual flags are the actual
+38 -4
View File
@@ -102,17 +102,32 @@
padding: calc(var(--u, 8px) * 4) 0 calc(var(--u, 8px) * 3);
text-align: center;
}
/* The empty state stopped being one line the day it started NAMING the default
* route (see liveDefaultRoute). Five centred lines of mono is a wall, so the plate
* stays centred while the prose inside is a measured, left-aligned column in the
* page's reading face — the route tag keeps the mono, because it is a value. */
.rt-empty p {
margin: 0;
font-family: var(--font-mono);
margin: 0 auto;
max-width: 58ch;
font-family: var(--font-sans);
font-size: 13px;
line-height: 1.55;
text-align: left;
color: var(--ink);
}
.rt-empty-sub {
margin-top: 6px !important;
font-size: 11.5px !important;
margin-top: 8px !important;
font-size: 12px !important;
color: var(--faint) !important;
}
/* Amber under BOTH branches, and that is not an oversight: with no rules at all,
* `block` means the LAN has no internet and `direct` means the LAN is on the plain
* WAN with its real address. Neither is a resting state, so neither gets the green
* or the accent that would say "this is fine". */
.rt-empty-route {
color: var(--amber);
font-weight: 600;
}
/* ---- the bus ---- */
.rt-list {
@@ -321,6 +336,25 @@
line-height: 1.45;
color: var(--dim);
}
/* ---- the per-rule kill policy (Rule.Kill) ----
*
* The same amber pill as .rt-badge.dead, and for the same reason: --amber is warn,
* --accent is ACTIVE. A rule that fails open is not broken and not an incident —
* it is a deliberate weakening — so it must be legible without reading as an
* alarm, and it must never wear the orange that means "this is working".
* Only `open` and an unreadable value are marked at all; the fail-closed default
* draws nothing (see killMark). */
.rt-badge.kill {
padding: 1px 7px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 999px;
background: color-mix(in srgb, var(--amber) 12%, transparent);
color: var(--amber);
}
.rt-kill-note {
margin: 0;
}
/* The target is still what the operator asked for, so it stays readable — just
* quiet, because the router is not using it. */
.rt-rule.dead .rt-target {
+239 -30
View File
@@ -13,6 +13,17 @@ import {
} from '../api'
import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
import { everyLabel, relFetch } from '../format'
import { effectiveTarget, liveDefaultRoute } from '../defaultRoute'
import { killSwitchClosed } from '../planeState'
import {
carryKill,
killMark,
killPolicy,
killSelectValue,
KILL_OPEN_FORM_WARN,
KILL_OPTIONS,
} from '../killPolicy'
import { carryRulesetFormat } from '../ruleset'
// ---------------------------------------------------------------------------
// The api.ts `Rule` is a deliberately thin subset (Name/Enabled/Order/Target/
@@ -60,16 +71,31 @@ type RuleForce = {
}
/**
* Everything `Proto` can match, and nothing else. The engine understands two
* transports and exactly ten application protocols its sniffers can name
* (generate/route.go sniffedProtocols); a value outside this set builds a rule
* that is perfectly valid and can never fire — so its traffic quietly falls
* Everything `Proto` can match, and nothing else. A value outside this set builds
* a rule that is perfectly valid and can never fire — so its traffic quietly falls
* through to whatever rule sits below it. That is why this is a closed list and
* not a text box.
*
* Split into two groups because they answer different questions: the transport is
* known the moment a packet arrives, while an app protocol is only known once the
* first bytes have been read and labelled.
* Three groups, because the engine reads them through three different matchers
* (generate/route.go, ruleMatchers) and they answer different questions:
*
* - Transport — the L4 network. Known the moment a packet arrives.
* - Detected protocol — the L7 label a sniffer puts on a connection once its
* first bytes have been read. This group, and ONLY this group, is the engine's
* `sniffedProtocols` set; anything else routed into that matcher is inert.
* - Layer 3 — ICMP. Not a sniffed label: it lands in the emitted rule's
* `network`, never in `protocol` (the sniffers are skipped outright for an
* ICMP flow, so they never report "icmp"). All three spellings are the SAME
* one network; `icmpv4`/`icmpv6` additionally pin `ip_version`, which the
* engine derives from the destination address.
*
* ICMP carries caveats the picker deliberately does not try to enforce, because
* the daemon reports each one against the whole config on apply: it reaches the
* engine only while globals l3_tunnel is on, it has no ports (a port matcher
* beside it can never be satisfied), `icmpv6` also needs globals ipv6 on, and it
* is DROPPED rather than falling through when routed at a target that cannot
* carry layer 3 — i.e. every proxy protocol. Only wireguard/AmneziaWG nodes and
* direct/interface egresses can carry a ping.
*/
const PROTO_TRANSPORT: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'tcp', label: 'TCP' },
@@ -87,10 +113,27 @@ const PROTO_APP: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'rdp', label: 'RDP' },
{ id: 'ntp', label: 'NTP' },
]
const PROTO_VALUES = new Set([...PROTO_TRANSPORT, ...PROTO_APP].map((p) => p.id))
/**
* The family-qualified spellings are offered next to plain `icmp` rather than
* hidden behind it: the engine treats them as first-class and the difference is
* observable (an `ip_version` item on the same rule), so hiding them would leave a
* capability reachable only by hand-editing /etc/config/shater — and would mean
* that anyone who edited such a rule here lost the narrowing on the next save.
*/
const PROTO_L3: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'icmp', label: 'ICMP (ping)' },
{ id: 'icmpv4', label: 'ICMP — IPv4 only' },
{ id: 'icmpv6', label: 'ICMP — IPv6 only' },
]
const PROTO_VALUES = new Set(
[...PROTO_TRANSPORT, ...PROTO_APP, ...PROTO_L3].map((p) => p.id),
)
/** The Proto picker's option list — shared by the inline add row and the editor. */
function ProtoOptions({ value }: { value: string }) {
// The engine lower-cases `Proto` before matching it, so a hand-written `ICMP`
// is a working rule; judge it the same way and flag only what really is inert.
const matches = PROTO_VALUES.has(value.trim().toLowerCase())
return (
<>
<option value="">any</option>
@@ -108,10 +151,18 @@ function ProtoOptions({ value }: { value: string }) {
</option>
))}
</optgroup>
{/* A stored value the engine can't detect is kept and flagged, never
silently rewritten — the rule it belongs to is live right now. */}
<optgroup label="Layer 3">
{PROTO_L3.map((p) => (
<option key={p.id} value={p.id}>
{p.label}
</option>
))}
</optgroup>
{/* A stored value none of the groups spells verbatim is kept and offered as
written, never silently rewritten — the rule it belongs to is live right
now. It is flagged only when the engine cannot match it either. */}
{value !== '' && !PROTO_VALUES.has(value) && (
<option value={value}>{value} — never matches</option>
<option value={value}>{matches ? value : `${value} — never matches`}</option>
)}
</>
)
@@ -311,13 +362,6 @@ function defaultRouteConsequence(
)
}
/** Effective routing target for a rule (Target wins; a bare Egress is a target too). */
function effectiveTarget(r: RRule): string {
if (r.Target && r.Target.trim()) return r.Target.trim()
if (r.Egress && r.Egress.trim()) return `egress:${r.Egress.trim()}`
return 'direct'
}
/** Semantic tone for a target chip — block is critical, direct is quiet, else accent. */
function targetTone(t: string): 'block' | 'direct' | 'proxy' {
if (t === 'block') return 'block'
@@ -365,6 +409,8 @@ interface AddForm {
rulesets: string[]
proto: string
target: string
/** Rule.Kill, as the picker spells it — see killPolicy.ts. */
kill: string
schedEnabled: boolean
schedDays: string[]
schedStart: string
@@ -376,7 +422,27 @@ const EMPTY_FORM: AddForm = {
port: '',
rulesets: [],
proto: '',
target: 'direct',
// NO DEFAULT TARGET, and that is the point.
//
// It was `direct`. A form with nothing typed into it is a rule with no matchers,
// i.e. the router's DEFAULT ROUTE — so pressing "Add rule" on an untouched form
// put the whole network onto the plain WAN, with the real address, in one
// keystroke. The page flagged the consequence honestly ("no matchers — matches
// everything", and after an apply a critical "nothing is going through the
// tunnel"), which is precisely what made the pre-filled value so expensive: the
// trap was not the warning, it was that a value nobody chose had already been
// chosen for them, and it was the most dangerous one on the list.
//
// `block` was the tempting fix and is the wrong one for the same reason: it also
// lets one keystroke reconfigure the router, just toward an outage instead of a
// leak. The recoverable side of THIS default is not a safer value — it is no
// value. An empty target refuses to save and asks the question (see onAdd).
target: '',
// Fail-closed is the daemon's own default for a rule with no `option kill`, and
// a new rule must not start life on the side that leaks. Unlike the target, this
// one HAS a safe side: it decides what happens on a failure, so leaving it unset
// cannot itself reconfigure anything.
kill: 'closed',
schedEnabled: false,
schedDays: [],
schedStart: '',
@@ -732,7 +798,10 @@ export default function Routing() {
[rules],
)
const killSwitch = (config?.Globals?.KillSwitch ?? 'closed') === 'open' ? 'open' : 'closed'
// Normalised by the daemon's own rule — `=== 'open'` read "OPEN" as fail-CLOSED,
// so this dialog would have described a blocking router that isn't blocking.
// See planeState.killSwitchClosed.
const killSwitch = killSwitchClosed(config?.Globals?.KillSwitch) ? 'closed' : 'open'
const onToggle = useCallback(
async (name: string) => {
@@ -861,6 +930,15 @@ export default function Routing() {
setFormError('A rule with that name already exists.')
return
}
// The target has no default (see EMPTY_FORM) and an empty one is not a rule
// shape the daemon reads as "nothing": generate resolves an empty target to
// `direct`, so saving one would write the leak this refusal exists to stop.
if (!form.target.trim()) {
setFormError(
'Pick where this rule sends its traffic — a rule with no target routes it direct, out of the tunnel.',
)
return
}
setFormError(null)
const rule: RRule = {
Name: name,
@@ -874,6 +952,10 @@ export default function Routing() {
Proto: form.proto,
Target: form.target,
Egress: '',
// Written explicitly, including the safe default: `option kill 'closed'`
// says which side of the kill-switch this rule chose, where a missing
// option says only that nobody was asked.
Kill: form.kill,
// Schedule (Phase 7): only carried when the admin enabled it. Empty days
// = every day; empty start = 00:00; empty/equal end = all-day. The engine
// (generate) emits the rule only inside this window; cron reconciles at
@@ -949,13 +1031,23 @@ export default function Routing() {
const enabledCount = force.filter((f) => f.on).length
const overridden = force.filter((f) => f.profile !== null)
const overrideProfile = overridden[0]?.profile ?? null
// The route unmatched traffic takes right now — named in the lead and in the
// empty state, because "the default route" on its own is what let a blocked
// network read as normal. See liveDefaultRoute.
const def = liveDefaultRoute(rules, isLiveDefault, killSwitch)
return (
<section className="page" aria-label="Routing rules">
<div className="rt-intro">
<p className="rt-lead">
Rules run top to bottom on the bus — the <strong>first match wins</strong>. Traffic that
reaches the bottom follows the default route.
reaches the bottom follows the default route, which right now is{' '}
<strong className="mono">{def.target}</strong>
{def.rule ? (
<> — rule “{def.rule}”.</>
) : (
<> — no rule claims it, so the kill-switch decides.</>
)}
</p>
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules in force`}>
{enabledCount}
@@ -996,8 +1088,39 @@ export default function Routing() {
{rules.length === 0 ? (
<div className="rt-empty">
<p>No rules — all traffic follows the default route.</p>
<p className="rt-empty-sub">Add a rule below to steer a destination list, source, or port.</p>
{/* Names the route and says what it costs. The kill-switch is the only
thing deciding it here — with no rules there is no catch-all to
inherit Final — so each branch says which setting the operator would
have to change, not just what is happening. */}
{killSwitch === 'open' ? (
<>
<p>
No rules — so everything from your network follows the default route,{' '}
<strong className="mono rt-empty-route">direct</strong>: it leaves through your normal internet
connection with your real IP address, unproxied and unfiltered. The kill-switch is
set to <strong>fail-open</strong>, which is what chooses that.
</p>
<p className="rt-empty-sub">
Add a rule below. A rule with no matchers becomes the default route and takes it
over.
</p>
</>
) : (
<>
<p>
No rules — so everything from your network follows the default route,{' '}
<strong className="mono rt-empty-route">block</strong>:{' '}
<strong>devices have no internet</strong>{' '}
until a rule says where their traffic should go. The kill-switch is set to{' '}
<strong>fail-closed</strong>, and with no rule claiming the default, blocking is
what that means.
</p>
<p className="rt-empty-sub">
Add a rule below with no matchers — it becomes the default route — and point it at a
group, a node or <span className="mono">direct</span>.
</p>
</>
)}
</div>
) : (
<ol className="rt-list" aria-label="Routing rules in first-match order">
@@ -1120,6 +1243,12 @@ function RuleRow({
const isDefault = isCatchAll(rule) && !dead
const target = effectiveTarget(rule)
const tone = targetTone(target)
// What this rule does if its target cannot be built (Rule.Kill). null on the
// fail-closed default — see killMark: badging every row would price a deliberate
// bypass the same as the safe side, which is the reading this mark exists to
// prevent. Drawn even on an inert row: the policy is what the config SAYS, and a
// rule that is inert today is one edit away from being live.
const kill = killMark(rule)
// Dimmed by the EFFECTIVE state, never by the saved one. A rule the active
// profile switched off is not in force, and the row has to read that way even
// though its switch — which edits the saved setting — is still on.
@@ -1166,6 +1295,7 @@ function RuleRow({
faceplate's active state, and a force-enabled rule is exactly that. */}
{force.dir === 'disabled' && <span className="rt-badge dead">off · by profile</span>}
{force.dir === 'enabled' && <span className="rt-badge prof-on">on · by profile</span>}
{kill && <span className="rt-badge kill">{kill.badge}</span>}
</div>
<div className="rt-match">
{unmigrated ? (
@@ -1201,6 +1331,10 @@ function RuleRow({
overrides several rules would otherwise repeat that paragraph on every
one of them. What is left is the part only this row can say: whether it
is in force, and what its own switch is showing instead. */}
{/* The badge names the policy; this says what it costs. Under the matchers,
like the profile note, because it is a consequence of the row rather
than one of its conditions. */}
{kill && <p className="rt-dead-note rt-kill-note">{kill.note}</p>}
{force.profile && (
<p className="rt-dead-note rt-prof-note">
{force.dir === 'disabled' ? 'Not in force' : 'In force'} — profile{' '}
@@ -1353,7 +1487,16 @@ function Matchers({ rule }: { rule: RRule }): ReactNode {
* Interfaces-egresses, and the huge Nodes list LAST. `current` re-surfaces a
* value that isn't in the live config (e.g. a target pointing at a since-removed
* node) as its own option so editing a rule can never silently drop its target. */
function TargetOptions({ targets, current }: { targets: TargetGroups; current?: string }): ReactNode {
function TargetOptions({
targets,
current,
placeholder,
}: {
targets: TargetGroups
current?: string
/** Add form only: the unchosen state, which the submit handler refuses. */
placeholder?: string
}): ReactNode {
const known = useMemo(() => {
const s = new Set<string>()
for (const g of [targets.simple, targets.groups, targets.chains, targets.egresses, targets.nodes])
@@ -1362,6 +1505,7 @@ function TargetOptions({ targets, current }: { targets: TargetGroups; current?:
}, [targets])
return (
<>
{placeholder && <option value="">{placeholder}</option>}
{current && current.trim() && !known.has(current) && (
<option value={current}>{current} · current</option>
)}
@@ -1420,6 +1564,47 @@ function TargetOptions({ targets, current }: { targets: TargetGroups; current?:
* Rulesets panel's job, deliberately kept out of the rule editor so a list is
* created in exactly one place.
*/
/**
* The per-rule fallback picker — `Rule.Kill`.
*
* It sits next to Target because it is a statement ABOUT the target: what this
* rule does on the day that target cannot be built. Two options only, because the
* daemon reads only two outcomes (killPolicy.ts). A stored value that is neither
* is re-surfaced as its own option, exactly as TargetOptions re-surfaces a target
* pointing at a since-removed node: a picker that quietly dropped it would be
* rewriting a field it never showed, and would erase the evidence of the typo the
* daemon is warning about.
*/
function KillField({
value,
busy,
onChange,
}: {
value: string
busy: boolean
onChange: (v: string) => void
}) {
const unknown = killPolicy(value) === 'unknown'
return (
<label className="rt-field">
<span className="rt-flabel">If target fails</span>
<select
className="rt-input mono"
value={value}
onChange={(e) => onChange(e.target.value)}
disabled={busy}
>
{unknown && <option value={value}>“{value}” · not recognised, blocks</option>}
{KILL_OPTIONS.map((o) => (
<option key={o.value} value={o.value}>
{o.label}
</option>
))}
</select>
</label>
)
}
function RulesetPicker({
options,
selected,
@@ -1645,9 +1830,11 @@ function AddRule({
value={form.target}
onChange={(e) => set('target', e.target.value)}
>
<TargetOptions targets={targets} />
<TargetOptions targets={targets} placeholder="— pick a target —" />
</select>
</label>
<KillField value={form.kill} busy={busy} onChange={(v) => set('kill', v)} />
</div>
<RulesetPicker
@@ -1676,6 +1863,9 @@ function AddRule({
{noMatchers && !error && (
<span className="rt-edit-warn">no matchers — matches everything</span>
)}
{killPolicy(form.kill) === 'open' && !error && (
<span className="rt-edit-warn">{KILL_OPEN_FORM_WARN}</span>
)}
{error && (
<span className="rt-add-error" role="alert">
{error}
@@ -1714,6 +1904,9 @@ function RuleEditForm({
const [port, setPort] = useState(initial.DstPort ?? '')
const [proto, setProto] = useState(initial.Proto ?? '')
const [target, setTarget] = useState(effectiveTarget(initial))
// The stored spelling, shown as the picker spells it. An unrecognised value is
// carried through verbatim so the form cannot rewrite what it did not offer.
const [kill, setKill] = useState(killSelectValue(initial.Kill))
const [rulesets, setRulesets] = useState<string[]>([...(initial.DstRuleset ?? [])])
const [schedEnabled, setSchedEnabled] = useState(!!initial.SchedEnabled)
const [schedDays, setSchedDays] = useState<string[]>([...(initial.SchedDays ?? [])])
@@ -1747,8 +1940,8 @@ function RuleEditForm({
setErr('A rule with that name already exists.')
return
}
// Spread carries Order/Enabled/Kill/Egress + any off-form field through
// untouched; the form fields below overwrite exactly the matchers/target/schedule.
// Spread carries Order/Enabled/Egress + any off-form field through untouched;
// the form fields below overwrite exactly the matchers/target/kill/schedule.
// An empty field clears its matcher on purpose (you can strip a matcher this way).
// Saving re-anchors the schedule to THIS browser's UTC offset: the times in the
// form are taken as your wall clock (the router evaluates at UTC+offset).
@@ -1760,6 +1953,10 @@ function RuleEditForm({
DstRuleset: rulesets,
Proto: proto,
Target: target,
// carryKill keeps the stored spelling when the policy did not change: ''
// and 'default' mean 'closed', and rewriting one into the other would put an
// explicit option on every rule anyone ever opened this form for.
Kill: carryKill(initial.Kill, kill),
SchedEnabled: schedEnabled,
SchedDays: schedEnabled ? schedDays : [],
SchedStart: schedEnabled ? schedStart : '',
@@ -1848,6 +2045,8 @@ function RuleEditForm({
<TargetOptions targets={targets} current={target} />
</select>
</label>
<KillField value={kill} busy={busy} onChange={setKill} />
</div>
<RulesetPicker options={rulesetOptions} selected={rulesets} busy={busy} onToggle={toggleRuleset} />
@@ -1875,6 +2074,9 @@ function RuleEditForm({
{noMatchers && !err && (
<span className="rt-edit-warn">no matchers — matches everything</span>
)}
{killPolicy(kill) === 'open' && !err && (
<span className="rt-edit-warn">{KILL_OPEN_FORM_WARN}</span>
)}
{err && (
<span className="rt-add-error" role="alert">
{err}
@@ -2030,9 +2232,16 @@ function RulesetForm({
return
}
// Type is meaningless for geosite/geoip (the remote .srs is self-describing).
const rs: Ruleset = isGeoSource(source)
? { Name: nm, Source: source }
: { Name: nm, Type: type, Source: source }
//
// `Format` is CARRIED, not rebuilt: this form has no control for it, so every
// save that dropped it destroyed a value only SSH could put back — and
// renaming a list came through here. See ruleset.ts for what it costs.
const rs: Ruleset = carryRulesetFormat(
isGeoSource(source)
? { Name: nm, Source: source }
: { Name: nm, Type: type, Source: source },
initial,
)
if (source === 'inline') {
const entries = parseLines(text)
if (entries.length === 0) {
+12
View File
@@ -309,6 +309,18 @@
.set-field-ctl {
min-width: 0;
}
/* A <select> shrink-wraps to its WIDEST option, and nothing capped it: the geo
provider labels ("Auto — country codes from SagerNet, the rest from
Loyalsoldier") pushed the element to 559px inside a 390px viewport and the
whole page scrolled sideways. The labels are load-bearing — they are where
the coverage and cost difference between providers is stated — so the fix is
to cap the control, not to shorten what it says. Truncation is the browser's
job once there is a definite width to truncate against. */
.set-field-ctl select {
width: 100%;
max-width: 100%;
min-width: 0;
}
.set-dl-row {
justify-content: flex-start;
}
+277 -28
View File
@@ -2,8 +2,17 @@ import './Settings.css'
import { useCallback, useEffect, useRef, useState } from 'react'
import type { ReactNode } from 'react'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { AlertsSection } from './Alerts'
import { apply as apiApply, downloadLog, getConfig, putConfig, ApiError } from '../api'
import type { Globals, LogRange, Model } from '../api'
import { killSwitchClosed } from '../planeState'
import {
GEO_PROVIDERS,
checkGeoTemplate,
isKnownGeoProvider,
normGeoProvider,
usesCustomTemplates,
} from '../geoProvider'
// The Settings page is a thin editor over the desired-state Model's Globals —
// same save→apply split as DNS.tsx: every edit rewrites model.Globals in-place,
@@ -15,13 +24,12 @@ import type { Globals, LogRange, Model } from '../api'
// on blur/Enter after validation — an invalid value shows an inline error and is
// NOT saved.
//
// api.ts types Globals without PanelPort (it rides through the untyped Model
// passthrough), so we surface it via a local structural extension.
// ---- local Globals extension (promote to api.ts) --------------------------
/** Globals plus the admin-panel port, carried through the Model passthrough. */
type GlobalsX = Globals & { PanelPort?: number }
// There is no local extension of `Globals` here. There used to be one for
// PanelPort, with a note saying api.ts did not type the field — api.ts has typed
// it for a long time now, and the stale note was an invitation to declare the
// next field twice, in two shapes that could drift. Every key this page writes is
// on `Globals` in api.ts, and it has to be: PUT /api/config decodes with
// DisallowUnknownFields, so a key that exists only here fails the WHOLE write.
// ---- helpers ---------------------------------------------------------------
@@ -103,6 +111,30 @@ function parseDuration(raw: string): ParseResult<string> {
return { ok: true, value: s.toLowerCase() }
}
/**
* Parse a custom `{category}` geo URL template.
*
* EMPTY IS LEGAL and is not an error to refuse: it means "this source keeps using
* the built-in Auto chain", which is a state the operator is allowed to return to
* by clearing the box. A non-empty template without the placeholder is refused
* here, because the daemon will not splice the category in and would otherwise
* fetch a plausible-looking wrong URL that only fails at update time.
*/
function parseGeoTemplate(raw: string): ParseResult<string> {
const s = raw.trim()
if (s === '') return { ok: true, value: '' }
const v = checkGeoTemplate(s)
return v.ok ? { ok: true, value: s } : { ok: false, error: v.reason }
}
/** Parse an optional http(s):// URL — blank clears it. */
function parseUrlOrBlank(raw: string): ParseResult<string> {
const s = raw.trim()
if (s === '') return { ok: true, value: '' }
if (!HTTP_RE.test(s)) return { ok: false, error: 'Enter an http(s):// URL, or leave it blank.' }
return { ok: true, value: s }
}
// `none` really does silence the engine log — it is emitted as the engine's own
// log-disable switch, not as a quieter level.
const LOG_LEVELS: ReadonlyArray<{ value: string; label: string }> = [
@@ -116,10 +148,18 @@ const LOG_LEVELS: ReadonlyArray<{ value: string; label: string }> = [
// Logging/stats backend. "off" collects nothing; "memory" keeps aggregates in RAM
// (lost on restart); "sqlite" persists logs to /etc/shater/stats.db so they survive
// a restart, bounded by the retention row caps + the disk-limit knob below.
//
// THE VALUE `sqlite` IS A HISTORICAL NAME AND THE LABELS NO LONGER REPEAT IT. There
// is no SQLite in the daemon: the store is bbolt (stats/boltring.go) — pure Go, no
// CGO, already linked into the binary via experimental/cachefile — and it was chosen
// precisely to be rid of "the stop-the-world window the sqlite VACUUM used to
// impose", in that file's own words. The wire value has to stay (it is in every
// shipped config, and the daemon still matches on it); what the operator READS
// should describe where the logs go, which is the disk.
const STATS_BACKENDS: ReadonlyArray<{ value: string; label: string }> = [
{ value: 'off', label: 'Off — no logging' },
{ value: 'memory', label: 'Memory (RAM)' },
{ value: 'sqlite', label: 'SQLite · persistent' },
{ value: 'sqlite', label: 'Disk · survives a restart' },
]
// ---- page ------------------------------------------------------------------
@@ -201,15 +241,15 @@ export default function Settings() {
// ---- one setter for every Globals field -----------------------------------
const setGlobal = useCallback(
<K extends keyof GlobalsX>(key: K, value: GlobalsX[K], okMsg: string) => {
<K extends keyof Globals>(key: K, value: Globals[K], okMsg: string) => {
if (!config) return
const nextGlobals = { ...(config.Globals as GlobalsX), [key]: value }
const nextGlobals: Globals = { ...config.Globals, [key]: value }
void save({ ...config, Globals: nextGlobals }, okMsg)
},
[config, save],
)
const globals = config?.Globals as GlobalsX | undefined
const globals = config?.Globals
const busy = saving || applying
const ready = !!config
@@ -218,14 +258,16 @@ export default function Settings() {
const ringUnlimited = (globals?.StatsRingSize ?? 0) === 0
const timelineUnlimited = (globals?.StatsTimelineMinutes ?? 0) === 0
const domainsUnlimited = (globals?.StatsMaxDomains ?? 0) === 0
// SQLite disk cap: 0 = unlimited (stats.db grows with the disk).
// Disk cap: 0 = unlimited (stats.db grows with the disk).
const diskUnlimited = (globals?.StatsDiskLimitMB ?? 0) === 0
// Logging backend. Default "memory" when the field is absent (older config). When
// "off", nothing is collected, so the retention sizes below don't apply — dim them.
const statsBackend = globals?.StatsBackend || 'memory'
const loggingOff = statsBackend === 'off'
const loggingSqlite = statsBackend === 'sqlite'
// The wire value is still `sqlite` (historical — see STATS_BACKENDS); the store
// is bbolt on disk, so everything the operator reads calls it the disk backend.
const loggingDisk = statsBackend === 'sqlite'
// Retention controls are meaningless with logging off; disable them there.
const retentionDisabledCtl = busy || !ready || loggingOff
@@ -257,7 +299,29 @@ export default function Settings() {
// the Targets page. Absent ⇒ enabled (older config), so read it as `!== false`.
const groupHealthOn = globals?.GroupHealth !== false
const killSwitch = globals?.KillSwitch === 'open' ? 'open' : 'closed'
// ---- geo data --------------------------------------------------------------
// An absent GeoProvider reads as `auto`, because auto is what the router runs.
// An UNKNOWN one is preserved verbatim and marked: opening this page must not be
// an edit, and the daemon degrades it to auto with a warning rather than
// refusing to start — so the panel says which of the two is true instead of
// drawing a setting that is not in effect.
const geoProvider = normGeoProvider(globals?.GeoProvider)
const geoProviderKnown = isKnownGeoProvider(geoProvider)
const geoCustom = usesCustomTemplates(geoProvider)
const geoProviderNote =
GEO_PROVIDERS.find((p) => p.id === geoProvider)?.blurb ??
'The router has no such provider and is running Auto instead — pick one from the list.'
// "in effect" here means exactly the daemon's test: non-empty AND carrying the
// placeholder. A template failing either keeps that ONE source on the built-in
// chain; the other source is unaffected, which is why they are two verdicts.
const geoSiteTemplateOk = checkGeoTemplate(globals?.GeositeURL ?? '').ok
const geoIpTemplateOk = checkGeoTemplate(globals?.GeoipURL ?? '').ok
// Normalised the daemon's way (planeState.killSwitchClosed), not by string
// equality: `kill_switch 'OPEN'` is fail-OPEN on the router, and `=== 'open'`
// read it as closed — the panel would have drawn the protective setting over a
// router that has none.
const killSwitch = killSwitchClosed(globals?.KillSwitch) ? 'closed' : 'open'
/**
* The master switch, which is the most destructive control in the panel and was
@@ -300,10 +364,16 @@ export default function Settings() {
},
[confirm, killSwitch, setGlobal],
)
// "so nothing leaks unproxied" claimed more than the holding plane promises.
// netplane/nft.go states its own contract as "No client TRAFFIC reaches the WAN"
// and names the exception in the same paragraph: clients still reach the router's
// resolver and dnsmasq forwards those lookups to the ISP in the clear. It cannot
// be closed — blocking it would also cut the daemon's own name resolution, and
// with it any chance of recovering unattended.
const killNote =
killSwitch === 'open'
? 'Fail-open — if the engine stops, traffic falls back to the direct WAN. Stays online, but unprotected.'
: 'Fail-closed — if the engine stops, LAN→WAN is blocked so nothing leaks unproxied.'
: 'Fail-closed — if the engine stops, LAN→WAN is blocked so no traffic from your devices reaches the internet. DNS is the exception: lookups sent to the router still go out to your provider in the clear, which is what lets the router recover on its own.'
const loading = config === null && loadError === null
@@ -374,7 +444,16 @@ export default function Settings() {
<Field
label="Panel port"
note="Admin-panel port. 0 uses the default 8088. A change needs a restart to rebind."
// "needs a restart to rebind" read as a promise that the rebind
// happens. It is not one the panel can make: cmd/shaterd/main.go
// treats a failed panel listen as `logger.Warn("panel server
// unavailable (daemon continues)")` and carries on — the daemon keeps
// routing traffic and the panel simply is not there. Nothing reports
// it in the UI either, because the UI is what went missing, and
// `status.panel_port` keeps naming the CONFIGURED port regardless
// (which is also what LuCI builds its "Open panel" button from). So
// the note names the failure and where the answer actually is.
note="Admin-panel port. 0 uses the default 8088. A change takes effect on restart — and if the new port is already taken the panel does not come back at all: the daemon keeps running and only says so in its log."
>
<InlineEdit<number>
value={globals?.PanelPort ?? 0}
@@ -406,7 +485,7 @@ export default function Settings() {
<Field
label="Log level"
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level."
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level — set up where they go in the Alerts section below."
>
<Select
value={globals?.LogLevel || 'warning'}
@@ -568,6 +647,163 @@ export default function Settings() {
</Field>
</Group>
{/* ---- GEO DATA ---- */}
{/* Where every geosite/geoip list on the Routing and DNS pages fetches
from. Five real Globals fields that had no control at all: the only
way to move `geoip:us` off SagerNet was to edit the config over SSH,
while this panel used the data those settings pick. */}
<Group title="Geo data" count={geoProvider}>
<p className="set-group-note">
The catalogue behind every <strong>geosite</strong> and <strong>geoip</strong> list —
the rule-sets on Routing, the category blocklists on DNS, and the suggestions their
pickers offer. Changing it changes where those lists are downloaded from, and how big
they are: <span className="mono">netflix</span> is about 108 address prefixes, while
the country <span className="mono">us</span> is about 159 000.
</p>
<Field label="Provider" note={geoProviderNote}>
<Select
value={geoProvider}
options={GEO_PROVIDERS.map((p) => ({ value: p.id, label: p.label }))}
ariaLabel="Geo data provider"
busy={busy}
disabled={!ready}
onChange={(v) => setGlobal('GeoProvider', v, `Geo provider → ${v}`)}
/>
</Field>
{!geoProviderKnown && (
<p className="set-warn" role="status">
The router does not recognise{' '}
<span className="mono">{globals?.GeoProvider}</span> and is using{' '}
<strong>Auto</strong> instead — it warns and carries on rather than refusing to
start. Pick a provider above to make the setting real.
</p>
)}
{geoCustom && (
<>
<p className="set-group-note">
Each template must contain <span className="mono">{'{category}'}</span> exactly
once — the router splices the category in there and will not append it, because
appending would build a wrong URL that only fails at update time. Leave one blank
to keep that source on the built-in Auto chain.
</p>
<Field
label="Geosite URL template"
note="Where a domain category (youtube, category-ads-all…) is fetched from."
>
<InlineEdit<string>
value={globals?.GeositeURL ?? ''}
format={(s) => s}
parse={parseGeoTemplate}
inputMode="url"
width="26rem"
placeholder="https://mirror.example/geosite/{category}.srs"
ariaLabel="Custom geosite URL template"
busy={busy}
disabled={!ready}
onCommit={(v) =>
setGlobal(
'GeositeURL',
v,
v ? 'Geosite template saved' : 'Geosite template cleared — back to Auto',
)
}
/>
</Field>
{!geoSiteTemplateOk && (
<p className="set-warn" role="status">
No geosite template — geosite lists keep using the built-in Auto chain
(SagerNet), whatever this provider says.
</p>
)}
<Field
label="Geoip URL template"
note="Where an address category (a country code, an ASN…) is fetched from."
>
<InlineEdit<string>
value={globals?.GeoipURL ?? ''}
format={(s) => s}
parse={parseGeoTemplate}
inputMode="url"
width="26rem"
placeholder="https://mirror.example/geoip/{category}.srs"
ariaLabel="Custom geoip URL template"
busy={busy}
disabled={!ready}
onCommit={(v) =>
setGlobal(
'GeoipURL',
v,
v ? 'Geoip template saved' : 'Geoip template cleared — back to Auto',
)
}
/>
</Field>
{!geoIpTemplateOk && (
<p className="set-warn" role="status">
No geoip template — geoip lists keep using the built-in Auto chain, whatever
this provider says.
</p>
)}
<Field
label="Geosite category index"
// Suggestions only. Saying more would be a promise the daemon
// does not make: an empty or unreachable index leaves the field
// free text, exactly as it already degrades on a failed fetch.
note="A GitHub git-trees API URL listing the published files. Used only to suggest category names in the pickers — leave it blank and you simply type the category yourself."
>
<InlineEdit<string>
value={globals?.GeositeIndexURL ?? ''}
format={(s) => s}
parse={parseUrlOrBlank}
inputMode="url"
width="26rem"
placeholder="https://api.github.com/repos/…/git/trees/…"
ariaLabel="Geosite category index URL"
busy={busy}
disabled={!ready}
onCommit={(v) =>
setGlobal(
'GeositeIndexURL',
v,
v ? 'Geosite index saved' : 'Geosite index cleared — no suggestions',
)
}
/>
</Field>
<Field
label="Geoip category index"
note="The same, for address categories. Blank means no suggestions for geoip."
>
<InlineEdit<string>
value={globals?.GeoipIndexURL ?? ''}
format={(s) => s}
parse={parseUrlOrBlank}
inputMode="url"
width="26rem"
placeholder="https://api.github.com/repos/…/git/trees/…"
ariaLabel="Geoip category index URL"
busy={busy}
disabled={!ready}
onCommit={(v) =>
setGlobal(
'GeoipIndexURL',
v,
v ? 'Geoip index saved' : 'Geoip index cleared — no suggestions',
)
}
/>
</Field>
</>
)}
</Group>
{/* ---- HEALTH CHECK ---- */}
<Group title="Health check">
<p className="set-group-note">
@@ -620,6 +856,13 @@ export default function Settings() {
</Field>
</Group>
{/* ---- ALERTS ---- */}
{/* Extracted from the DNS page — out-of-band notifications belong with
the appliance-wide knobs, next to the log level whose note points
here. Renders its own section header (same plate as a Group); all
writes go through `save`, so the dirty banner and toast stay one. */}
<AlertsSection config={config} busy={busy} loading={loading} onSave={save} />
{/* ---- STATISTICS & LOGGING ---- */}
<Group
title="Statistics &amp; logging"
@@ -627,7 +870,7 @@ export default function Settings() {
>
<Field
label="Logging backend"
note="Off: collect nothing. Memory: fast, lost on restart, RAM-bounded. SQLite: survives restart, disk-bounded."
note="Off: collect nothing. Memory: fast, lost on restart, RAM-bounded. Disk: survives a restart, disk-bounded."
>
<Select
value={statsBackend}
@@ -642,7 +885,7 @@ export default function Settings() {
v === 'off'
? 'Logging off — collecting nothing'
: v === 'sqlite'
? 'Logging backend → SQLite (persistent)'
? 'Logging backend → disk (survives a restart)'
: 'Logging backend → memory',
)
}
@@ -655,13 +898,13 @@ export default function Settings() {
Insights page shows an off state. The retention limits below apply once logging
is turned back on.
</p>
) : loggingSqlite ? (
) : loggingDisk ? (
<p className="set-group-note">
Logs persist to <span className="mono">/etc/shater/stats.db</span> and survive a
restart. The <strong>entries</strong> limit below caps rows kept per log table; the{' '}
<strong>disk limit</strong> caps the whole <span className="mono">stats.db</span>{' '}
file (oldest rows are pruned to stay under it). Set any size to <strong>0</strong>{' '}
for <strong>Unlimited</strong>.
<strong>disk limit</strong> aims the whole <span className="mono">stats.db</span>{' '}
file at a size (oldest rows are deleted and the file rebuilt to stay near it). Set
any size to <strong>0</strong> for <strong>Unlimited</strong>.
</p>
) : (
<p className="set-group-note">
@@ -698,10 +941,16 @@ export default function Settings() {
</p>
)}
{loggingSqlite && (
{loggingDisk && (
<Field
label="SQLite disk limit (MB) (0 = unlimited)"
note="Hard cap on the on-disk stats.db file. A positive number is the ceiling — oldest rows are pruned and the DB vacuumed to stay under it; 0 lets it grow with the disk."
label="Disk limit (MB) (0 = unlimited)"
// Was: "oldest rows are pruned and the DB vacuumed". There is no
// SQLite and no VACUUM here — the store is bbolt, and reclaiming
// space means rebuilding the file (bbolt.Compact + atomic swap).
// The rebuild is SKIPPED when the filesystem cannot fit the
// transient second copy, so "ceiling" was a promise too: the DB
// then sits over the cap until space frees up. Both are said.
note="Target size for the on-disk stats.db file. Above it, the oldest rows are deleted and the file is rebuilt to give the space back — the rebuild needs room for a temporary second copy, so on a full disk the file stays over the limit until space frees up. 0 lets it grow with the disk."
>
<InlineEdit<number>
value={globals?.StatsDiskLimitMB ?? 0}
@@ -711,7 +960,7 @@ export default function Settings() {
inputMode="numeric"
width="9rem"
placeholder="Unlimited"
ariaLabel="SQLite disk limit in MB (0 = unlimited)"
ariaLabel="Stats database disk limit in MB (0 = unlimited)"
busy={busy}
disabled={retentionDisabledCtl}
onCommit={(v) =>
@@ -720,7 +969,7 @@ export default function Settings() {
/>
</Field>
)}
{loggingSqlite && diskUnlimited && (
{loggingDisk && diskUnlimited && (
<p className="set-warn" role="status">
Unlimited — stats.db grows with disk; set a cap (MB) to bound it.
</p>

Some files were not shown because too many files have changed in this diff Show More