Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3c7536dba0 | ||
|
|
3c92e1cbfd | ||
|
|
bcc9df9282 | ||
|
|
ee3641fe45 | ||
|
|
439f62238f | ||
|
|
d971eb85ee | ||
|
|
a0f6083e28 | ||
|
|
77369aedfe | ||
|
|
515ae6d1b7 | ||
|
|
a2ffbb1292 | ||
|
|
1746d4d0ef | ||
|
|
f86501bf77 |
+162
-315
@@ -1,36 +1,45 @@
|
||||
# Shater v0.2 — build the 4-package signed opkg feed and publish it as a rolling
|
||||
# Gitea release consumable as an `src/gz` feed.
|
||||
# Shater v0.2 — build the 4-package signed **apk** feed and publish it as
|
||||
# per-arch Gitea releases consumable as an apk repository.
|
||||
#
|
||||
# WHAT CHANGED FROM v0.1
|
||||
# v0.1 shipped 3 packages: xrayctl (SDK-compiled Go) + shater-core +
|
||||
# luci-app-shater (hand-packed data .ipk). v0.2 collapses the runtime into ONE
|
||||
# forked binary and ships 4 packages, all built the canonical SDK way:
|
||||
# WHAT WE SHIP
|
||||
# ONE forked binary plus its OpenWrt glue, 4 packages, all built the canonical
|
||||
# SDK way:
|
||||
# - shaterd PREBUILT static-musl + SPA-embedded + UPX binary. Built
|
||||
# OUT OF TREE by scripts/build-shaterd.sh (Go + Node + UPX)
|
||||
# and staged into openwrt/shaterd/files/ BEFORE the SDK
|
||||
# build; the openwrt/shaterd package just $(INSTALL_BIN)s
|
||||
# the arch-matched artifact. (arch-specific .ipk)
|
||||
# the arch-matched artifact. (arch-specific .apk)
|
||||
# - shater-core data glue, PKGARCH=all
|
||||
# - luci-app-shater LuCI thin launcher, PKGARCH=all (uses feeds/luci/luci.mk)
|
||||
# - byedpi ciadpi, C cross-compiled from source by the SDK (arch-specific)
|
||||
#
|
||||
# TARGET HARDWARE / ARCH MATRIX
|
||||
# x86_64 -> the QEMU testbed VM (generic x86-64).
|
||||
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 + BPI-R4, mediatek/filogic).
|
||||
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 mini + BPI-R4,
|
||||
# mediatek/filogic), both on 25.12 with apk-tools 3.
|
||||
# Only shaterd + byedpi are arch-specific; shater-core + luci-app-shater are
|
||||
# PKGARCH=all, so one build of each covers every device. opkg filters by
|
||||
# Architecture at install time, so a single combined feed URL serves all.
|
||||
# PKGARCH=all, so one build of each covers every device — but the RELEASES
|
||||
# are still per-arch (see the release-apk job for why).
|
||||
#
|
||||
# FEED SIGNING (opkg / usign — OpenWrt 24.10 is opkg, not apk; apk lands at 25.12)
|
||||
# The feed index (Packages) is usign-signed with the SECRET key in the Gitea
|
||||
# repo secret KEY_BUILD; routers verify it with the committed public key
|
||||
# dist/shater-feed.pub (fingerprint 5ac4b177689cb8e0). Do NOT regenerate the
|
||||
# key — that invalidates every deployed router's trust.
|
||||
# FORMAT: apk ONLY (25.12+)
|
||||
# The fleet runs OpenWrt/ImmortalWrt 25.12, where opkg is replaced by Alpine
|
||||
# apk (.apk files, binary packages.adb index, EC keys in /etc/apk/keys/). The
|
||||
# old .ipk lane was removed in 2026-07 (docs-shater/DECISIONS.md D22): no
|
||||
# device we serve has an opkg binary at all, so building and signing a second
|
||||
# feed served nobody.
|
||||
#
|
||||
# FEED SIGNING (EC / apk)
|
||||
# packages.adb is signed with the EC (prime256v1) SECRET key in the Gitea repo
|
||||
# secret KEY_APK; routers verify it with the committed public key
|
||||
# dist/shater-apk.pem (ci/gen-apk-key.sh). Do NOT regenerate the key — that
|
||||
# invalidates every deployed router's trust.
|
||||
#
|
||||
# AUTO-RELEASE
|
||||
# push a tag `vX.Y.Z` -> versioned release. workflow_dispatch / (optional) main
|
||||
# -> rolling `latest` pre-release (always-fresh feed). Publish uses the Gitea
|
||||
# API via curl (ci/gitea-release.sh) — no external action needed.
|
||||
# push a tag `vX.Y.Z` -> versioned per-arch releases `apk-vX.Y.Z-<arch>`.
|
||||
# workflow_dispatch -> rolling per-arch `apk-latest-<arch>` (always-fresh
|
||||
# feed). Publish uses the Gitea API via curl (ci/gitea-release.sh) — no
|
||||
# external action needed. NOTE: the apk release tags deliberately do NOT start
|
||||
# with `v` so publishing them cannot re-trigger this workflow's `v*` filter.
|
||||
#
|
||||
# PACKAGE VERSIONING (bug B4)
|
||||
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
|
||||
@@ -41,28 +50,13 @@
|
||||
# exported via $GITHUB_ENV):
|
||||
# tag `vX.Y.Z` -> X.Y.Z-r1
|
||||
# anything else -> <nearest tag>-r<commits since it + 1>
|
||||
# and hands them to the SDK builds as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
|
||||
# and hands them to the SDK build as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
|
||||
# $SHATER_VERSION (the same numbers, plus the short sha off-tag) is stamped
|
||||
# into the binary's constant.Version. ci/sdk-build*.sh then ASSERT that the
|
||||
# built .ipk/.apk really carry that version, so the failure can never be
|
||||
# silent again. This is also why both build jobs check out with fetch-depth: 0
|
||||
# into the binary's constant.Version. ci/sdk-build-apk.sh then ASSERTS that the
|
||||
# built .apk really carry that version, so the failure can never be silent
|
||||
# again. This is also why the build job checks out with fetch-depth: 0
|
||||
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
|
||||
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
|
||||
#
|
||||
# APK LANE (25.12+, ADDITIVE — T2)
|
||||
# The fleet is migrating to BananaWRT 25.12-mtk-vendor (= ImmortalWrt 25.12
|
||||
# base), where opkg is replaced by Alpine apk (.apk, binary packages.adb
|
||||
# index, EC keys in /etc/apk/keys/). The `build-apk` + `release-apk` jobs
|
||||
# below build the SAME 4 packages through the ImmortalWrt 25.12 apk-SDK and
|
||||
# publish PER-ARCH apk repos as releases `apk-latest-<arch>` (rolling) /
|
||||
# `apk-<tag>-<arch>` (versioned). Per-arch because apk filenames carry no
|
||||
# architecture (shaterd-0.2.0-r1.apk would collide across arches in one flat
|
||||
# release) and apk fetches packages relative to the packages.adb URL.
|
||||
# Signed with the EC key in the Gitea secret KEY_APK; trust anchor
|
||||
# dist/shater-apk.pem (ci/gen-apk-key.sh). The usign/opkg lane above is
|
||||
# UNCHANGED and keeps serving the 24.10 fleet. NOTE: the apk release tags
|
||||
# deliberately do NOT start with `v` so publishing them cannot re-trigger
|
||||
# this workflow's `v*` tag filter.
|
||||
|
||||
# CACHING (T3 — fast CI)
|
||||
# All caches use actions/cache pinned to v3.3.2: the LAST release speaking the
|
||||
@@ -80,34 +74,33 @@
|
||||
# (PKG_VERSION/PKG_HASH live there). Stale-safe: the buildroot verifies
|
||||
# PKG_HASH on every dl/ file and re-downloads on mismatch, so restore-keys
|
||||
# prefix fallback is allowed.
|
||||
# - Go module + build cache — key = hash of go.sum; shared by all 4 build
|
||||
# - Go module + build cache — key = hash of go.sum; shared by both build
|
||||
# jobs (each builds both GOARCHes).
|
||||
# - panel/node_modules — key = hash of panel/package-lock.json, exact-only
|
||||
# (a lockfile change MUST miss); on hit build-shaterd.sh gets --fast.
|
||||
# - apt .deb archives for the apk lane's debian:bookworm host-deps
|
||||
# (.cache/apt) — key = hash of ci/sdk-build-apk.sh (the apt list is in it).
|
||||
# - usign binary (.cache/tools) — static helper, fixed key.
|
||||
# - apt .deb archives for the debian:bookworm host-deps of the apk SDK
|
||||
# container (.cache/apt) — key = hash of ci/sdk-build-apk.sh (the apt list
|
||||
# is in it).
|
||||
# - SDK feeds/ git checkouts (.cache/feeds) — the single biggest recurring
|
||||
# cost: `scripts/feeds update -a` cloned base+packages+luci+routing+
|
||||
# telephony EVERY run (~7 min/job; github.com is ~1 MB/s from this
|
||||
# runner — run 51 evidence). The feeds dir is symlinked into the SDK
|
||||
# container from the workspace cache; `feeds update` on an existing clone
|
||||
# is a fast fetch+checkout of the pinned revs. Correctness-safe: update
|
||||
# always checks out feeds.conf's pins, and ci/sdk-build*.sh wipes the
|
||||
# always checks out feeds.conf's pins, and ci/sdk-build-apk.sh wipes the
|
||||
# cache + re-clones fresh if update ever fails on a cached checkout.
|
||||
# Key = lane + SDK release (shared across the two arch jobs of a lane —
|
||||
# same release pins identical feed revs; the sequential runner means the
|
||||
# second arch restores what the first saved). restore-keys lets an SDK
|
||||
# version bump start from the old clones (git fetch delta, not re-clone).
|
||||
# Key = lane + SDK release (shared across the two arch jobs — the same
|
||||
# release pins identical feed revs; the sequential runner means the second
|
||||
# arch restores what the first saved). restore-keys lets an SDK version
|
||||
# bump start from the old clones (git fetch delta, not re-clone).
|
||||
# Act_runner facts this design leans on (verified in run 51 logs):
|
||||
# - the cache backend works: restores/saves confirmed, hashFiles() works;
|
||||
# - docker images (openwrt/sdk, debian:bookworm, runner-images) live on the
|
||||
# PERSISTENT host daemon — "Image is up to date" each run, no re-download;
|
||||
# - docker images (debian:bookworm, runner-images) live on the PERSISTENT
|
||||
# host daemon — "Image is up to date" each run, no re-download;
|
||||
# - each actions/cache SAVE is followed by an exact 3-minute act_runner
|
||||
# stall (node process lingers; hit→no-save→no stall). Steady state saves
|
||||
# nothing, so adding cache entries is fine, but keys that change every
|
||||
# run (e.g. github.sha) would cost +3 min/entry/run — do NOT do that.
|
||||
|
||||
name: release
|
||||
|
||||
on:
|
||||
@@ -125,148 +118,11 @@ concurrency:
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
build:
|
||||
name: ${{ matrix.arch }}
|
||||
runs-on: ubuntu-latest
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
include:
|
||||
- { arch: x86_64, sdk: x86_64-24.10.4 } # testbed VM (generic x86-64)
|
||||
- { arch: aarch64_cortex-a53, sdk: mediatek-filogic-24.10.4 } # BPI-R3 + BPI-R4 (mediatek/filogic)
|
||||
steps:
|
||||
# fetch-depth: 0 — the package version is DERIVED from the git tag
|
||||
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
|
||||
# shallow checkout has neither tags nor ancestry, so `git describe` would
|
||||
# fail and every dispatch build would fall back to 0.0.0.
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
# scripts/build-shaterd.sh builds the engine via a go.mod
|
||||
# `replace => ./submodules/wireguard-go` (AmneziaWG fork), so that submodule
|
||||
# must be present or `go build` dies with "no such file or directory".
|
||||
# actions/checkout does not fetch submodules by default; init ONLY this one
|
||||
# (clients/apple+android are large and unused here).
|
||||
- name: Init wireguard-go submodule (awg)
|
||||
run: git submodule update --init --depth 1 submodules/wireguard-go
|
||||
|
||||
# THE version step (bug B4). One computation, used by both the binary
|
||||
# (constant.Version) and the three tag-versioned packages, exported to
|
||||
# every later step of this job:
|
||||
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
|
||||
- name: Compute version from git tag
|
||||
run: bash ci/version.sh --env >> "$GITHUB_ENV"
|
||||
|
||||
# Toolchain for scripts/build-shaterd.sh: Go (daemon), Node (Vite SPA), UPX.
|
||||
- name: Set up Go
|
||||
uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: go.mod # pins Go 1.24.7 (go.mod `go` line)
|
||||
cache: false # explicit actions/cache@v3.3.2 below (setup-go's
|
||||
# built-in cache uses the new API act_runner lacks)
|
||||
|
||||
- name: Set up Node
|
||||
uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: '20' # Vite 5 needs Node 18+; 20 LTS
|
||||
|
||||
# ---- caches (see the header comment for keys + version pin rationale) ----
|
||||
- name: Cache Go modules + build cache
|
||||
uses: actions/cache@v3.3.2
|
||||
with:
|
||||
path: |
|
||||
~/go/pkg/mod
|
||||
~/.cache/go-build
|
||||
key: go-${{ hashFiles('go.sum') }}
|
||||
restore-keys: |
|
||||
go-
|
||||
|
||||
- name: Cache panel node_modules
|
||||
id: npm-cache
|
||||
uses: actions/cache@v3.3.2
|
||||
with:
|
||||
path: panel/node_modules
|
||||
key: npm-${{ hashFiles('panel/package-lock.json') }}
|
||||
# NO restore-keys: node_modules must exactly match the lockfile;
|
||||
# on any lockfile change this misses and `npm ci` runs fresh.
|
||||
|
||||
- name: Cache SDK dl/ (package sources)
|
||||
uses: actions/cache@v3.3.2
|
||||
with:
|
||||
path: .cache/dl
|
||||
key: dl-${{ hashFiles('openwrt/*/Makefile') }}
|
||||
restore-keys: |
|
||||
dl-
|
||||
|
||||
# feeds git checkouts (see header): both 24.10.4 arch jobs share one entry
|
||||
# (same release = same feeds.conf.default pins), so derive the release
|
||||
# from the matrix sdk tag (x86_64-24.10.4 -> 24.10.4).
|
||||
- name: Compute feeds cache key
|
||||
id: feedskey
|
||||
run: echo "ver=$(echo '${{ matrix.sdk }}' | sed 's/.*-//')" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Cache SDK feeds checkouts
|
||||
uses: actions/cache@v3.3.2
|
||||
with:
|
||||
path: .cache/feeds
|
||||
key: feeds-opkg-${{ steps.feedskey.outputs.ver }}
|
||||
restore-keys: |
|
||||
feeds-opkg-
|
||||
|
||||
- name: Cache CI tools (usign)
|
||||
uses: actions/cache@v3.3.2
|
||||
with:
|
||||
path: .cache/tools
|
||||
key: tools-usign-v1
|
||||
|
||||
- name: Install UPX
|
||||
run: sudo apt-get update -qq && sudo apt-get install -y -qq upx-ucl
|
||||
|
||||
# Build the SPA-embedded, static-musl, UPX'd shaterd for BOTH arches and
|
||||
# stage dist/shaterd-<a>.upx into openwrt/shaterd/files/. MUST run before
|
||||
# the SDK package build (the openwrt/shaterd package installs the staged
|
||||
# artifact). $SHATER_VERSION (from the version step above) is stamped into
|
||||
# constant.Version, so the binary and the package agree. On an exact
|
||||
# node_modules cache hit, --fast skips the redundant `npm ci`.
|
||||
- name: Build & stage shaterd artifact
|
||||
env:
|
||||
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
|
||||
run: |
|
||||
set -eu
|
||||
FAST=""
|
||||
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
|
||||
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
|
||||
bash scripts/build-shaterd.sh $FAST
|
||||
|
||||
# Compile the 4 packages through the arch-matched OpenWrt SDK and produce a
|
||||
# signed per-arch opkg feed (Packages + Packages.gz + Packages.sig + .ipk).
|
||||
# SHATER_PKG_VERSION/SHATER_PKG_RELEASE reach the package Makefiles through
|
||||
# the SDK container; ci/sdk-build.sh asserts the .ipk really carry them.
|
||||
- name: Build signed feed (SDK)
|
||||
env:
|
||||
KEY_BUILD: ${{ secrets.KEY_BUILD }}
|
||||
run: bash ci/build-feed.sh "${{ matrix.arch }}" "${{ matrix.sdk }}" "out/${{ matrix.arch }}"
|
||||
|
||||
- name: Show feed
|
||||
run: ls -l "out/${{ matrix.arch }}" && cat "out/${{ matrix.arch }}/Packages"
|
||||
|
||||
- name: Upload feed artifact
|
||||
# v4 uses an artifact backend Gitea Actions does not implement
|
||||
# (GHESNotSupportedError); v3 works on Gitea's act_runner.
|
||||
uses: actions/upload-artifact@v3
|
||||
with:
|
||||
name: shater-${{ matrix.arch }}
|
||||
path: out/${{ matrix.arch }}/*
|
||||
if-no-files-found: error
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# APK lane (additive): the same 4 packages through the ImmortalWrt 25.12
|
||||
# apk-SDK for the 25.12/apk fleet (BananaWRT 25.12-mtk-vendor routers + the
|
||||
# future 25.12 VM). Produces a per-arch apk repo dir: *.apk + EC-signed
|
||||
# packages.adb + shater-apk.pem. Artifact prefix `apkfeed-` (NOT `shater-`)
|
||||
# so the opkg release job's `artifacts/shater-*` glob never picks these up.
|
||||
# Build the 4 packages through the ImmortalWrt 25.12 apk-SDK for the 25.12/apk
|
||||
# fleet (BPI-R3 mini on BananaWRT 25.12-mtk-vendor, BPI-R4 on OpenWrt 25.12,
|
||||
# and the testbed VM). Produces a per-arch apk repo dir: *.apk + EC-signed
|
||||
# packages.adb + shater-apk.pem, uploaded as the artifact `apkfeed-<arch>`.
|
||||
build-apk:
|
||||
name: apk ${{ matrix.arch }}
|
||||
runs-on: ubuntu-latest
|
||||
@@ -281,8 +137,10 @@ jobs:
|
||||
- arch: aarch64_cortex-a53 # BPI-R3 mini (BananaWRT 25.12-mtk-vendor) + BPI-R4
|
||||
sdk_url: https://downloads.immortalwrt.org/releases/25.12.1/targets/mediatek/filogic/immortalwrt-sdk-25.12.1-mediatek-filogic_gcc-14.3.0_musl.Linux-x86_64.tar.zst
|
||||
steps:
|
||||
# fetch-depth: 0 — see the opkg lane: the package version comes from
|
||||
# `git describe`, which needs tags + ancestry.
|
||||
# fetch-depth: 0 — the package version is DERIVED from the git tag
|
||||
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
|
||||
# shallow checkout has neither tags nor ancestry, so `git describe` would
|
||||
# fail and every dispatch build would fall back to 0.0.0.
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
@@ -295,8 +153,10 @@ jobs:
|
||||
- name: Init wireguard-go submodule (awg)
|
||||
run: git submodule update --init --depth 1 submodules/wireguard-go
|
||||
|
||||
# Same single version computation as the opkg lane — both lanes MUST agree
|
||||
# on the version, they package the identical tree.
|
||||
# THE version step (bug B4). One computation, used by both the binary
|
||||
# (constant.Version) and the three tag-versioned packages, exported to
|
||||
# every later step of this job:
|
||||
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
|
||||
- name: Compute version from git tag
|
||||
run: bash ci/version.sh --env >> "$GITHUB_ENV"
|
||||
|
||||
@@ -373,11 +233,27 @@ jobs:
|
||||
restore-keys: |
|
||||
feeds-apk-
|
||||
|
||||
# D23 — the shipped tag set is a TRIMMED subset (scripts/router-tags.sh);
|
||||
# everything else in CI builds with the full upstream set, so without this
|
||||
# step the one combination we actually ship is never exercised. That is how
|
||||
# `with_gvisor` was trimmed while `with_wireguard` stayed and every shipped
|
||||
# binary answered a WireGuard node with "gVisor is not included in this
|
||||
# build" (2026-07-25). The check runs the declared-feature/tag comparison
|
||||
# and then constructs one node of every declared protocol through box.New
|
||||
# UNDER THE SHIPPED TAGS. It runs before the artifact build so a tag trim
|
||||
# that breaks a feature fails the release instead of shipping.
|
||||
- name: Verify the shipped build-tag set (D23)
|
||||
run: bash scripts/check-router-tags.sh
|
||||
|
||||
- name: Install UPX
|
||||
run: sudo apt-get update -qq && sudo apt-get install -y -qq upx-ucl
|
||||
|
||||
# Same artifact-order contract as the opkg lane: the SPA-embedded shaterd
|
||||
# binary is built OUT of the SDK and staged before the package build.
|
||||
# Artifact-order contract: the SPA-embedded shaterd binary is built OUT of
|
||||
# the SDK and staged into openwrt/shaterd/files/ BEFORE the package build
|
||||
# (the openwrt/shaterd package only installs the staged artifact).
|
||||
# $SHATER_VERSION (from the version step above) is stamped into
|
||||
# constant.Version, so the binary and the package agree. On an exact
|
||||
# node_modules cache hit, --fast skips the redundant `npm ci`.
|
||||
- name: Build & stage shaterd artifact
|
||||
env:
|
||||
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
|
||||
@@ -409,117 +285,21 @@ jobs:
|
||||
if-no-files-found: error
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Publish once both arches are built. Rolling `latest` on dispatch, a versioned
|
||||
# release on a `vX.Y.Z` tag. Self-contained (curl -> Gitea API).
|
||||
release:
|
||||
name: release
|
||||
needs: build
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Download all arch feeds
|
||||
uses: actions/download-artifact@v3
|
||||
with:
|
||||
path: artifacts
|
||||
|
||||
- name: Assemble release assets
|
||||
id: assets
|
||||
run: |
|
||||
set -eu
|
||||
mkdir -p release
|
||||
# For each downloaded arch feed: one ready-to-serve tarball + loose ipks.
|
||||
for d in artifacts/shater-*; do
|
||||
[ -d "$d" ] || continue
|
||||
arch="${d#artifacts/shater-}"
|
||||
tar -C "$d" -czf "release/shater-feed-${arch}.tar.gz" .
|
||||
# loose .ipk for direct `opkg install <url>` (dedupe shared _all ipks by name)
|
||||
for ipk in "$d"/*.ipk; do
|
||||
[ -e "$ipk" ] || continue
|
||||
cp -n "$ipk" "release/$(basename "$ipk")"
|
||||
done
|
||||
done
|
||||
# ship the feed's public key so routers can verify (see docs-shater/INSTALL.md)
|
||||
cp -f dist/shater-feed.pub release/shater-feed.pub
|
||||
ls -l release
|
||||
echo "count=$(ls release | wc -l)" >> "$GITHUB_OUTPUT"
|
||||
|
||||
# restore the prebuilt usign binary (skips apt + cmake + clone + build)
|
||||
- name: Cache CI tools (usign)
|
||||
uses: actions/cache@v3.3.2
|
||||
with:
|
||||
path: .cache/tools
|
||||
key: tools-usign-v1
|
||||
|
||||
- name: Install usign (feed signer)
|
||||
run: bash ci/install-usign.sh
|
||||
|
||||
- name: Build & sign combined opkg feed index
|
||||
# One Packages/Packages.gz over ALL loose .ipk (every arch + arch=all),
|
||||
# with basename Filenames. opkg filters by Architecture, so a single
|
||||
# release URL serves every device: BPI routers pick aarch64_cortex-a53 +
|
||||
# all, the x86 testbed picks x86_64 + all. Signed with KEY_BUILD so
|
||||
# routers keep check_signature on. This is what makes the release directly
|
||||
# consumable as an `src/gz` feed (see docs-shater/INSTALL.md).
|
||||
env:
|
||||
KEY_BUILD: ${{ secrets.KEY_BUILD }}
|
||||
run: bash ci/make-index.sh release
|
||||
|
||||
- name: Determine release identity
|
||||
id: rel
|
||||
run: |
|
||||
set -eu
|
||||
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
|
||||
echo "tag=${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
|
||||
echo "name=shater ${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
|
||||
echo "prerelease=false" >> "$GITHUB_OUTPUT"
|
||||
echo "rolling=false" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "tag=latest" >> "$GITHUB_OUTPUT"
|
||||
echo "name=shater latest (main)" >> "$GITHUB_OUTPUT"
|
||||
echo "prerelease=true" >> "$GITHUB_OUTPUT"
|
||||
echo "rolling=true" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
- name: Publish Gitea release
|
||||
env:
|
||||
TOKEN: ${{ secrets.RELEASE_TOKEN != '' && secrets.RELEASE_TOKEN || github.token }}
|
||||
TAG: ${{ steps.rel.outputs.tag }}
|
||||
NAME: ${{ steps.rel.outputs.name }}
|
||||
PRERELEASE: ${{ steps.rel.outputs.prerelease }}
|
||||
ROLLING: ${{ steps.rel.outputs.rolling }}
|
||||
BODY: |
|
||||
Automated build. Packages: shaterd + byedpi (per-arch), shater-core +
|
||||
luci-app-shater (arch=all).
|
||||
Targets: x86_64 (testbed) and aarch64_cortex-a53 (BPI-R3 + BPI-R4, mediatek/filogic).
|
||||
|
||||
── Add as an opkg feed (recommended — then updating is one command) ──
|
||||
This release is itself a SIGNED package feed; opkg filters by
|
||||
architecture, so the same lines work on every device:
|
||||
wget -O /etc/opkg/keys/5ac4b177689cb8e0 https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
|
||||
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" >> /etc/opkg/customfeeds.conf
|
||||
opkg update
|
||||
opkg install luci-app-shater # pulls shater-core + shaterd too
|
||||
The public-key install is one-time; after it, `opkg update/upgrade`
|
||||
verify the signature with check_signature left on. Full guide: docs-shater/INSTALL.md.
|
||||
|
||||
── Update (name our packages — never a bare `opkg upgrade`) ──
|
||||
opkg update
|
||||
opkg upgrade shaterd shater-core luci-app-shater byedpi
|
||||
|
||||
── Or install the loose .ipk directly / from the tarball feed ──
|
||||
wget -O /tmp/f.tgz <this release>/shater-feed-aarch64_cortex-a53.tar.gz
|
||||
mkdir -p /tmp/shater && tar -C /tmp/shater -xzf /tmp/f.tgz
|
||||
opkg install /tmp/shater/luci-app-shater_*_all.ipk
|
||||
run: bash ci/gitea-release.sh release/*
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Publish the apk lane: ONE release PER ARCH (apk package filenames carry no
|
||||
# arch, and apk fetches `<name>-<ver>.apk` relative to the packages.adb URL —
|
||||
# a flat multi-arch release would collide). Rolling `apk-latest-<arch>` on
|
||||
# dispatch, `apk-<tag>-<arch>` on a version tag. The tags do NOT match the
|
||||
# workflow's `v*` trigger, so publishing them cannot re-trigger the build.
|
||||
# Publish: ONE release PER ARCH (apk package filenames carry no arch, and apk
|
||||
# fetches `<name>-<ver>.apk` relative to the packages.adb URL — a flat
|
||||
# multi-arch release would collide). Every run refreshes the ROLLING pointer
|
||||
# `apk-latest-<arch>`; a `vX.Y.Z` tag run ALSO publishes the pinnable
|
||||
# `apk-vX.Y.Z-<arch>`. The tags do NOT match the workflow's `v*` trigger, so
|
||||
# publishing them cannot re-trigger the build.
|
||||
#
|
||||
# WHY THE ROLLING RELEASE IS PUBLISHED ON TAG RUNS TOO (fixed 2026-07-25):
|
||||
# it used to be an either/or — `TAG=apk-latest-<arch>` on dispatch, ELSE
|
||||
# `TAG=apk-<ver>-<arch>` — so once releases moved to tag pushes the rolling
|
||||
# pointer was never written again. It froze at 0.2.0 (published 2026-07-24)
|
||||
# while v0.2.9/v0.2.10 published fine, and every router whose
|
||||
# /etc/apk/repositories.d/shater.list points at the rolling URL kept getting a
|
||||
# successful, silent `apk update` with nothing new. Rolling is the whole point
|
||||
# of that URL, so it is now written unconditionally and asserted afterwards.
|
||||
release-apk:
|
||||
name: release apk
|
||||
needs: build-apk
|
||||
@@ -537,6 +317,9 @@ jobs:
|
||||
with:
|
||||
path: artifacts
|
||||
|
||||
# Identity of the VERSIONED release only. The rolling pointer is published
|
||||
# on every run with fixed prerelease=true/rolling=true, so it needs nothing
|
||||
# from here.
|
||||
- name: Determine release identity
|
||||
id: rel
|
||||
run: |
|
||||
@@ -558,21 +341,39 @@ jobs:
|
||||
PRERELEASE: ${{ steps.rel.outputs.prerelease }}
|
||||
ROLLING: ${{ steps.rel.outputs.rolling }}
|
||||
run: |
|
||||
set -eu
|
||||
set -euo pipefail
|
||||
for d in artifacts/apkfeed-*; do
|
||||
[ -d "$d" ] || continue
|
||||
arch="${d#artifacts/apkfeed-}"
|
||||
if [ "$VER" = latest ]; then TAG="apk-latest-$arch"; else TAG="apk-$VER-$arch"; fi
|
||||
ROLL="apk-latest-$arch"
|
||||
|
||||
# The version we just built, read straight off the artifact
|
||||
# (`shaterd-<ver>-r<rel>.apk`). NOT recomputed with ci/version.sh:
|
||||
# this job checks out shallow, so it has no tags to describe from.
|
||||
pkg=""
|
||||
for a in "$d"/shaterd-*.apk; do
|
||||
if [ -f "$a" ]; then pkg="$(basename "$a")"; fi
|
||||
done
|
||||
[ -n "$pkg" ] || { echo "[release-apk] ERROR: no shaterd-*.apk in $d"; exit 11; }
|
||||
want="${pkg#shaterd-}"; want="${want%.apk}"
|
||||
echo "[release-apk] arch=$arch built version=$want"
|
||||
|
||||
BODY="Automated apk (OpenWrt/ImmortalWrt 25.12+) package repo for \`$arch\`.
|
||||
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
|
||||
This build: \`$want\`.
|
||||
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
|
||||
|
||||
── Add as an apk repository ──
|
||||
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$TAG/shater-apk.pem
|
||||
── Add as an apk repository (rolling — install once, then just update) ──
|
||||
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$ROLL/shater-apk.pem
|
||||
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
|
||||
apk update
|
||||
apk add luci-app-shater # pulls shater-core + shaterd too
|
||||
apk add byedpi # optional: ByeDPI desync egress
|
||||
\`apk-latest-<arch>\` is a MOVING pointer: every release run replaces its
|
||||
assets, so the same repo line keeps serving the newest build. To pin a
|
||||
version instead, point the repo line at
|
||||
\`.../download/apk-vX.Y.Z-\$(cat /etc/apk/arch)/packages.adb\` — then the
|
||||
file must be edited by hand for each upgrade.
|
||||
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
|
||||
apk update
|
||||
apk upgrade shaterd shater-core luci-app-shater byedpi
|
||||
@@ -580,9 +381,55 @@ jobs:
|
||||
configured repo and can downgrade unrelated system packages; naming them
|
||||
upgrades only those (apk-tools 3: \"If list of packages is provided, only
|
||||
those packages are upgraded along with needed dependencies\").
|
||||
Full guide: docs-shater/INSTALL.md §6. The opkg/24.10 feed lives in the \`latest\` release."
|
||||
echo "[release-apk] publishing $TAG from $d"
|
||||
TAG="$TAG" NAME="shater apk $VER ($arch)" BODY="$BODY" \
|
||||
PRERELEASE="$PRERELEASE" ROLLING="$ROLLING" \
|
||||
Full guide: docs-shater/INSTALL.md §5."
|
||||
|
||||
# 1) the pinnable versioned release (tag runs only)
|
||||
if [ "$VER" != latest ]; then
|
||||
echo "[release-apk] publishing apk-$VER-$arch from $d"
|
||||
TAG="apk-$VER-$arch" NAME="shater apk $VER ($arch)" BODY="$BODY" \
|
||||
PRERELEASE="$PRERELEASE" ROLLING="$ROLLING" \
|
||||
bash ci/gitea-release.sh "$d"/*
|
||||
fi
|
||||
|
||||
# 2) the rolling pointer — ALWAYS, tag run included. ci/gitea-release.sh
|
||||
# deletes the existing release before recreating it, so the old
|
||||
# version's assets are REPLACED, never accumulated (two versions of
|
||||
# one package in one index would let apk choose, not us).
|
||||
echo "[release-apk] publishing $ROLL from $d"
|
||||
TAG="$ROLL" NAME="shater apk latest ($arch)" BODY="$BODY" \
|
||||
PRERELEASE=true ROLLING=true \
|
||||
bash ci/gitea-release.sh "$d"/*
|
||||
|
||||
# 3) ASSERT the rolling release really serves THIS build — same class
|
||||
# of check as ci/sdk-build-apk.sh's package-version assert, and for
|
||||
# the same reason: the previous failure mode was silent. Reads the
|
||||
# published release back over the API and requires our three
|
||||
# tag-versioned packages at $want, the index, the key — and NO
|
||||
# left-over package asset at any other version.
|
||||
api="$GITHUB_SERVER_URL/api/v1/repos/$GITHUB_REPOSITORY/releases/tags/$ROLL"
|
||||
got="$(curl -fsS -H "Authorization: token $TOKEN" "$api" \
|
||||
| tr '{},' '\n\n\n' \
|
||||
| sed -n 's/.*"name"[[:space:]]*:[[:space:]]*"\([^"]*\)".*/\1/p' | sort -u)" || {
|
||||
echo "[release-apk] ERROR: cannot read back $ROLL from the API"; exit 12; }
|
||||
echo "[release-apk] $ROLL assets: $(printf '%s ' $got)"
|
||||
# here-string, NOT `printf | grep -q`: under `pipefail` the early
|
||||
# exit of grep -q can SIGPIPE the writer and fail a passing check.
|
||||
for f in "shaterd-$want.apk" "shater-core-$want.apk" \
|
||||
"luci-app-shater-$want.apk" packages.adb shater-apk.pem; do
|
||||
grep -qxF "$f" <<<"$got" || {
|
||||
echo "[release-apk] ERROR: $ROLL does not contain '$f' after publish."
|
||||
echo " A router pinned to the rolling URL would have silently"
|
||||
echo " stayed on its old version with a successful apk update."
|
||||
exit 13; }
|
||||
done
|
||||
stale="$(grep -E '^(shaterd|shater-core|luci-app-shater)-.*\.apk$' <<<"$got" \
|
||||
| grep -vxF -e "shaterd-$want.apk" -e "shater-core-$want.apk" \
|
||||
-e "luci-app-shater-$want.apk" || true)"
|
||||
[ -z "$stale" ] || {
|
||||
echo "[release-apk] ERROR: $ROLL still holds stale package assets:"
|
||||
printf ' %s\n' $stale
|
||||
echo " Two versions of one package in one feed = apk picks by its"
|
||||
echo " own rules, not by our intent."
|
||||
exit 14; }
|
||||
echo "[release-apk] OK — $ROLL serves $want"
|
||||
done
|
||||
|
||||
+1
-1
@@ -63,7 +63,7 @@ nul
|
||||
/venv/
|
||||
/test/cache.db
|
||||
|
||||
# feed artifacts (tracked public key dist/shater-feed.pub is force-added)
|
||||
# feed artifacts (the tracked apk trust anchor dist/shater-apk.pem is force-added)
|
||||
/dist/
|
||||
|
||||
# local agent config (CLAUDE.md is deliberately tracked; .claude local settings are not)
|
||||
|
||||
+16
-22
@@ -49,29 +49,22 @@ Full list with MVP/T1/T2 tags — [`docs-shater/FEATURES.md`](docs-shater/FEATUR
|
||||
|
||||
## Install
|
||||
|
||||
Two signed feeds. Pick by the router's OpenWrt version. Verbatim commands and the
|
||||
manual `.ipk`/`.apk` install are in [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
|
||||
|
||||
**opkg (OpenWrt 24.10):**
|
||||
One signed **apk** feed (OpenWrt / ImmortalWrt / BananaWRT **25.12+**), one
|
||||
release per arch. Verbatim commands, the manual `.apk` install and the
|
||||
rolling-vs-pinned choice are in
|
||||
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
|
||||
|
||||
```sh
|
||||
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
|
||||
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
|
||||
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
|
||||
>> /etc/opkg/customfeeds.conf
|
||||
opkg update && opkg install luci-app-shater # -> shater-core -> shaterd
|
||||
```
|
||||
|
||||
**apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+):**
|
||||
|
||||
```sh
|
||||
wget -O /etc/apk/keys/shater-apk.pem \
|
||||
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
|
||||
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
|
||||
> /etc/apk/repositories.d/shater.list
|
||||
wget -O /etc/apk/keys/shater-apk.pem "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
|
||||
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" > /etc/apk/repositories.d/shater.list
|
||||
apk update && apk add luci-app-shater # -> shater-core -> shaterd
|
||||
```
|
||||
|
||||
`apk-latest-<arch>` is a moving pointer refreshed by every release run — install
|
||||
once and `apk update && apk upgrade shaterd shater-core luci-app-shater byedpi`
|
||||
keeps the router current. Point the repo line at `apk-vX.Y.Z-<arch>` instead to
|
||||
pin a build; that file then has to be edited by hand for every upgrade.
|
||||
|
||||
shater ships **inert** (globals off) so install never breaks connectivity. After
|
||||
configuring nodes/rules: `uci set shater.globals.enabled=1 && uci commit shater`,
|
||||
then `shaterd apply` and `shaterd confirm`.
|
||||
@@ -91,16 +84,17 @@ into `openwrt/shaterd/files/`. Details in
|
||||
| `panel/` | Admin SPA (Vite + React + TS) and its Go server |
|
||||
| `openwrt/` | Packages: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
|
||||
| `docs-shater/` | Product documentation |
|
||||
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, feed/release scripts, CI |
|
||||
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, apk feed/release scripts, CI |
|
||||
| `SPECS/`, `docs-lx/` | Engine-fork constitution/specs and feature-config reference |
|
||||
| `docs/`, `mkdocs.yml` | **Upstream** sing-box docs (mkdocs) — kept as-is |
|
||||
| `adapter/ cmd/ dns/ route/ option/ protocol/ transport/ …` | sing-box-lx engine tree |
|
||||
|
||||
## CI, upstream & license
|
||||
|
||||
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes signed
|
||||
feeds: opkg (usign, key `5ac4b177689cb8e0`) and apk (EC key `shater-apk.pem`). A
|
||||
`vX.Y.Z` tag → versioned release; `workflow_dispatch` → rolling `latest`.
|
||||
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes a signed
|
||||
per-arch apk repo (EC key `shater-apk.pem`). A `vX.Y.Z` tag → the pinnable
|
||||
`apk-vX.Y.Z-<arch>`; every run also refreshes the rolling `apk-latest-<arch>` and
|
||||
asserts over the API that it really serves the version just built.
|
||||
|
||||
The engine is the **sing-box-lx** fork — a thin downstream of upstream sing-box that
|
||||
lives by **rebase, never merge**; its constitution is
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
|
||||
[](LICENSE)
|
||||

|
||||

|
||||

|
||||
|
||||
---
|
||||
|
||||
@@ -128,42 +128,15 @@ data-plane, DNS-flow, apply-flow) — в [`docs-shater/ARCHITECTURE.md`](docs-sh
|
||||
|
||||
## Установка
|
||||
|
||||
shater поставляется двумя подписанными фидами. Выберите по версии OpenWrt на роутере:
|
||||
|
||||
- **OpenWrt 24.10** → фид **opkg** (`.ipk`, `Packages.gz`, ключ usign).
|
||||
- **OpenWrt / ImmortalWrt / BananaWRT 25.12+** → фид **apk** (`.apk`, `packages.adb`,
|
||||
EC-ключ).
|
||||
shater поставляется одним подписанным **apk-фидом** (OpenWrt / ImmortalWrt /
|
||||
BananaWRT **25.12+**: `.apk`, индекс `packages.adb`, EC-ключ в `/etc/apk/keys/`).
|
||||
Старый opkg-фид (`.ipk`, 24.10) снят — оба наших роутера на 25.12 с apk-tools 3,
|
||||
бинаря `opkg` там просто нет (`docs-shater/DECISIONS.md` D22).
|
||||
|
||||
Пакеты ставятся по зависимостям: `shaterd` → `shater-core` → `luci-app-shater`
|
||||
(+ опциональный `byedpi`). `shaterd` подтягивается автоматически как зависимость.
|
||||
|
||||
### Путь A — фид opkg (OpenWrt 24.10)
|
||||
|
||||
```sh
|
||||
# 1) доверяем ключу фида — ИМЯ файла обязано равняться отпечатку usign-ключа.
|
||||
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
|
||||
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
|
||||
|
||||
# 2) добавляем фид (один URL обслуживает все арки).
|
||||
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
|
||||
>> /etc/opkg/customfeeds.conf
|
||||
|
||||
# 3) обновляемся и ставим (shaterd подтянется как зависимость).
|
||||
opkg update
|
||||
opkg install luci-app-shater # -> shater-core -> shaterd
|
||||
opkg install byedpi # опционально: ByeDPI desync-egress
|
||||
```
|
||||
|
||||
Обновление — **только наши пакеты, никогда голый `opkg upgrade`** (без аргументов
|
||||
он тянет обновления и на системные пакеты, это классический способ окирпичить
|
||||
роутер):
|
||||
|
||||
```sh
|
||||
opkg update
|
||||
opkg upgrade shaterd shater-core luci-app-shater byedpi
|
||||
```
|
||||
|
||||
### Путь B — фид apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+)
|
||||
### Фид apk
|
||||
|
||||
`/etc/apk/arch` сам выбирает нужный per-arch релиз (apk-релизы раздельны по арке):
|
||||
|
||||
@@ -198,13 +171,22 @@ apk upgrade shaterd shater-core luci-app-shater byedpi
|
||||
only those packages are upgraded along with needed dependencies»*. Проверить
|
||||
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
|
||||
|
||||
> **Роллинг или фиксация — это выбор URL в `shater.list`.** `apk-latest-<arch>`
|
||||
> — движущийся указатель: каждый релизный прогон заменяет его ассеты, поэтому
|
||||
> «поставил и забыл»: `apk update` сам видит новую сборку. `apk-vX.Y.Z-<arch>` —
|
||||
> фиксация на конкретной сборке: роутер не получит ничего нового, пока
|
||||
> `/etc/apk/repositories.d/shater.list` не отредактируют руками — на каждом
|
||||
> роутере и на каждый релиз. На `mini_router` сознательно прописан
|
||||
> версионированный URL, и ручная правка — его цена. Подробнее —
|
||||
> [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) §5.1.
|
||||
|
||||
> Версии пакетов CI берёт из git-тега (`vX.Y.Z` → `X.Y.Z-r1`, сборка вне тега →
|
||||
> `X.Y.Z-r<коммитов+1>`), поэтому каждая новая сборка действительно видна
|
||||
> менеджеру пакетов как новая. Подробности — `docs-shater/INSTALL.md` §2.1.
|
||||
|
||||
> Полные инструкции — раздельная установка из `.ipk`/`.apk` вручную, закрепление
|
||||
> версии (`vX.Y.Z` / `apk-vX.Y.Z-<arch>`), совместимость с BananaWRT
|
||||
> `25.12-mtk-vendor` — в [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
|
||||
> Полные инструкции — ручная установка из `.apk`, фиксация версии
|
||||
> (`apk-vX.Y.Z-<arch>`), совместимость с BananaWRT `25.12-mtk-vendor` — в
|
||||
> [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
|
||||
|
||||
### Включение
|
||||
|
||||
@@ -258,8 +240,8 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
|
||||
| `openwrt/` | Пакеты: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
|
||||
| `docs-shater/` | Документация продукта (см. таблицу ниже) |
|
||||
| `scripts/` | `build-shaterd.sh` — сборка ship-артефакта |
|
||||
| `ci/` | Скрипты сборки фидов и релизов (SDK, usign/EC, Gitea API) |
|
||||
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанные фиды opkg/apk |
|
||||
| `ci/` | Скрипты сборки apk-фида и релизов (SDK, EC-подпись, Gitea API) |
|
||||
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанный apk-фид |
|
||||
| `SPECS/` | Конституция форка движка и спеки (Spec Kit) |
|
||||
| `docs-lx/` | Справочник конфигурации фич движка (`lx-config.md`, `.ru.md`) |
|
||||
| `lx-test/`, `submodules/` | Примеры конфигов движка и submodule AmneziaWG-рантайма |
|
||||
@@ -273,16 +255,17 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
|
||||
CI на **Gitea Actions** (`.gitea/workflows/release.yml`) собирает все 4 пакета и
|
||||
публикует **подписанные фиды**:
|
||||
|
||||
- **opkg (24.10):** один комбинированный релиз, подписан usign-ключом (публичный
|
||||
`dist/shater-feed.pub`, отпечаток `5ac4b177689cb8e0`; секрет — в Gitea-secret
|
||||
`KEY_BUILD`).
|
||||
- **apk (25.12+):** параллельная линия, **по релизу на арку**, подписан EC-ключом
|
||||
(`dist/shater-apk.pem`; секрет — `KEY_APK`).
|
||||
- **apk (25.12+)** — единственный формат: **по релизу на арку**, индекс
|
||||
`packages.adb` подписан EC-ключом (публичный `dist/shater-apk.pem`; секрет — в
|
||||
Gitea-secret `KEY_APK`).
|
||||
|
||||
Триггеры: push тега **`vX.Y.Z`** → версионный релиз; `workflow_dispatch` →
|
||||
плавающий `latest`/`apk-latest-<arch>` (всегда свежий фид). Публикация — через
|
||||
Gitea API (`ci/gitea-release.sh`). Ключи **никогда не перегенерируются** — это
|
||||
инвалидировало бы доверие на всех развёрнутых роутерах.
|
||||
Триггеры: push тега **`vX.Y.Z`** → версионный релиз `apk-vX.Y.Z-<arch>`;
|
||||
`workflow_dispatch` → только роллинг. Роллинг `apk-latest-<arch>` обновляется
|
||||
**на каждом прогоне**, включая теговый, и после публикации проверяется через API:
|
||||
в нём обязаны лежать наши три пакета ровно собранной версии и ни одного ассета
|
||||
другой версии. Публикация — через Gitea API (`ci/gitea-release.sh`). Ключ
|
||||
**никогда не перегенерируется** — это инвалидировало бы доверие на всех
|
||||
развёрнутых роутерах.
|
||||
|
||||
---
|
||||
|
||||
@@ -307,7 +290,7 @@ build-тегами и живущий **ребейзом на каждый upstre
|
||||
| Документ | О чём |
|
||||
|----------|-------|
|
||||
| [`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, testbed/инфра |
|
||||
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка обоих фидов (opkg/apk) |
|
||||
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка apk-фида (роллинг/фиксация) |
|
||||
| [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md) | One-binary дизайн, auth-handoff, data/DNS/apply-потоки (диаграммы) |
|
||||
| [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
|
||||
| [`docs-shater/ROADMAP.md`](docs-shater/ROADMAP.md) | Фазовый план |
|
||||
|
||||
+19
-20
@@ -1,6 +1,6 @@
|
||||
#!/bin/sh
|
||||
# ci/build-feed-apk.sh — build the signed **apk** feed for ONE arch (the 25.12
|
||||
# lane — additive next to ci/build-feed.sh, which stays the opkg/24.10 lane).
|
||||
# ci/build-feed-apk.sh — build the signed **apk** feed for ONE arch (25.12+;
|
||||
# the only packaging lane shater has — see docs-shater/DECISIONS.md D22).
|
||||
#
|
||||
# Usage: ci/build-feed-apk.sh <ARCH> <SDK_URL> <OUTDIR>
|
||||
# e.g. ci/build-feed-apk.sh aarch64_cortex-a53 \
|
||||
@@ -10,14 +10,14 @@
|
||||
# This is the per-arch entrypoint the Gitea workflow's `build-apk` job calls.
|
||||
# It runs on the CI RUNNER and:
|
||||
# 1. asserts the prebuilt shaterd binary for this arch was already staged by
|
||||
# scripts/build-shaterd.sh (same artifact-order contract as the opkg lane);
|
||||
# 2. drives a plain `debian:bookworm` container (workspace shared via
|
||||
# `--volumes-from`, same trick as ci/build-feed.sh) that downloads the
|
||||
# ImmortalWrt 25.12 apk-SDK tarball and runs ci/sdk-build-apk.sh in it:
|
||||
# compile the 4 packages as .apk, then `apk mkndx --sign` the per-arch
|
||||
# `packages.adb` index. Unlike the usign lane (index signed on the runner),
|
||||
# apk indexing NEEDS the SDK's host `apk` tool, so index+sign happen inside
|
||||
# the container.
|
||||
# scripts/build-shaterd.sh (the artifact-order contract);
|
||||
# 2. drives a plain `debian:bookworm` container (the job's workspace volume is
|
||||
# shared into it with `--volumes-from $(hostname)`; a bare `-v $PWD:...`
|
||||
# points at a host path that does not exist under act_runner's DinD) that
|
||||
# downloads the ImmortalWrt 25.12 apk-SDK tarball and runs
|
||||
# ci/sdk-build-apk.sh in it: compile the 4 packages as .apk, then
|
||||
# `apk mkndx --sign` the per-arch `packages.adb` index. Indexing NEEDS the
|
||||
# SDK's host `apk` tool, so index+sign happen inside the container.
|
||||
#
|
||||
# Why the ImmortalWrt SDK (not openwrt/sdk images): the 25.12 fleet runs
|
||||
# BananaWRT 25.12-mtk-vendor = ImmortalWrt 25.12 base (target mediatek/filogic,
|
||||
@@ -25,9 +25,9 @@
|
||||
# mediatek-filogic 25.12 tag — hence the official SDK tarball.
|
||||
#
|
||||
# Env:
|
||||
# KEY_APK EC (prime256v1) PRIVATE key PEM (Gitea repo secret — the apk analog
|
||||
# of KEY_BUILD). If set, packages.adb carries an embedded signature
|
||||
# verifiable by dist/shater-apk.pem (routers: /etc/apk/keys/).
|
||||
# KEY_APK EC (prime256v1) PRIVATE key PEM (Gitea repo secret). If set,
|
||||
# packages.adb carries an embedded signature verifiable by
|
||||
# dist/shater-apk.pem (routers: /etc/apk/keys/).
|
||||
# If unset, an UNSIGNED index is produced (warning; not shippable —
|
||||
# apk signatures are effectively mandatory).
|
||||
set -eu
|
||||
@@ -56,10 +56,9 @@ fi
|
||||
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
|
||||
|
||||
# --- 0.4) package version from the git tag ------------------------------------
|
||||
# Same contract as the opkg lane (ci/build-feed.sh): the workflow puts these in
|
||||
# the job env via `ci/version.sh --env >> $GITHUB_ENV`; recompute here when run
|
||||
# standalone. Passed into the container below and re-exported to the
|
||||
# unprivileged build user in ci/sdk-build-apk.sh.
|
||||
# The workflow puts these in the job env via `ci/version.sh --env >>
|
||||
# $GITHUB_ENV`; recompute here when run standalone. Passed into the container
|
||||
# below and re-exported to the unprivileged build user in ci/sdk-build-apk.sh.
|
||||
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
|
||||
eval "$(sh "$REPO/ci/version.sh" --env)"
|
||||
fi
|
||||
@@ -73,7 +72,7 @@ echo "[apk-feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
|
||||
# SDK; PKG_HASH still verifies every file, so stale = re-downloaded.
|
||||
# apt/ debian:bookworm .deb archives for the host-deps install.
|
||||
# The nested container runs the build as an unprivileged user -> must be writable
|
||||
# (same reason as the chmod 0777 "$OUT" in ci/build-feed.sh).
|
||||
# (same reason as the chmod 0777 "$OUT" above).
|
||||
CACHE="$REPO/.cache"
|
||||
mkdir -p "$CACHE/sdk" "$CACHE/dl" "$CACHE/apt"
|
||||
chmod -R a+rwX "$CACHE/dl" "$CACHE/apt" 2>/dev/null || true
|
||||
@@ -96,8 +95,8 @@ sh "$REPO/ci/fetch-sdk.sh" "$SDK_URL" "$SDK_TAR"
|
||||
|
||||
# --- 1) SDK build + index + sign inside a debian container -------------------
|
||||
# `--volumes-from $(hostname)` shares THIS job container's workspace volume into
|
||||
# the nested container (see ci/build-feed.sh for why a bare -v does not work on
|
||||
# the act_runner DinD setup).
|
||||
# the nested container: a bare `-v $PWD:...` points at a host path that does not
|
||||
# exist under the act_runner DinD setup.
|
||||
echo "[apk-feed] SDK build arch=$ARCH (ImmortalWrt 25.12 apk-SDK)"
|
||||
docker pull -q debian:bookworm
|
||||
docker run --rm --volumes-from "$(hostname)" \
|
||||
|
||||
@@ -1,106 +0,0 @@
|
||||
#!/bin/sh
|
||||
# ci/build-feed.sh — build the signed opkg feed for ONE arch.
|
||||
#
|
||||
# Usage: ci/build-feed.sh <ARCH> <SDK_DOCKER_TAG> <OUTDIR>
|
||||
# e.g. ci/build-feed.sh x86_64 x86_64-24.10.4 out/x86_64
|
||||
# ci/build-feed.sh aarch64_cortex-a53 mediatek-filogic-24.10.4 out/aarch64_cortex-a53
|
||||
#
|
||||
# This is the reusable per-arch entrypoint the Gitea workflow calls. It runs on
|
||||
# the CI RUNNER and:
|
||||
# 1. asserts the prebuilt shaterd binary for this arch was already staged by
|
||||
# scripts/build-shaterd.sh (into openwrt/shaterd/files/) — proving artifact
|
||||
# order: SPA+shaterd build BEFORE the SDK package build;
|
||||
# 2. drives the arch-matched `openwrt/sdk` docker image to compile all 4
|
||||
# packages (ci/sdk-build.sh) and collect their .ipk into OUTDIR;
|
||||
# 3. builds + usign-signs the opkg `Packages` index over OUTDIR
|
||||
# (ci/install-usign.sh + ci/make-index.sh; signs iff $KEY_BUILD is set).
|
||||
#
|
||||
# Env:
|
||||
# KEY_BUILD usign SECRET key (Gitea repo secret). If set, the feed index is
|
||||
# signed and verifiable by dist/shater-feed.pub (fp 5ac4b177689cb8e0).
|
||||
# If unset, an UNSIGNED feed is produced (make-index warns).
|
||||
set -eu
|
||||
|
||||
ARCH="${1:?arch required (x86_64 | aarch64_cortex-a53)}"
|
||||
SDK_TAG="${2:?sdk docker tag required (e.g. x86_64-24.10.4)}"
|
||||
OUT="${3:?output dir required}"
|
||||
|
||||
REPO="$(cd "$(dirname "$0")/.." && pwd)"
|
||||
mkdir -p "$OUT"; OUT="$(cd "$OUT" && pwd)"
|
||||
# $OUT is created here as ROOT on the runner, but the nested `openwrt/sdk`
|
||||
# container runs as the unprivileged `buildbot` (uid 1000) — so it must be able
|
||||
# to write the collected .ipk into $OUT. World-writable is set HERE (a chmod
|
||||
# from inside the container, as buildbot, cannot fix a root-owned dir).
|
||||
chmod 0777 "$OUT"
|
||||
|
||||
# --- 0) the prebuilt shaterd binary must already be staged for this arch ------
|
||||
case "$ARCH" in
|
||||
x86_64) sfx=amd64 ;;
|
||||
aarch64_cortex-a53) sfx=arm64 ;;
|
||||
*) echo "[feed] ERROR: unsupported ARCH '$ARCH'"; exit 2 ;;
|
||||
esac
|
||||
if [ ! -f "$REPO/openwrt/shaterd/files/shaterd-$sfx.upx" ]; then
|
||||
echo "[feed] ERROR: openwrt/shaterd/files/shaterd-$sfx.upx not staged."
|
||||
echo " Run scripts/build-shaterd.sh BEFORE ci/build-feed.sh." >&2
|
||||
exit 3
|
||||
fi
|
||||
|
||||
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
|
||||
|
||||
# --- 0.4) package version from the git tag ------------------------------------
|
||||
# The workflow normally puts these in the job env (ci/version.sh --env >>
|
||||
# $GITHUB_ENV); recompute here when this script is run standalone so a manual
|
||||
# `ci/build-feed.sh ...` produces the same versions as CI. They are handed to the
|
||||
# SDK container below and read by openwrt/*/Makefile (bug B4 — versions used to
|
||||
# be hand-written literals that nobody bumped, so v0.2.2…v0.2.6 all shipped as
|
||||
# 0.2.0-r3 and no router could ever see an update).
|
||||
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
|
||||
eval "$(sh "$REPO/ci/version.sh" --env)"
|
||||
fi
|
||||
echo "[feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
|
||||
|
||||
# --- 0.5) persistent dl/ (package source tarballs) ----------------------------
|
||||
# Workspace dir restored/saved by actions/cache in the workflow and shared into
|
||||
# the nested SDK container via --volumes-from; becomes CONFIG_DOWNLOAD_FOLDER
|
||||
# there (ci/sdk-build.sh). PKG_HASH still verifies every file, so a stale cache
|
||||
# can never produce a wrong build. Must be writable by the container's
|
||||
# unprivileged buildbot user (same reason as the $OUT chmod above).
|
||||
DL_DIR="$REPO/.cache/dl"
|
||||
mkdir -p "$DL_DIR"
|
||||
chmod -R a+rwX "$DL_DIR" 2>/dev/null || true
|
||||
|
||||
# --- 0.6) persistent feeds/ git checkouts -------------------------------------
|
||||
# Workspace dir restored/saved by actions/cache (key: feeds-opkg-<release>) and
|
||||
# symlinked over the SDK's feeds/ inside the container (ci/sdk-build.sh), so
|
||||
# `scripts/feeds update -a` fetches deltas instead of re-cloning base+packages+
|
||||
# luci from scratch (~7 min/run on this runner's slow github.com link).
|
||||
# Top-level chmod only: the contents are created by the container's uid-1000
|
||||
# build user and restored with the same ownership (tar-as-root preserves it).
|
||||
FEEDS_CACHE="$REPO/.cache/feeds/opkg"
|
||||
mkdir -p "$FEEDS_CACHE"
|
||||
chmod a+rwX "$REPO/.cache" "$REPO/.cache/feeds" "$FEEDS_CACHE" 2>/dev/null || true
|
||||
|
||||
# --- 1) SDK package build (4 packages) in the arch-matched SDK image ----------
|
||||
# We drive the `openwrt/sdk` docker image directly (not openwrt/gh-action-sdk):
|
||||
# on a self-hosted Gitea act_runner the marketplace action fetch can be
|
||||
# unavailable, and we need a CLEAN single-feed layout. `--volumes-from
|
||||
# $(hostname)` shares THIS job container's workspace volume into the nested SDK
|
||||
# container — a bare `-v $PWD:...` points at a host path that does not exist
|
||||
# under the act_runner DinD setup. (Requires the job to run inside a container,
|
||||
# which Gitea Actions does by default.)
|
||||
echo "[feed] SDK build arch=$ARCH image=openwrt/sdk:$SDK_TAG"
|
||||
docker pull "openwrt/sdk:$SDK_TAG"
|
||||
docker run --rm --volumes-from "$(hostname)" \
|
||||
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e DL_DIR="$DL_DIR" \
|
||||
-e FEEDS_CACHE="$FEEDS_CACHE" \
|
||||
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
|
||||
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
|
||||
"openwrt/sdk:$SDK_TAG" \
|
||||
sh "$REPO/ci/sdk-build.sh"
|
||||
|
||||
# --- 2) index + sign the per-arch feed (usign, KEY_BUILD passed through) -------
|
||||
sh "$REPO/ci/install-usign.sh"
|
||||
KEY_BUILD="${KEY_BUILD:-}" bash "$REPO/ci/make-index.sh" "$OUT"
|
||||
|
||||
echo "[feed] done arch=$ARCH -> $OUT"
|
||||
ls -l "$OUT"
|
||||
+7
-10
@@ -2,24 +2,21 @@
|
||||
# ci/gen-apk-key.sh — generate the Shater **apk** feed signing keypair (25.12 lane).
|
||||
#
|
||||
# apk (OpenWrt/ImmortalWrt 25.12+) verifies package indexes with EC keys
|
||||
# (prime256v1 PEM), NOT usign — the existing usign identity
|
||||
# (dist/shater-feed.pub, fp 5ac4b177689cb8e0) keeps signing the opkg/24.10 feed
|
||||
# and is NOT touched by this script. This generates a SEPARATE, second identity:
|
||||
# (prime256v1 PEM). This is the ONLY feed identity shater has since the opkg
|
||||
# lane was removed (D22) — the old usign key is history, not a second lane.
|
||||
#
|
||||
# dist/shater-apk.key EC PRIVATE key. NEVER commit (dist/ is gitignored).
|
||||
# Paste its full PEM contents into the Gitea repo secret
|
||||
# KEY_APK (the apk analog of the usign secret KEY_BUILD).
|
||||
# Then delete the local file (or keep it in a password
|
||||
# manager as the offline backup — losing it means every
|
||||
# deployed router must re-trust a new key).
|
||||
# dist/shater-apk.pem PUBLIC key. Commit it next to shater-feed.pub:
|
||||
# KEY_APK. Then delete the local file (or keep it in a
|
||||
# password manager as the offline backup — losing it
|
||||
# means every deployed router must re-trust a new key).
|
||||
# dist/shater-apk.pem PUBLIC key. Commit it:
|
||||
# git add -f dist/shater-apk.pem
|
||||
# (-f because /dist/ is gitignored). Routers install it
|
||||
# as /etc/apk/keys/shater-apk.pem.
|
||||
#
|
||||
# Run ONCE. Refuses to overwrite: regenerating the key invalidates the trust of
|
||||
# every router that already installed shater-apk.pem (same rule as D7 for the
|
||||
# usign key).
|
||||
# every router that already installed shater-apk.pem (see D22).
|
||||
set -eu
|
||||
|
||||
REPO="$(cd "$(dirname "$0")/.." && pwd)"
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
#!/bin/bash
|
||||
# Make `usign` available on the CI runner so ci/make-index.sh can sign the opkg
|
||||
# feed index. The OpenWrt SDK ships usign, but the index/signing step runs on the
|
||||
# bare runner (outside the SDK container), so we build the tiny standalone tool
|
||||
# from source (no libubox — it is intentionally dependency-free so it can
|
||||
# bootstrap a build system). No-op if usign is already on PATH.
|
||||
#
|
||||
# Ported unchanged from Shater v0.1 (ci/install-usign.sh): usign is
|
||||
# format-agnostic and the signing story is identical for the v0.2 4-package feed.
|
||||
#
|
||||
# CI cache: a previously-built binary is reused from $USIGN_CACHE (default:
|
||||
# <repo>/.cache/tools — a workspace dir the workflow persists via actions/cache),
|
||||
# skipping the apt + cmake + clone + build (~1 min). After a fresh build the
|
||||
# binary is copied there so the NEXT run hits the cache. usign is a tiny static
|
||||
# helper with no versioned protocol — a stale cached binary cannot mis-sign.
|
||||
set -eu
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
|
||||
TOOLS="${USIGN_CACHE:-$REPO_ROOT/.cache/tools}"
|
||||
|
||||
# place <binary> — install onto PATH (system-wide if we can, else ~/bin)
|
||||
place() {
|
||||
local SUDO=""; [ "$(id -u)" = 0 ] || SUDO="sudo"
|
||||
if $SUDO install -m0755 "$1" /usr/local/bin/usign 2>/dev/null; then
|
||||
:
|
||||
else
|
||||
mkdir -p "$HOME/bin"
|
||||
install -m0755 "$1" "$HOME/bin/usign"
|
||||
echo "$HOME/bin" >> "${GITHUB_PATH:-/dev/null}"
|
||||
export PATH="$HOME/bin:$PATH"
|
||||
fi
|
||||
}
|
||||
|
||||
if command -v usign >/dev/null 2>&1; then
|
||||
echo "[usign] already present: $(command -v usign)"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ -x "$TOOLS/usign" ]; then
|
||||
place "$TOOLS/usign"
|
||||
echo "[usign] restored from cache: $(command -v usign || echo "$HOME/bin/usign")"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
SUDO=""; [ "$(id -u)" = 0 ] || SUDO="sudo"
|
||||
if ! command -v cmake >/dev/null 2>&1 || ! command -v cc >/dev/null 2>&1; then
|
||||
$SUDO apt-get update -qq
|
||||
$SUDO apt-get install -y -qq cmake gcc git
|
||||
fi
|
||||
|
||||
tmp="$(mktemp -d)"
|
||||
# Canonical source; fall back to the GitHub mirror if git.openwrt.org is flaky.
|
||||
git clone --depth 1 https://git.openwrt.org/project/usign.git "$tmp/usign" \
|
||||
|| git clone --depth 1 https://github.com/openwrt/usign.git "$tmp/usign"
|
||||
( cd "$tmp/usign" && cmake -DCMAKE_BUILD_TYPE=Release . >/dev/null && make >/dev/null )
|
||||
|
||||
place "$tmp/usign/usign"
|
||||
# seed the cache for the next run (best-effort)
|
||||
mkdir -p "$TOOLS" 2>/dev/null && install -m0755 "$tmp/usign/usign" "$TOOLS/usign" 2>/dev/null || true
|
||||
echo "[usign] built: $(command -v usign || echo "$HOME/bin/usign")"
|
||||
@@ -1,39 +0,0 @@
|
||||
#!/bin/bash
|
||||
# Build the opkg feed index (Packages + Packages.gz) with SHA256 for a dir of
|
||||
# .ipk files, then optionally usign-sign it if $KEY_BUILD (the Gitea repo secret)
|
||||
# is set and usign is present. Arg $1 = feed dir.
|
||||
#
|
||||
# Ported from Shater v0.1 (ci/make-index.sh), unchanged. It is package-count and
|
||||
# package-name agnostic: it indexes whatever .ipk are in the dir, so it serves
|
||||
# BOTH the per-arch feed built by ci/build-feed.sh AND the combined release feed
|
||||
# assembled in the release job (shaterd + byedpi per-arch, shater-core +
|
||||
# luci-app-shater = _all). opkg filters by Architecture at install time, so one
|
||||
# combined URL serves every device.
|
||||
#
|
||||
# Feed format: opkg `src/gz` (.ipk + text Packages index, usign signature).
|
||||
# OpenWrt 24.10 (our SDK) still uses opkg; apk arrives at 25.12. The committed
|
||||
# trust anchor dist/shater-feed.pub is a usign (Ed25519) key, matching this.
|
||||
set -euo pipefail
|
||||
OUT="${1:?feed dir required}"; cd "$OUT"
|
||||
: > Packages
|
||||
for ipk in *.ipk; do
|
||||
[ -e "$ipk" ] || continue
|
||||
ctrl=$(tar -xzOf "$ipk" ./control.tar.gz | tar -xzO ./control)
|
||||
sz=$(wc -c < "$ipk"); sha=$(sha256sum "$ipk" | cut -d' ' -f1)
|
||||
printf '%s\n' "$ctrl" | sed '/^[[:space:]]*$/d' >> Packages
|
||||
printf 'Filename: %s\nSize: %s\nSHA256sum: %s\n\n' "$ipk" "$sz" "$sha" >> Packages
|
||||
done
|
||||
gzip -kf Packages
|
||||
|
||||
if [ -n "${KEY_BUILD:-}" ]; then
|
||||
# Signing was requested — a missing/broken signer must FAIL the build, not
|
||||
# silently ship an unsigned feed that routers with check_signature on reject.
|
||||
command -v usign >/dev/null 2>&1 || { echo "[index] ERROR: KEY_BUILD set but usign not found" >&2; exit 1; }
|
||||
umask 077; printf '%s\n' "$KEY_BUILD" > /tmp/usign.sec
|
||||
usign -S -m Packages -s /tmp/usign.sec || { rm -f /tmp/usign.sec; echo "[index] ERROR: usign signing failed" >&2; exit 1; }
|
||||
rm -f /tmp/usign.sec
|
||||
echo "[index] signed -> Packages.sig ($(head -1 Packages.sig))"
|
||||
else
|
||||
echo "[index] no KEY_BUILD -> UNSIGNED feed (opkg needs check_signature off, or set the secret)"
|
||||
fi
|
||||
echo "[index] contents:"; ls -l
|
||||
+3
-4
@@ -10,7 +10,6 @@
|
||||
# the target fleet (BananaWRT 25.12-mtk-vendor = ImmortalWrt 25.12 base, its
|
||||
# distfeeds even point at downloads.immortalwrt.org/releases/25.12-SNAPSHOT) is
|
||||
# ImmortalWrt — so we extract the official ImmortalWrt SDK tarball ourselves.
|
||||
# Same --volumes-from workspace-sharing pattern as ci/sdk-build.sh (opkg lane).
|
||||
#
|
||||
# The OpenWrt buildsystem refuses to run as root, so the SDK build itself runs
|
||||
# as an unprivileged `build` user created here.
|
||||
@@ -36,8 +35,8 @@ echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallba
|
||||
test -f "$REPO/openwrt/shaterd/Makefile" || {
|
||||
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
|
||||
|
||||
# The prebuilt shaterd artifact must already be staged for this arch (same
|
||||
# contract as the opkg lane — scripts/build-shaterd.sh runs first).
|
||||
# The prebuilt shaterd artifact must already be staged for this arch
|
||||
# (artifact-order contract — scripts/build-shaterd.sh runs first).
|
||||
case "$ARCH" in
|
||||
x86_64) sfx=amd64 ;;
|
||||
aarch64_cortex-a53) sfx=arm64 ;;
|
||||
@@ -113,7 +112,7 @@ export HOME=/home/build
|
||||
cd "$SDKDIR"
|
||||
|
||||
# Register this repo's openwrt/ as a src-link feed named `shater` (absolute
|
||||
# path required) — identical to the opkg lane (ci/sdk-build.sh).
|
||||
# path required).
|
||||
cp -f feeds.conf.default feeds.conf
|
||||
grep -q '^src-link shater ' feeds.conf || echo "src-link shater $REPO/openwrt" >> feeds.conf
|
||||
|
||||
|
||||
-143
@@ -1,143 +0,0 @@
|
||||
#!/bin/sh
|
||||
# Runs INSIDE an `openwrt/sdk:<target>-<ver>` container (CWD = SDK root
|
||||
# /builder). The job's workspace is shared into this container via
|
||||
# `docker run --volumes-from`, so the repo is visible at $REPO and output goes
|
||||
# to $OUT (a dir under the repo, hence also visible to the runner afterwards).
|
||||
#
|
||||
# Unlike Shater v0.1 (which compiled ONLY xrayctl in the SDK and hand-packed the
|
||||
# pure-data packages with tar), v0.2 builds ALL FOUR packages the canonical way,
|
||||
# via the SDK feed + `make package/<p>/compile`:
|
||||
#
|
||||
# shaterd prebuilt binary — Build/Compile only VALIDATES that
|
||||
# openwrt/shaterd/files/shaterd-<amd64|arm64>.upx was staged
|
||||
# by scripts/build-shaterd.sh on the runner BEFORE this ran.
|
||||
# (arch-specific .ipk: RSTRIP/STRIP disabled — packed ELF.)
|
||||
# shater-core PKGARCH=all data glue (procd init, sysctl, uci-defaults).
|
||||
# luci-app-shater PKGARCH=all LuCI thin launcher — its Makefile does
|
||||
# `include $(TOPDIR)/feeds/luci/luci.mk`, so the `luci` feed
|
||||
# MUST be updated first (that is what creates feeds/luci/luci.mk).
|
||||
# byedpi arch-specific C — the SDK cross-compiles ciadpi from the
|
||||
# upstream tarball (needs network for PKG_SOURCE_URL).
|
||||
#
|
||||
# Env (required): ARCH, REPO, OUT.
|
||||
set -eu
|
||||
ARCH="${ARCH:?ARCH env required}"
|
||||
REPO="${REPO:?REPO env required}"
|
||||
OUT="${OUT:?OUT env required}"
|
||||
mkdir -p "$OUT"
|
||||
|
||||
echo "[sdk] arch=$ARCH repo=$REPO out=$OUT"
|
||||
# Package version, derived from the git tag by ci/version.sh and handed in by
|
||||
# ci/build-feed.sh. openwrt/{shaterd,shater-core,luci-app-shater}/Makefile read
|
||||
# these straight out of the environment ($(if $(SHATER_PKG_VERSION),...)); make
|
||||
# imports every environment variable as a variable, and it propagates through
|
||||
# `make package/<p>/compile`, the metadata dump and the sub-makes alike.
|
||||
# byedpi deliberately keeps its own upstream version (see its Makefile).
|
||||
echo "[sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
|
||||
test -f "$REPO/openwrt/shaterd/Makefile" || {
|
||||
echo "[sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
|
||||
|
||||
# The prebuilt shaterd artifact must already be staged for this arch.
|
||||
case "$ARCH" in
|
||||
x86_64) sfx=amd64 ;;
|
||||
aarch64_cortex-a53) sfx=arm64 ;;
|
||||
*) echo "[sdk] ERROR: unsupported ARCH '$ARCH'"; exit 2 ;;
|
||||
esac
|
||||
test -f "$REPO/openwrt/shaterd/files/shaterd-$sfx.upx" || {
|
||||
echo "[sdk] ERROR: openwrt/shaterd/files/shaterd-$sfx.upx not staged."
|
||||
echo " scripts/build-shaterd.sh must run on the runner before the SDK build."; exit 3; }
|
||||
|
||||
# --- register this repo's openwrt/ as a src-link feed named `shater` ---------
|
||||
# src-link REQUIRES an absolute path; $REPO/openwrt is exactly a feed root (it
|
||||
# contains the 4 package dirs and nothing else that looks like a package).
|
||||
cp -f feeds.conf.default feeds.conf
|
||||
grep -q '^src-link shater ' feeds.conf || echo "src-link shater $REPO/openwrt" >> feeds.conf
|
||||
|
||||
# Update metadata for ALL feeds: our `shater` feed + the SDK defaults (base,
|
||||
# luci, packages, routing, telephony). We need `luci` for feeds/luci/luci.mk and
|
||||
# `base`/`packages` for the runtime deps (kmod-nft-tproxy, kmod-nft-socket,
|
||||
# ip-full, rpcd, luci-base) to resolve.
|
||||
#
|
||||
# Persistent feeds checkouts: $FEEDS_CACHE (a workspace dir the runner restores
|
||||
# via actions/cache, shared into this container via --volumes-from) replaces
|
||||
# the SDK's ephemeral feeds/ dir, so `feeds update` git-fetches deltas instead
|
||||
# of re-cloning base+packages+luci every run (~7 min on the runner's slow
|
||||
# github.com link). Correctness-safe: update always checks out feeds.conf's
|
||||
# pinned revisions; if it ever fails on a cached checkout (e.g. a force-pushed
|
||||
# upstream), the cache is wiped and the update retried with fresh clones.
|
||||
if [ -n "${FEEDS_CACHE:-}" ] && mkdir -p "$FEEDS_CACHE" 2>/dev/null; then
|
||||
rm -rf feeds
|
||||
ln -s "$FEEDS_CACHE" feeds
|
||||
echo "[sdk] feeds/ -> $FEEDS_CACHE (persistent cache)"
|
||||
fi
|
||||
echo "[sdk] feeds update -a"
|
||||
if ! ./scripts/feeds update -a; then
|
||||
[ -L feeds ] || { echo "[sdk] ERROR: feeds update failed"; exit 8; }
|
||||
echo "[sdk] WARNING: feeds update failed on cached checkouts — wiping cache, cloning fresh"
|
||||
find "$FEEDS_CACHE" -mindepth 1 -maxdepth 1 -exec rm -rf {} + 2>/dev/null || true
|
||||
./scripts/feeds update -a
|
||||
fi
|
||||
|
||||
echo "[sdk] feeds install (prefer shater feed)"
|
||||
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
|
||||
|
||||
# Select our packages, then defconfig. `make package/<p>/compile` builds the
|
||||
# explicit target regardless, but selecting first makes deps visible to defconfig.
|
||||
for p in shaterd shater-core byedpi luci-app-shater; do
|
||||
echo "CONFIG_PACKAGE_$p=m" >> .config
|
||||
done
|
||||
# Route source downloads through OpenWrt's fast CDN mirror FIRST — sourceware.org
|
||||
# (elfutils) and other upstreams intermittently stall mid-transfer, and curl's
|
||||
# --connect-timeout doesn't cover a stalled stream, so the SDK download hangs the
|
||||
# build. LOCALMIRROR is tried before each package's own PKG_SOURCE_URL. (lx CI)
|
||||
echo 'CONFIG_LOCALMIRROR="https://sources.cdn.openwrt.org"' >> .config
|
||||
# Persistent dl/ across runs: $DL_DIR is a workspace dir the runner restores via
|
||||
# actions/cache (see ci/build-feed.sh). Correctness-safe: the buildroot verifies
|
||||
# PKG_HASH on every file already in dl/ and re-downloads on mismatch, so a stale
|
||||
# cache can never leak a wrong source into the build.
|
||||
if [ -n "${DL_DIR:-}" ]; then
|
||||
echo "CONFIG_DOWNLOAD_FOLDER=\"$DL_DIR\"" >> .config
|
||||
fi
|
||||
echo "[sdk] defconfig"
|
||||
make defconfig >/dev/null
|
||||
|
||||
# --- compile the 4 packages --------------------------------------------------
|
||||
for p in shaterd shater-core byedpi luci-app-shater; do
|
||||
echo "[sdk] === build $p ==="
|
||||
make "package/$p/compile" V=s -j"$(nproc)"
|
||||
done
|
||||
|
||||
# --- collect ONLY our 4 packages' .ipk (per-arch shaterd/byedpi + _all core/luci)
|
||||
# NOT `find bin -name '*.ipk'`: the openwrt/sdk image ships HUNDREDS of prebuilt
|
||||
# kmod/base .ipk under bin/, which a blanket copy would pull into the feed and
|
||||
# get signed under OUR key. Match each package's own `<name>_<ver>_<arch>.ipk`.
|
||||
found=0
|
||||
for p in shaterd shater-core byedpi luci-app-shater; do
|
||||
for ipk in $(find bin -type f -name "${p}_*.ipk"); do
|
||||
cp -f "$ipk" "$OUT/"; found=$((found+1))
|
||||
done
|
||||
done
|
||||
[ "$found" -ge 4 ] || { echo "[sdk] ERROR: expected >=4 of OUR .ipk, collected $found"; echo "[sdk] (all .ipk under bin/:)"; find bin -type f -name '*.ipk' | head -20; exit 4; }
|
||||
|
||||
# --- assert the tag-derived version actually reached the packages -------------
|
||||
# The whole point of B4 is that a WRONG-but-plausible version ships silently. The
|
||||
# env -> make hand-off has several layers (docker -e, make's env import, the
|
||||
# metadata dump), so verify the result instead of trusting it: every one of our
|
||||
# three tag-versioned packages must be named `<name>_<ver>-r<rel>_<arch>.ipk`.
|
||||
# byedpi is excluded on purpose — it keeps upstream ByeDPI's own version.
|
||||
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
|
||||
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
|
||||
for p in shaterd shater-core luci-app-shater; do
|
||||
ls "$OUT/${p}_${want}_"*.ipk >/dev/null 2>&1 || {
|
||||
echo "[sdk] ERROR: $p was not built as version '$want'."
|
||||
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
|
||||
echo " Makefile — the build would have shipped a stale version (bug B4)."
|
||||
echo "[sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
|
||||
exit 12; }
|
||||
done
|
||||
echo "[sdk] version check OK — our 3 packages are $want"
|
||||
fi
|
||||
|
||||
chmod -R a+rwX "$OUT" 2>/dev/null || true
|
||||
echo "[sdk] OK arch=$ARCH — collected $found of our .ipk:"
|
||||
ls -l "$OUT"
|
||||
+7
-9
@@ -6,9 +6,9 @@
|
||||
# PKG_VERSION/PKG_RELEASE used to be hand-written literals in the four package
|
||||
# Makefiles, and nobody remembered to bump them: v0.2.2 … v0.2.6 all shipped as
|
||||
# `shaterd 0.2.0-r3` with DIFFERENT binaries inside (v0.2.6's ELF is 5 491 616 B
|
||||
# vs r2's 5 488 336 B). Since both opkg and apk offer an upgrade only when the
|
||||
# feed's version string differs from the installed one, `apk update` saw nothing
|
||||
# new and the routers could not be updated through the normal path at all.
|
||||
# vs r2's 5 488 336 B). Since apk offers an upgrade only when the feed's version
|
||||
# string differs from the installed one, `apk update` saw nothing new and the
|
||||
# routers could not be updated through the normal path at all.
|
||||
#
|
||||
# So the version is now DERIVED, in CI, from the git tag, and the package
|
||||
# Makefiles only carry a fallback for manual/offline builds.
|
||||
@@ -21,14 +21,12 @@
|
||||
# rolling `latest`)
|
||||
# no tag / no git at all -> PKG_VERSION=0.0.0 PKG_RELEASE=1 (+ warning)
|
||||
#
|
||||
# Both managers compare `<upstream>-r<rel>` the same way: the dotted upstream
|
||||
# part first (numerically, component by component), the `r<rel>` only as a
|
||||
# tie-break. Verified against the real tools, not from memory:
|
||||
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
|
||||
# apk compares `<upstream>-r<rel>` as: the dotted upstream part first
|
||||
# (numerically, component by component), the `r<rel>` only as a tie-break.
|
||||
# Verified against the real tool, not from memory —
|
||||
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
|
||||
# 0.2.6-r1 > 0.2.0-r3 0.2.6-r12 > 0.2.6-r1
|
||||
# 0.2.7-r1 > 0.2.6-r12 0.0.0-r1 < 0.2.0-r3
|
||||
# opkg 38eccbb1 from openwrt/rootfs:x86-64-24.10.4 (`opkg compare-versions`):
|
||||
# identical results (opkg implements the Debian algorithm).
|
||||
# That is exactly the ordering this scheme needs:
|
||||
# * a release always outranks every rolling build that preceded it
|
||||
# (0.2.7-r1 > 0.2.6-rN for any N — the dotted part decides), and
|
||||
|
||||
Vendored
-2
@@ -1,2 +0,0 @@
|
||||
untrusted comment: shater feed signing key
|
||||
RWRaxLF3aJy44JbcxSFujtrFFEQ8lIsnTkd1K5TdjIhdlC2c0wa0fv4V
|
||||
+10
-11
@@ -31,12 +31,11 @@ Do not delete it — we port proven pieces from it. What v0.1 has:
|
||||
- **`luci-app-shater`** — a custom "instrument panel" LuCI app (client-side JS +
|
||||
ucode/rpcd ubus backend): Overview with a live Signal Path, Simple/Advanced
|
||||
toggle, quick-start wizard, Nodes/Subs/Rules/DNS/Live/Profiles/Settings pages.
|
||||
- **CI + signed opkg feed** on Gitea: builds per-arch, signs the feed index with
|
||||
usign, publishes a rolling `latest` Gitea release consumable as `src/gz`. **Feed
|
||||
signing key fingerprint `5ac4b177689cb8e0`**; public key `dist/shater-feed.pub`,
|
||||
secret in the Gitea repo secret `KEY_BUILD`.
|
||||
- **CI + a signed package feed** on Gitea: builds per-arch, signs the feed index,
|
||||
publishes a rolling `latest` Gitea release the router consumes as a feed.
|
||||
(v0.1 shipped `.ipk` signed with a usign key — that lane is retired, D22.)
|
||||
- Verified end-to-end on the VM: real LAN client proxied, DNS anti-leak, honest
|
||||
fail-closed, opkg install/upgrade from the signed feed.
|
||||
fail-closed, install/upgrade from the signed feed.
|
||||
|
||||
v0.1 is engine-locked to **xray-core**; its generator, share-link parser and
|
||||
`run.json` are xray-shaped.
|
||||
@@ -91,7 +90,7 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
|
||||
- **`shater` branch `v0.1`** = the standalone xray-based version (frozen, ported
|
||||
from).
|
||||
- Until Phase 1 merges the engine in, `main` is the docs-first overlay seed you
|
||||
are reading now (LICENSE, README, `docs-shater/`, `dist/shater-feed.pub`).
|
||||
are reading now (LICENSE, README, `docs-shater/`, the feed signing key).
|
||||
|
||||
## What to port from v0.1 (don't rewrite these ideas)
|
||||
|
||||
@@ -105,8 +104,8 @@ overlay, don't redo:
|
||||
- **Subscription fetch** (HAPP emulation, fingerprint reconcile, per-sub cache)
|
||||
and the flexible **ruleset/list** model — though sing-box has its own share-link
|
||||
parser and config schema we now target.
|
||||
- **CI feed build + usign signing + Gitea release** (adapt to the single forked
|
||||
binary; keep key `5ac4b177689cb8e0`).
|
||||
- **CI feed build + index signing + Gitea release** (adapted to the single forked
|
||||
binary; the format is apk, signed with the EC key — D22).
|
||||
- The LuCI **design system** (the "instrument panel" identity) — reused for the
|
||||
mini-dashboard and as the panel's visual language.
|
||||
|
||||
@@ -122,9 +121,9 @@ filter/stats engine wired into sing-box's DNS.
|
||||
`https://github.com/SagerNet/sing-box`).
|
||||
- **CI:** Gitea Actions (act_runner + Docker). v0.1's workflow was removed from
|
||||
`main`; new CI is added when the v0.2 build exists.
|
||||
- **Feed signing:** usign key `5ac4b177689cb8e0`; secret in repo secret
|
||||
`KEY_BUILD`; public key `dist/shater-feed.pub` (kept so existing installs keep
|
||||
verifying).
|
||||
- **Feed signing:** EC (prime256v1) key for the apk index; secret in the repo
|
||||
secret `KEY_APK`; public key `dist/shater-apk.pem`, installed on routers as
|
||||
`/etc/apk/keys/shater-apk.pem`. Never regenerate it (D22).
|
||||
- **Test VM:** OpenWrt 24.10.3 x86_64 in Docker (`docker ps --filter
|
||||
name=openwrt-vm`). SSH via the ssh-manager MCP server `local_openwrt`
|
||||
(localhost:2222, root/openwrt). LuCI at `http://127.0.0.1:8080` (root/openwrt),
|
||||
|
||||
+138
-1
@@ -64,11 +64,17 @@ sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files m
|
||||
stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is
|
||||
GPL-3.0 for clarity.
|
||||
|
||||
## D7 — Keep the v0.1 feed signing identity
|
||||
## D7 — Keep the v0.1 feed signing identity *(SUPERSEDED by D22)*
|
||||
The usign feed key `5ac4b177689cb8e0` (public key in `dist/shater-feed.pub`,
|
||||
secret in Gitea secret `KEY_BUILD`) carries over, so routers that already trust it
|
||||
keep verifying v0.2 packages. Do not regenerate it without a documented rotation.
|
||||
|
||||
> **Superseded 2026-07-25 (D22).** The opkg feed this identity signed no longer
|
||||
> exists, so there is nothing left for the key to verify. It was never rotated or
|
||||
> compromised — it is simply unused. `dist/shater-feed.pub` was deleted from the
|
||||
> tree; the reasoning, and how to resurrect the identity if it is ever needed
|
||||
> again, is in D22.
|
||||
|
||||
## D8 — Preserve, don't destroy: v0.1 lives on its branch
|
||||
The reset moved the full working xray-based project to the `v0.1` branch and
|
||||
cleaned `main`. Nothing is lost; reusable logic (reliability layer, nft/routing,
|
||||
@@ -95,6 +101,10 @@ runtime, forcing an ELF with `PT_INTERP=/lib64/ld-linux-x86-64.so.2` + `PT_DYNAM
|
||||
plane is tproxy/redirect (netplane); generate never emits a tun inbound, so
|
||||
the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever
|
||||
appears it falls back to the system stack — re-add the tag then.
|
||||
**REVERTED 2026-07-25 — that reasoning was wrong and shipped a dead feature.**
|
||||
gVisor is not only the tun stack: it is the netstack of the **WireGuard
|
||||
endpoint**, which we do emit and do declare [MVP]. See D23; the tag is back and
|
||||
is now held there by a test.
|
||||
- 2026-07-23: `with_clash_api` also dropped. The admin panel is shater's own
|
||||
web server and generate never emits a `clash_api` service; the desktop/CLI
|
||||
`LX_TAGS` keeps the tag for external dashboards.
|
||||
@@ -529,3 +539,130 @@ the rule editor**, because a second place to author a list is a second place for
|
||||
its semantics and its duplicate-name rules to drift, and the whole point of this
|
||||
decision was to stop having two.
|
||||
|
||||
## D22 — One packaging lane: apk. The opkg/`.ipk` lane is deleted, not disabled
|
||||
Decided 2026-07-25 (product owner). CI built and published TWO signed feeds from
|
||||
every run: opkg/usign (`.ipk` + `Packages.gz`, OpenWrt 24.10) and apk/EC (`.apk` +
|
||||
`packages.adb`, OpenWrt/ImmortalWrt 25.12). The opkg half served nobody. Checked
|
||||
on the actual hardware, not inferred:
|
||||
|
||||
| Device | Firmware | pkg arch | package manager |
|
||||
|---|---|---|---|
|
||||
| `mini_router` (BPi-R3 Mini) | ImmortalWrt 25.12.1 | `aarch64_cortex-a53` | apk-tools 3.0.5 |
|
||||
| `main_router` (BPi-R4) | OpenWrt 25.12.0 | `aarch64_cortex-a53` | apk-tools 3.0.5 — **no `opkg` binary on the system at all** |
|
||||
|
||||
**Decision: delete the opkg lane outright.** Removed: the `build` + `release`
|
||||
jobs from `.gitea/workflows/release.yml`; `ci/build-feed.sh`, `ci/sdk-build.sh`,
|
||||
`ci/make-index.sh`, `ci/install-usign.sh`; and the trust anchor
|
||||
`dist/shater-feed.pub`. The Gitea secret `KEY_BUILD` is now referenced by
|
||||
nothing and can be deleted from the repo settings. `ci/version.sh`,
|
||||
`ci/gitea-release.sh` and `ci/fetch-sdk.sh` are shared or apk-only and stay.
|
||||
|
||||
- **Rejected: keep the lane but stop triggering it** (comment it out / gate it on
|
||||
a dispatch input). Dead code in CI is worse than no code: it keeps a second SDK
|
||||
matrix, a second signing key and a second feed layout alive in everyone's head
|
||||
and in every future edit, and it silently rots because nothing runs it. The
|
||||
24.10 SDK images it pins are themselves a frozen dependency.
|
||||
- **Rejected: keep `dist/shater-feed.pub` as a historical artifact.** A committed
|
||||
trust anchor is an instruction — it invites someone to follow the old install
|
||||
path for a feed that is no longer produced. Nothing is lost by removing it:
|
||||
git history still holds the file, the SECRET half is untouched in `KEY_BUILD`,
|
||||
and a usign secret key blob contains its own public half, so the identity can
|
||||
be reconstructed if a 24.10 device ever has to be served again. Deleting the
|
||||
file is reversible; a stale trust anchor pointing at an unmaintained feed is
|
||||
the thing that quietly misleads.
|
||||
- **Not done: revoking or rotating the usign key.** There is no incident. It is
|
||||
retired, not burned (D7).
|
||||
|
||||
Consequence: one SDK, one key, one feed layout, one set of install instructions.
|
||||
It also makes the rolling release `apk-latest-<arch>` the *only* install path
|
||||
that does not require hand-editing a file per release — which is why the same
|
||||
change fixed it: publishing was an either/or (`apk-latest-<arch>` on dispatch,
|
||||
ELSE `apk-vX.Y.Z-<arch>` on a tag), so once releases moved to tag pushes the
|
||||
rolling pointer stopped being written and froze at `0.2.0` while v0.2.9/v0.2.10
|
||||
shipped — routers on the rolling URL got a successful, silent `apk update` with
|
||||
nothing new. `release-apk` now writes the rolling pointer on every run and
|
||||
asserts, by reading the published release back over the Gitea API, that it holds
|
||||
our three tag-versioned packages at exactly the version just built and no asset
|
||||
at any other version.
|
||||
|
||||
|
||||
## D23 — The router tag set is a checked contract, not a string literal
|
||||
`with_gvisor` was trimmed from the router set on 2026-07-23 (D9) as "unreachable
|
||||
code: we never emit a tun inbound". True about tun — and irrelevant, because
|
||||
gVisor is also the netstack of the **WireGuard endpoint**, which shater emits and
|
||||
FEATURES.md declares [MVP] (AmneziaWG is called *"a driving requirement"*). Every
|
||||
binary shipped between then and 2026-07-25 answered a configured WireGuard node
|
||||
with:
|
||||
|
||||
```
|
||||
create instance: initialize endpoint[0]: create WireGuard device:
|
||||
gVisor is not included in this build, rebuild with -tags with_gvisor
|
||||
```
|
||||
|
||||
`transport/wireguard/device_stack_stub.go` (`//go:build !with_gvisor`) returns
|
||||
`tun.ErrGVisorNotIncluded` from **both** device constructors, so
|
||||
`system_interface: true` is not an escape hatch either: WireGuard was 100% dead
|
||||
in the shipped artifact while the panel offered it, the parser accepted `wg://`,
|
||||
`awg://` and wg-quick `.conf` imports, and the owner had 7 WireGuard sections in
|
||||
UCI on a production router.
|
||||
|
||||
- **Decision:** `with_gvisor` is part of the router tag set and stays there for
|
||||
as long as we ship WireGuard. It costs **~2.8 MB raw / ~0.65 MB UPX per arch**
|
||||
(measured 2026-07-25, both arches; `/overlay` on the production router is
|
||||
6.9 GB with 205 MB used). A tag whose absence turns a declared feature into a
|
||||
runtime error is not "dead weight" — it is the feature.
|
||||
|
||||
### Why the bug was invisible, and what now makes it visible
|
||||
The defect was not a typo in a tag list. It was that **nothing connected the tag
|
||||
list to the feature list**, and the shipped tag combination was the one build
|
||||
configuration nothing exercised: the whole test suite compiles with the FULL
|
||||
upstream set (`with_gvisor` included), so `TestAmneziaWGEndpoint` passed happily
|
||||
while the artifact it was supposed to vouch for could not create a WireGuard
|
||||
device. Tests proved the code was right; they never proved the *build* was.
|
||||
|
||||
Three pieces now hold it together:
|
||||
|
||||
1. **One definition of the set** — `scripts/router-tags.sh` (`SHATER_ROUTER_TAGS`
|
||||
+ `SHATER_ROUTER_LDFLAGS`), sourced by `scripts/build-shaterd.sh` and by the
|
||||
checker. The tag list used to live as a literal inside the build script, i.e.
|
||||
in a file no test reads. A second copy is a second truth.
|
||||
2. **A declared-feature table** — `shater/buildtags`: every tag-gated capability
|
||||
we promise, with the exact tags it needs *to run* and why (the code anchor).
|
||||
`TestRouterTagSetCoversDeclaredFeatures` parses the shell file and fails if a
|
||||
declared feature lost a tag. It needs no build tags, no Linux, no network and
|
||||
no privileges, so it runs in every plain `go test ./...` — including on the
|
||||
Windows dev host, where nothing else can see the shipped configuration.
|
||||
3. **A construction test under the shipped tags** —
|
||||
`shater/generate.TestShippedTagSetConstructsDeclaredProtocols` drives one node
|
||||
of every declared protocol (ss/vmess/trojan/vless ws-grpc-httpupgrade-quic-
|
||||
xhttp/REALITY/uTLS-fp/hysteria2/tuic/**wg**/**awg**) through `box.New`+`Start`.
|
||||
`scripts/check-router-tags.sh` runs it **with `SHATER_ROUTER_TAGS`**, and CI
|
||||
runs that script (`.gitea/workflows/release.yml`) *before* the artifact is
|
||||
built. In a router-tag-set run nothing may be skipped: a protocol that is not
|
||||
compiled in fails the run instead of quietly disappearing from it.
|
||||
|
||||
(2) catches a trim the moment it is made and names the feature it kills; (3)
|
||||
catches what a list comparison cannot — a tag that is present but insufficient.
|
||||
Neither is a substitute for the other. A new protocol in `shater/parse` +
|
||||
`shater/generate` means a new row in `buildtags.Features` and a new probe case;
|
||||
`TestEveryTagGatedFeatureIsProbed` fails until both exist.
|
||||
|
||||
- **Rejected: "just add the tag".** The one-line fix restores WireGuard and
|
||||
leaves the mechanism that hid it fully intact — the next size-driven trim is
|
||||
equally invisible. The tag is the smallest part of this decision.
|
||||
- **Rejected: run the WHOLE test suite with the router tag set in CI.** It is the
|
||||
obvious move and it does not work: parts of the suite legitimately depend on
|
||||
upstream-only tags, and the run costs a second full compile of a 25 MB binary's
|
||||
worth of packages on every release. A focused, unprivileged construction test
|
||||
buys the same evidence for ~10 s and, unlike a full run, can be *required* to
|
||||
skip nothing.
|
||||
- **Rejected: assert the tag set against upstream's `DEFAULT_BUILD_TAGS`.** That
|
||||
makes any trim a failure, which turns the check into noise and re-litigates D9
|
||||
on every upstream rebase. The contract is with our own feature list, not with
|
||||
upstream's.
|
||||
- **Not done: dropping `with_lx_command`.** It is inert for `shaterd` — nothing
|
||||
under `shater/` imports `sing-box/daemon` or `experimental/libbox`, and
|
||||
`go list -deps ./shater/cmd/shaterd` links neither, so it costs zero bytes. It
|
||||
stays only so the router set remains a subset of the lx desktop set. Noted
|
||||
because "a tag that buys nothing" is the mirror image of this bug and should be
|
||||
removed deliberately, not silently.
|
||||
|
||||
@@ -93,8 +93,8 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
|
||||
SIM uplink → different egress); backup/restore; i18n (EN + RU).
|
||||
|
||||
## Ops & distribution
|
||||
- **[MVP]** Single signed binary; signed opkg feed on Gitea (reuse key
|
||||
`5ac4b177689cb8e0`); one-line install; `opkg upgrade`.
|
||||
- **[MVP]** Single signed binary; signed apk feed on Gitea (EC key
|
||||
`dist/shater-apk.pem`); one-line install; named-package `apk upgrade`.
|
||||
- **[T1]** Upstream-rebase cadence (track sing-box-lx tags) with a smoke suite.
|
||||
- **[T2]** apk (OpenWrt 25.x) packaging; multi-router fleet management; REST/gRPC
|
||||
external API; Telegram bot.
|
||||
- **[T2]** Multi-router fleet management; REST/gRPC external API; Telegram bot.
|
||||
(apk packaging landed and is now the only lane — D22.)
|
||||
|
||||
+98
-107
@@ -1,7 +1,7 @@
|
||||
# Shater v0.2 — Build & Install
|
||||
|
||||
How to build the ship artifact (the SPA-embedded `shaterd` binary) and install
|
||||
the OpenWrt feed onto a router.
|
||||
the signed apk repo onto a router.
|
||||
|
||||
## 1. Build the `shaterd` binary
|
||||
|
||||
@@ -28,7 +28,7 @@ Arg / env:
|
||||
- `VERSION` — stamped into `constant.Version`. Resolution: positional arg →
|
||||
`$SHATER_VERSION` → `ci/version.sh --binary` → `v0.2.0-dev`. `ci/version.sh` is
|
||||
the **same** computation the package version comes from (§2.1), so the string
|
||||
the panel shows always matches what `apk info shaterd` / `opkg status` report.
|
||||
the panel shows always matches what `apk list -I shaterd` reports.
|
||||
- `--fast` — skip `npm ci` when `panel/node_modules` already exists.
|
||||
- `UPX=/path/to/upx` — override the UPX binary (default `upx` on `PATH`). UPX is
|
||||
cross-arch, so one host packs both the amd64 and aarch64 ELFs. (Note: UPX also
|
||||
@@ -43,22 +43,44 @@ UPX="…/scratchpad/upx-4.2.4-win64/upx.exe" scripts/build-shaterd.sh v0.2.0 --f
|
||||
The `dist/*` and `openwrt/shaterd/files/shaterd-*.upx` outputs are gitignored —
|
||||
they are release artifacts, not source.
|
||||
|
||||
Tag set (D9 — keep in sync with `docs-shater/DECISIONS.md`):
|
||||
Tag set (D9/D23) — defined in **one** place, `scripts/router-tags.sh`, which
|
||||
documents every tag and is sourced by the build:
|
||||
|
||||
```
|
||||
with_quic,with_wireguard,with_utls,
|
||||
with_gvisor,with_quic,with_wireguard,with_utls,
|
||||
badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command
|
||||
```
|
||||
|
||||
We drop `with_purego,with_naive_outbound`: they pull cronet-go, which forces a
|
||||
glibc `PT_INTERP` even under `CGO_ENABLED=0`, making the binary unusable on musl.
|
||||
We drop `with_gvisor`: the shater data plane is tproxy/redirect and generate
|
||||
never emits a tun inbound, so the userspace gvisor netstack is unreachable code.
|
||||
We drop `with_clash_api`: the admin panel is shater's own web server and the
|
||||
generator never emits a `clash_api` service, so the Clash server is dead code.
|
||||
We drop `with_dhcp`: shater resolver types are `udp/tcp/doh/dot/local/fakeip`;
|
||||
a `dhcp://` DNS transport is never generated or registered.
|
||||
|
||||
`with_gvisor` was dropped in 2026-07 as "unreachable — we emit no tun inbound"
|
||||
and **put back on 2026-07-25**: gVisor is also the netstack of the WireGuard
|
||||
endpoint, so without it every `wg://`/`awg://` node died at apply time with
|
||||
*"gVisor is not included in this build"* while the panel still offered the
|
||||
feature. It costs ~2.8 MB raw / ~0.65 MB UPX per arch. Full story: `DECISIONS.md`
|
||||
D23.
|
||||
|
||||
### Changing the tag set
|
||||
|
||||
Run the guard — it is what stands between a size trim and a silently dead
|
||||
feature, and CI runs it before the artifact is built:
|
||||
|
||||
```sh
|
||||
scripts/check-router-tags.sh # from Windows/macOS it re-execs itself in golang:1.26
|
||||
```
|
||||
|
||||
It (1) fails if a feature declared in `FEATURES.md` lost a build tag it needs to
|
||||
run (`shater/buildtags`, no tags/OS/network required) and (2) constructs one node
|
||||
of every declared protocol through `box.New` **compiled with the shipped tag
|
||||
set** — nothing may be skipped in that run. Adding a protocol to
|
||||
`shater/parse`+`shater/generate` means adding a row to `buildtags.Features` and a
|
||||
probe case in `shater/generate/shipped_tags_linux_test.go`.
|
||||
|
||||
## 2. Packages
|
||||
|
||||
Four OpenWrt packages live under `openwrt/`:
|
||||
@@ -91,9 +113,9 @@ See `openwrt-package-build-ci` for SDK/feed mechanics.
|
||||
`PKG_VERSION`/`PKG_RELEASE` are **not** maintained by hand. They used to be, and
|
||||
nobody bumped them: **v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3`** with
|
||||
different binaries inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B).
|
||||
Both package managers offer an upgrade only when the feed's version string
|
||||
differs from the installed one, so `apk update` saw nothing new and the routers
|
||||
could not be updated through the normal path at all.
|
||||
apk offers an upgrade only when the feed's version string differs from the
|
||||
installed one, so `apk update` saw nothing new and the routers could not be
|
||||
updated through the normal path at all.
|
||||
|
||||
`ci/version.sh` now derives them from `git describe`, once per CI job:
|
||||
|
||||
@@ -103,9 +125,8 @@ could not be updated through the normal path at all.
|
||||
| dispatch, 3 commits past `v0.2.7` | `0.2.7` | `4` | `v0.2.7-r4-g<sha>` |
|
||||
| no reachable tag / no git | `0.0.0` | `1` | `v0.0.0-r1` |
|
||||
|
||||
Ordering is what makes this safe, and both managers agree on it (checked with
|
||||
`apk version -t` on apk-tools 3.0.3 and `opkg compare-versions` on opkg
|
||||
38eccbb1): the dotted part decides first, `-rN` only breaks ties — so
|
||||
Ordering is what makes this safe (checked with `apk version -t` on apk-tools
|
||||
3.0.3): the dotted part decides first, `-rN` only breaks ties — so
|
||||
`0.2.7-r1 > 0.2.6-r12 > 0.2.6-r1 > 0.2.0-r3`. A release therefore always
|
||||
outranks every rolling build before it, rolling builds between two releases grow
|
||||
monotonically, and an untagged build (`0.0.0`) can never masquerade as an
|
||||
@@ -113,8 +134,9 @@ upgrade.
|
||||
|
||||
The value travels as `SHATER_PKG_VERSION`/`SHATER_PKG_RELEASE` in the SDK build
|
||||
environment; the Makefiles read it with a literal fallback for manual/offline
|
||||
builds. Both lanes then **assert** the produced `.ipk`/`.apk` really carries it,
|
||||
so a lost variable fails the build instead of shipping a stale version.
|
||||
builds. `ci/sdk-build-apk.sh` then **asserts** the produced `.apk` really carries
|
||||
it, so a lost variable fails the build instead of shipping a stale version. The
|
||||
release job asserts the same version again on the published rolling repo (§5.1).
|
||||
|
||||
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
|
||||
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
|
||||
@@ -124,21 +146,27 @@ when our packaging of it changes.
|
||||
|
||||
## 3. Install on a router
|
||||
|
||||
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`):
|
||||
**The normal path is the signed apk repo — §5.** This section is the manual
|
||||
fallback (a router with no route to the Gitea host, or a hand-carried build).
|
||||
|
||||
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`).
|
||||
apk filenames carry no architecture, so make sure you copied the `.apk` built for
|
||||
*this* router's arch (`cat /etc/apk/arch`):
|
||||
|
||||
```sh
|
||||
# <ver> = the release version, e.g. 0.2.7-r1 (§2.1 — it comes from the git tag)
|
||||
opkg install shaterd_<ver>_<arch>.ipk # or: apk add shaterd (25.12+)
|
||||
opkg install shater-core_<ver>_all.ipk
|
||||
opkg install luci-app-shater_<ver>_all.ipk
|
||||
opkg install byedpi_0.17.3-r1_<arch>.ipk # optional: ByeDPI egress
|
||||
# --allow-untrusted: our member .apk are unsigned by design — trust lives in the
|
||||
# signed packages.adb index (§5), which a loose file install does not consult.
|
||||
apk add --allow-untrusted ./shaterd-<ver>.apk
|
||||
apk add --allow-untrusted ./shater-core-<ver>.apk
|
||||
apk add --allow-untrusted ./luci-app-shater-<ver>.apk
|
||||
apk add --allow-untrusted ./byedpi-0.17.3-r1.apk # optional: ByeDPI egress
|
||||
```
|
||||
|
||||
Installing from a signed feed instead:
|
||||
From the repo instead (§5 sets it up once), deps pull the rest in:
|
||||
|
||||
```sh
|
||||
# add the feed (customfeeds.conf / apk repositories), then:
|
||||
opkg update && opkg install shater-core luci-app-shater # shaterd pulled in as a dep
|
||||
apk update && apk add luci-app-shater # -> shater-core -> shaterd
|
||||
```
|
||||
|
||||
## 4. Enable
|
||||
@@ -158,92 +186,55 @@ daemon (`shaterd run`), which owns the engine, the `inet shater` data plane, pol
|
||||
routing, in-process DNS, and the admin panel (default `:8088`). The LuCI app's
|
||||
"Open panel" button mints a single-use token and hands the browser off to the panel.
|
||||
|
||||
## 5. Add the signed feed (recommended — then `opkg upgrade` just works)
|
||||
## 5. The signed apk repo (the normal install path)
|
||||
|
||||
CI (`.gitea/workflows/release.yml`) publishes every build as a **rolling `latest`
|
||||
Gitea release** that is itself a signed opkg `src/gz` feed: the release holds the
|
||||
`.ipk` for all arches, a `Packages`/`Packages.gz` index, a usign `Packages.sig`,
|
||||
and the public key `shater-feed.pub`. opkg filters by `Architecture`, so the **same
|
||||
two lines work on every device** (x86 testbed picks `x86_64 + all`; the BPI routers
|
||||
pick `aarch64_cortex-a53 + all`).
|
||||
OpenWrt/ImmortalWrt **25.12** packages with Alpine's **apk**: `.apk` files, a
|
||||
binary `packages.adb` index, EC (prime256v1) keys in `/etc/apk/keys/`, and
|
||||
effectively mandatory signatures (unsigned needs `--allow-untrusted`). This is
|
||||
the only format shater publishes — the `.ipk`/opkg lane was removed in 2026-07
|
||||
(`DECISIONS.md` D22); every device we serve is on 25.12 with apk-tools 3.
|
||||
|
||||
> **Format:** OpenWrt 24.10 (our SDK) uses **opkg** (`.ipk`, `Packages.gz`, usign),
|
||||
> so the feed is `src/gz` and the trust anchor is the usign key
|
||||
> `dist/shater-feed.pub` (fingerprint **`5ac4b177689cb8e0`**). apk only replaces
|
||||
> opkg at OpenWrt **25.12** — see §6.
|
||||
|
||||
One-time setup on the router:
|
||||
|
||||
```sh
|
||||
# 1) trust the feed key — the FILENAME must equal the usign key fingerprint.
|
||||
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
|
||||
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
|
||||
|
||||
# 2) add the feed (one URL serves every arch).
|
||||
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
|
||||
>> /etc/opkg/customfeeds.conf
|
||||
|
||||
# 3) refresh + install (shaterd is pulled in as a dependency).
|
||||
opkg update
|
||||
opkg install luci-app-shater # -> shater-core -> shaterd
|
||||
opkg install byedpi # optional: ByeDPI desync egress
|
||||
```
|
||||
|
||||
With the key installed, opkg's default `check_signature 1` verifies the feed on
|
||||
every `opkg update`; no `--nocheck-signature` needed. A **tagged** release
|
||||
(`vX.Y.Z`) publishes the identical layout at
|
||||
`.../releases/download/vX.Y.Z` if you prefer to pin a version instead of tracking
|
||||
`latest`.
|
||||
|
||||
### Updating
|
||||
|
||||
Name the packages. **Never run a bare `opkg upgrade`** — with no arguments it
|
||||
tries to upgrade *every* installed package from *every* configured feed, which on
|
||||
OpenWrt means base/system packages on the overlay and is a well-known way to
|
||||
brick a router.
|
||||
|
||||
```sh
|
||||
opkg update
|
||||
opkg upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
|
||||
```
|
||||
|
||||
Drop `byedpi` from the list if you never installed it. An upgrade is offered only
|
||||
when the feed's `Version` differs from the installed one — that is exactly what
|
||||
bug B4 broke (v0.2.2…v0.2.6 all published as `0.2.0-r3`). Since then CI derives
|
||||
the version from the git tag on every build (§2.1), so there is nothing to bump
|
||||
by hand any more; check with:
|
||||
|
||||
```sh
|
||||
opkg list-installed | grep -E 'shaterd|shater-core|luci-app-shater|byedpi'
|
||||
```
|
||||
|
||||
## 6. apk feed (OpenWrt/ImmortalWrt 25.12+ — incl. BananaWRT 25.12-mtk-vendor)
|
||||
|
||||
OpenWrt/ImmortalWrt **25.12** replaces opkg with Alpine's **apk**: `.apk` files,
|
||||
a binary `packages.adb` index, EC (prime256v1) keys in `/etc/apk/keys/`, and
|
||||
effectively mandatory signatures (unsigned needs `--allow-untrusted`). The
|
||||
package **Makefiles are unchanged** — the SDK release decides the format.
|
||||
|
||||
CI builds this lane **in parallel** with the opkg feed (same manual triggers:
|
||||
`v*` tag push or `workflow_dispatch`): the `build-apk` jobs in
|
||||
`.gitea/workflows/release.yml` compile the same 4 packages through the official
|
||||
**ImmortalWrt 25.12 SDK** (tarballs from
|
||||
CI (`v*` tag push or `workflow_dispatch`) compiles the 4 packages through the
|
||||
official **ImmortalWrt 25.12 SDK** (tarballs from
|
||||
`downloads.immortalwrt.org/releases/25.12.1/targets/{x86/64,mediatek/filogic}/`)
|
||||
and publish **one release per arch** — rolling `apk-latest-x86_64` /
|
||||
`apk-latest-aarch64_cortex-a53`, or `apk-vX.Y.Z-<arch>` for a tagged version.
|
||||
Per-arch (unlike the combined opkg release) because apk filenames carry no
|
||||
architecture and packages are fetched relative to the `packages.adb` URL.
|
||||
and publishes **one release per arch**: the rolling `apk-latest-x86_64` /
|
||||
`apk-latest-aarch64_cortex-a53`, plus `apk-vX.Y.Z-<arch>` on a tag. Per-arch
|
||||
because apk filenames carry no architecture and packages are fetched *relative to
|
||||
the `packages.adb` URL*, so one flat multi-arch release would collide.
|
||||
|
||||
> **Key:** apk cannot use the usign key. The apk trust anchor is the separate EC
|
||||
> public key **`dist/shater-apk.pem`** (generated once by `ci/gen-apk-key.sh`;
|
||||
> private half lives ONLY in the Gitea secret **`KEY_APK`**, the apk analog of
|
||||
> `KEY_BUILD`). Never regenerate either key — that invalidates every deployed
|
||||
> router's trust. The usign identity `shater-feed.pub` keeps signing the
|
||||
> opkg/24.10 feed, untouched.
|
||||
> **Key:** the trust anchor is the EC public key **`dist/shater-apk.pem`**
|
||||
> (generated once by `ci/gen-apk-key.sh`; the private half lives ONLY in the
|
||||
> Gitea secret **`KEY_APK`**). Never regenerate it — that invalidates every
|
||||
> deployed router's trust.
|
||||
|
||||
One-time setup on a 25.12 router (BananaWRT `25.12-mtk-vendor` on the BPI-R3
|
||||
mini, BPI-R4 on 25.12, or the future 25.12 VM — `/etc/apk/arch` picks the right
|
||||
per-arch release automatically):
|
||||
### 5.1 Rolling or pinned — pick the repo URL deliberately
|
||||
|
||||
The repo line names an **index file**, and which one you name is the whole
|
||||
update policy:
|
||||
|
||||
| Repo line points at | Behaviour | Cost |
|
||||
|---|---|---|
|
||||
| `apk-latest-<arch>/packages.adb` (**rolling**) | Every release run REPLACES this release's assets, so `apk update && apk upgrade <our packages>` always sees the newest build. Install once, never touch the file again. | You get whatever CI published last; there is no per-router pin. |
|
||||
| `apk-vX.Y.Z-<arch>/packages.adb` (**pinned**) | The router stays on exactly that build. `apk update` will never offer a newer shater. | `/etc/apk/repositories.d/shater.list` must be edited **by hand on every upgrade**, on every router. |
|
||||
|
||||
`mini_router` is deliberately on a **pinned** URL — a considered choice, and the
|
||||
hand-edit per release is its price. Use rolling unless you specifically want to
|
||||
freeze a device.
|
||||
|
||||
> The rolling release used to go stale silently: publishing was an either/or, so
|
||||
> tag runs wrote only `apk-vX.Y.Z-<arch>` and `apk-latest-<arch>` was last
|
||||
> refreshed on 2026-07-24 at `0.2.0` while v0.2.9/v0.2.10 shipped. A router on
|
||||
> the rolling URL kept getting a successful `apk update` with nothing new. Fixed
|
||||
> 2026-07-25: `release-apk` writes the rolling pointer on **every** run and then
|
||||
> reads the release back over the Gitea API, asserting it holds our three
|
||||
> tag-versioned packages at exactly the version just built and **no** leftover
|
||||
> asset at another version (two versions of one package in one index would let
|
||||
> apk choose instead of us).
|
||||
|
||||
### 5.2 One-time setup on the router
|
||||
|
||||
BananaWRT `25.12-mtk-vendor` on the BPI-R3 mini, OpenWrt 25.12 on the BPI-R4, or
|
||||
the testbed VM — `/etc/apk/arch` picks the right per-arch release automatically:
|
||||
|
||||
```sh
|
||||
# 1) trust the apk feed key (any *.pem filename under /etc/apk/keys works).
|
||||
@@ -251,6 +242,7 @@ wget -O /etc/apk/keys/shater-apk.pem \
|
||||
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
|
||||
|
||||
# 2) add the repo — the line points at the packages.adb INDEX FILE itself.
|
||||
# (rolling; for a pinned router put apk-vX.Y.Z-$(cat /etc/apk/arch) here — §5.1)
|
||||
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
|
||||
> /etc/apk/repositories.d/shater.list
|
||||
|
||||
@@ -260,7 +252,7 @@ apk add luci-app-shater # -> shater-core -> shaterd
|
||||
apk add byedpi # optional: ByeDPI desync egress
|
||||
```
|
||||
|
||||
### Updating
|
||||
### 5.3 Updating
|
||||
|
||||
**Never run a bare `apk upgrade`.** With no arguments apk reconciles *every*
|
||||
installed package against *every* configured repository at once; on a router
|
||||
@@ -286,9 +278,8 @@ Drop `byedpi` from either list if you never installed it. Check what you are on
|
||||
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
|
||||
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
|
||||
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
|
||||
as `0.2.0-r3` and `apk update` offered nothing). Pin a version instead of tracking
|
||||
rolling by pointing the repo line at
|
||||
`.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
|
||||
as `0.2.0-r3` and `apk update` offered nothing). Rolling vs pinned repo URL —
|
||||
§5.1.
|
||||
|
||||
### BananaWRT `25.12-mtk-vendor` compatibility
|
||||
|
||||
|
||||
@@ -6,7 +6,7 @@ OpenWrt). Лицо репозитория и быстрый старт — в к
|
||||
| Документ | О чём |
|
||||
|----------|-------|
|
||||
| [CONTEXT.md](CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, решения в кратце, testbed/инфра |
|
||||
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка обоих фидов — opkg (24.10) и apk (25.12+) |
|
||||
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка apk-фида (25.12+): роллинг или фиксация версии |
|
||||
| [ARCHITECTURE.md](ARCHITECTURE.md) | One-binary дизайн, auth-handoff LuCI→панель, data/DNS/apply-потоки (диаграммы) |
|
||||
| [FEATURES.md](FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
|
||||
| [ROADMAP.md](ROADMAP.md) | Фазовый план |
|
||||
|
||||
@@ -104,7 +104,7 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
|
||||
|
||||
## Phase 8 — Ship it ✅ DONE
|
||||
- Adapt CI to build/sign the single forked binary for both arches; publish the
|
||||
signed opkg feed (reuse key `5ac4b177689cb8e0`); install/upgrade docs.
|
||||
signed feed (apk since D22, EC key `dist/shater-apk.pem`); install/upgrade docs.
|
||||
- Set an upstream-rebase cadence (merge new sing-box-lx tags, run the smoke suite).
|
||||
|
||||
## Cross-cutting (every phase)
|
||||
|
||||
@@ -21,8 +21,8 @@ PKG_NAME:=byedpi
|
||||
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
|
||||
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
|
||||
# actually installed. Stamping our tag on it would be both a lie and a
|
||||
# regression: our tags are 0.2.x, and every version comparator (apk-tools 3 and
|
||||
# opkg alike, verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
|
||||
# regression: our tags are 0.2.x, and the version comparator (apk-tools 3,
|
||||
# verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
|
||||
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
|
||||
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
|
||||
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
|
||||
|
||||
@@ -38,9 +38,9 @@ PKG_NAME:=shaterd
|
||||
# VERSIONING — derived from the git tag, NOT hand-maintained here (bug B4).
|
||||
# ci/version.sh turns `git describe` into SHATER_PKG_VERSION/SHATER_PKG_RELEASE
|
||||
# (tag vX.Y.Z -> X.Y.Z + r1; off-tag -> last tag + r<commits+1>), and
|
||||
# ci/build-feed.sh / ci/build-feed-apk.sh export them into the SDK build env of
|
||||
# both lanes. Both lanes then ASSERT that the produced .ipk/.apk really carries
|
||||
# that version, so a lost env can never silently ship a stale one again.
|
||||
# ci/build-feed-apk.sh exports them into the SDK build env. ci/sdk-build-apk.sh
|
||||
# then ASSERTS that the produced .apk really carries that version, so a lost env
|
||||
# can never silently ship a stale one again.
|
||||
# The literals below are ONLY the manual/offline fallback (no CI, no git) — they
|
||||
# are not "the release version"; releases are named by the tag.
|
||||
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
|
||||
@@ -104,8 +104,8 @@ define Package/shaterd/install
|
||||
$(INSTALL_BIN) $(CURDIR)/files/$(SHATERD_BIN) $(1)/usr/bin/shaterd
|
||||
endef
|
||||
|
||||
# This package ships ONLY the binary — no init script — so opkg's default
|
||||
# postinst never touches the running service. On `opkg upgrade shaterd` the new
|
||||
# This package ships ONLY the binary — no init script — so the package manager's
|
||||
# postinst never touches the running service. On `apk upgrade shaterd` the new
|
||||
# ELF lands at /usr/bin/shaterd while the OLD image keeps running from its
|
||||
# unlinked inode: the upgrade silently has no effect until the next reboot, and
|
||||
# meanwhile the new CLI (`shaterd reconcile`, `status`, `mint-token` — invoked by
|
||||
|
||||
@@ -18,6 +18,32 @@ type URLTestOutboundOptions struct {
|
||||
// lx: SPEC 019 v2 — load-balancing.
|
||||
Mode string `json:"mode,omitempty"` // least_test (default) | round_robin
|
||||
Balancer *URLTestBalancerOptions `json:"balancer,omitempty"`
|
||||
// lx: health board §5.C — SelfCheck stands the group's OWN background
|
||||
// health-check up or down. nil/absent == true, so every existing config keeps
|
||||
// today's behaviour.
|
||||
//
|
||||
// Why this exists at all: a urltest group probes its members BY ITSELF — a
|
||||
// warm-up sweep at PostStart and a ticker for as long as traffic keeps
|
||||
// touching it — and it dials the members' outbounds DIRECTLY, from the
|
||||
// router, over whatever the default WAN route is. For a group that traffic
|
||||
// actually flows through, that is exactly right: the probe travels the same
|
||||
// path the connections do. But for a group NO routing rule reaches, that
|
||||
// same probe measures a path nothing uses — and it stores the result under
|
||||
// the members' BASE tags, which every health consumer then reads as "the
|
||||
// node's health". A node that is blocked on the direct WAN and perfectly
|
||||
// alive behind a tunnel therefore reads "dead" the moment such a group
|
||||
// probes it; the reading is not merely stale, it is FALSE, and it poisons
|
||||
// the shared board for everyone (selection, the panel, the observatory's
|
||||
// freshness gate). SelfCheck=false is how the control plane stands such a
|
||||
// group's own schedule down: the shater engine computes which groups the
|
||||
// applied rules actually reach (the observatory's used-set) and disables
|
||||
// the self-check on the rest, so the ONLY prober left is the observatory —
|
||||
// which probes along the real dial paths and nothing else.
|
||||
//
|
||||
// The flag suppresses only the group's own SCHEDULE (the PostStart warm-up
|
||||
// and the Touch ticker). An EXPLICIT CheckOutbounds/URLTest call — the
|
||||
// adapter interface a human or an API invokes on purpose — still works.
|
||||
SelfCheck *bool `json:"self_check,omitempty"`
|
||||
}
|
||||
|
||||
// URLTestBalancerOptions configures round_robin: a fixed-size pool of live nodes, lazily
|
||||
|
||||
+2
-1
@@ -8,7 +8,8 @@
|
||||
"dev": "vite",
|
||||
"build": "tsc --noEmit && vite build",
|
||||
"preview": "vite preview",
|
||||
"typecheck": "tsc --noEmit"
|
||||
"typecheck": "tsc --noEmit",
|
||||
"test": "node --test src/*.test.ts"
|
||||
},
|
||||
"dependencies": {
|
||||
"react": "^18.3.1",
|
||||
|
||||
@@ -405,3 +405,73 @@
|
||||
color: var(--dim);
|
||||
max-width: 74ch;
|
||||
}
|
||||
|
||||
/* ---- inline rename (shared) ----
|
||||
The pencil-in-the-row interaction: click the ✎ beside a name, type over it,
|
||||
Enter commits / Esc cancels / blur commits. Lifted out of Devices.css when
|
||||
Nodes grew the same affordance — one interaction, one set of rules, so the two
|
||||
pages can never drift apart. `--locked` is the same control with the action
|
||||
withheld: it stays visible and focusable-looking so a missing rename reads as
|
||||
a stated rule, not a dead button. */
|
||||
.inline-rename {
|
||||
flex: none;
|
||||
display: inline-flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
width: 22px;
|
||||
height: 22px;
|
||||
padding: 0;
|
||||
border: 1px solid transparent;
|
||||
border-radius: 5px;
|
||||
background: none;
|
||||
color: var(--faint);
|
||||
font-size: 12px;
|
||||
line-height: 1;
|
||||
cursor: pointer;
|
||||
transition: color 0.15s, background 0.15s, border-color 0.15s;
|
||||
}
|
||||
.inline-rename:hover:not(:disabled) {
|
||||
color: var(--accent);
|
||||
background: color-mix(in srgb, var(--accent) 12%, transparent);
|
||||
}
|
||||
.inline-rename:focus-visible {
|
||||
color: var(--accent);
|
||||
border-color: var(--accent);
|
||||
outline: 2px solid var(--accent);
|
||||
outline-offset: 1px;
|
||||
}
|
||||
.inline-rename:disabled {
|
||||
opacity: 0.5;
|
||||
cursor: default;
|
||||
}
|
||||
/* Withheld, not broken: keep the glyph readable and let the cursor say "there is
|
||||
a reason" rather than dimming it into invisibility. */
|
||||
.inline-rename--locked {
|
||||
opacity: 0.75;
|
||||
cursor: help;
|
||||
}
|
||||
.inline-rename--locked:hover {
|
||||
color: var(--dim);
|
||||
background: none;
|
||||
}
|
||||
|
||||
.inline-rename-input {
|
||||
min-width: 0;
|
||||
max-width: 24ch;
|
||||
padding: 4px 8px;
|
||||
border: 1px solid var(--accent);
|
||||
border-radius: 6px;
|
||||
background: var(--sink);
|
||||
color: var(--ink);
|
||||
font-size: 13px;
|
||||
font-weight: 600;
|
||||
letter-spacing: 0.01em;
|
||||
box-shadow: 0 1px 2px var(--shadow) inset;
|
||||
}
|
||||
.inline-rename-input:focus-visible {
|
||||
outline: 2px solid var(--accent);
|
||||
outline-offset: 1px;
|
||||
}
|
||||
.inline-rename-input:disabled {
|
||||
opacity: 0.55;
|
||||
}
|
||||
|
||||
+172
-16
@@ -51,6 +51,38 @@ export class ApiError extends Error {
|
||||
*/
|
||||
export type Plane = 'full' | 'hold' | 'none'
|
||||
|
||||
/**
|
||||
* Where the router's traffic actually ENDS UP, decided by the daemon from the
|
||||
* engine config it is running (apply.Status.traffic ← generate.TrafficOf).
|
||||
*
|
||||
* tunnel — the default route goes into a tunnel: everything not matched by a
|
||||
* more specific rule is proxied.
|
||||
* split — the default leaves directly, but some rules do tunnel their traffic.
|
||||
* direct — the default leaves directly and nothing is tunnelled at all.
|
||||
* blocked — the default is the fail-closed backstop: unmatched traffic is
|
||||
* dropped, not let out. Nothing leaks.
|
||||
*
|
||||
* `plane` DOES NOT ANSWER THIS and must never be read as if it did. `plane` says
|
||||
* how much of the data plane is installed (nft table, policy routing, engine up);
|
||||
* a router whose only rule is `default → direct` has all of it and sends the whole
|
||||
* LAN out the plain WAN with its real address. That combination — plane "full",
|
||||
* traffic "direct" — was live on a user's router under a green "Protected" LED.
|
||||
*/
|
||||
export type TrafficVerdict = 'tunnel' | 'split' | 'direct' | 'blocked'
|
||||
|
||||
export interface Traffic {
|
||||
// '' or absent ⇒ not known (daemon that predates this field, nothing applied
|
||||
// yet, or the plane is on hold). NEVER treat unknown as 'tunnel'.
|
||||
verdict?: TrafficVerdict | ''
|
||||
// The outbound tag the engine's default route names, in the engine's own
|
||||
// vocabulary ("direct", "block", a node/group tag). Diagnostic — wording is
|
||||
// driven by `verdict`, never by parsing this.
|
||||
default?: string
|
||||
// How many of the engine's route rules send their matched traffic into a tunnel.
|
||||
// Separates "some of your traffic is protected" from "none of it is".
|
||||
tunnel_rules?: number
|
||||
}
|
||||
|
||||
/**
|
||||
* One thing the last apply could not do. Deliberately fail-OPEN with a warning
|
||||
* rather than refusing the whole config (the alternative was taking the network
|
||||
@@ -88,6 +120,10 @@ export interface Status {
|
||||
// How much of the data plane is installed. Absent on older daemons ⇒ unknown,
|
||||
// in which case the UI shows nothing rather than guessing "full".
|
||||
plane?: Plane
|
||||
// Where the traffic actually goes under the running config. Absent on older
|
||||
// daemons ⇒ unknown; see TrafficVerdict for why this is a separate question
|
||||
// from `plane`.
|
||||
traffic?: Traffic
|
||||
// Findings from the last apply. ALWAYS an array from the daemon (never null);
|
||||
// empty means the last apply was clean. Pre-sorted critical-first and capped at
|
||||
// 50, where a truncated list ends with an `info` entry saying "suppressed".
|
||||
@@ -354,19 +390,92 @@ export interface GroupHealth {
|
||||
* any more and nothing to report here beyond the groups themselves.
|
||||
*/
|
||||
|
||||
/**
|
||||
* One hop of one chain, measured where that hop actually sits in the path.
|
||||
*
|
||||
* This is the reading the daemon always took and never showed. A chain is not a
|
||||
* target with a single health — it is an ordered series of them, and the only
|
||||
* question an operator ever asks about a broken chain is WHICH hop broke. The
|
||||
* end-to-end exit reading cannot answer that: it says "the path is dead" for a
|
||||
* four-hop chain and leaves the person to guess between four suspects.
|
||||
*
|
||||
* WIRE ORDER. `index` is 1-based and counts hops in the order the router dials
|
||||
* them: hop 1 is the first physical hop, and each later hop is dialled THROUGH
|
||||
* the ones before it. The hop carrying `exit: true` — always the largest index —
|
||||
* is where traffic leaves for the internet. A leading `egress:` in the chain's
|
||||
* configured Hops is NOT a numbered hop: the daemon lifts it into the entry
|
||||
* detour of hop 1, so a chain written `egress:ewan → node:awgout → group:sub0`
|
||||
* reports two hops, not three. Anything zipping this against the model's Hops
|
||||
* must drop that leading egress first and give up on labelling entirely if the
|
||||
* counts still disagree — a chain that splices sub-chains gets flattened here,
|
||||
* and a confidently WRONG hop name is worse than no name.
|
||||
*
|
||||
* `tag` is the engine-side outbound (`chain-<name>-h2`). Debugging and tooltips
|
||||
* only; it is never a label to put in front of a person.
|
||||
*
|
||||
* NODE HOP vs GROUP HOP. For `kind: "node"` the hop IS the measurement: `total`
|
||||
* is 1, the counters follow its own state, and `selected` is ''. For
|
||||
* `kind: "group"` the counters roll up that hop's per-hop member COPIES — the
|
||||
* copies dialled through the hops in front of it, which is exactly why they can
|
||||
* read alive here while the same group's standalone card reads dead. Both
|
||||
* readings are true; they measure different dial paths. `selected` is the node
|
||||
* NAME the hop routes through right now, and `delay_ms` / `age_seconds` belong
|
||||
* to that selected member (or the freshest alive one).
|
||||
*
|
||||
* Invariants the daemon guarantees — never re-derive them, just read them:
|
||||
* `tested === alive + dead` and `alive + dead + untested === total`.
|
||||
*
|
||||
* `state` is a closed set. `untested` is NEVER "dead" and never "healthy": it
|
||||
* means nothing fresh enough is known, which for a used chain is seconds away
|
||||
* from resolving on its own. `age_seconds: -1` means the age is unknown.
|
||||
*/
|
||||
export interface ChainHopHealth {
|
||||
/** 1-based WIRE order. Hop 1 is dialled first; see the note above. */
|
||||
index: number
|
||||
/** Engine outbound tag (`chain-<name>-h2`) — tooltips/debugging, never a label. */
|
||||
tag: string
|
||||
/** `node` ⇒ the hop is the measurement. `group` ⇒ the counters roll up members. */
|
||||
kind: 'node' | 'group'
|
||||
/** This hop is where traffic leaves for the internet. Always the largest index. */
|
||||
exit: boolean
|
||||
/** Closed set — switch on it exhaustively. `untested` is never "dead". */
|
||||
state: 'alive' | 'dead' | 'untested'
|
||||
/** RTT of the selected/freshest alive member; 0 (meaningless) when not alive. */
|
||||
delay_ms: number
|
||||
/** Age of that measurement in seconds; -1 when unknown. */
|
||||
age_seconds: number
|
||||
/** Node name this GROUP hop routes through right now; '' for a node hop. */
|
||||
selected: string
|
||||
total: number
|
||||
tested: number
|
||||
alive: number
|
||||
dead: number
|
||||
untested: number
|
||||
}
|
||||
|
||||
/** Per-chain reachability, the chain analogue of {@link GroupHealth}.used (plan
|
||||
* §5.E): a chain no enabled routing rule routes through is outside the
|
||||
* observatory's plan, so its exit is never probed and the Targets card renders it
|
||||
* "unused" instead of an exit-test readout. A chain has no membership counters —
|
||||
* it is a fixed path, and its end-to-end health is the exit test's job. */
|
||||
* observatory's plan, so nothing probes it and the Targets card says so instead
|
||||
* of rendering a health reading. A chain has no membership counters of its own —
|
||||
* it is a fixed path, and its health lives on its {@link ChainHopHealth} hops. */
|
||||
export interface ChainHealth {
|
||||
name: string
|
||||
/** An enabled routing rule (the Final target, a DNS-resolver detour, a device
|
||||
* target, …) reaches this chain, so the observatory probes its exit in the
|
||||
* target, …) reaches this chain, so the observatory probes its hops in the
|
||||
* background. false ⇒ nothing routes through the chain: it is skipped by the
|
||||
* background probing and its end-to-end health stays untested. That is an
|
||||
* "unused" note about the ROUTING CONFIG, never a health problem. */
|
||||
* background probing and its health stays untested. That is an "unused" note
|
||||
* about the ROUTING CONFIG, never a health problem. */
|
||||
used: boolean
|
||||
/**
|
||||
* Per-hop health in wire order (see {@link ChainHopHealth}).
|
||||
*
|
||||
* MAY BE ABSENT, and absent does not mean "this chain has no hops". It means
|
||||
* the engine never materialised per-hop outbounds for it: the chain is unused,
|
||||
* or it collapses to a single hop and the daemon points traffic straight at
|
||||
* that target instead of building a copy of it. Read a missing key as "nothing
|
||||
* measured per hop", never as an empty path or as a fault.
|
||||
*/
|
||||
hops?: ChainHopHealth[]
|
||||
}
|
||||
|
||||
export interface GroupsHealth {
|
||||
@@ -1226,6 +1335,17 @@ export interface RuleReach {
|
||||
shadowed_by_order?: number
|
||||
/** Operator-facing sentence; absent when `unreachable` is false. */
|
||||
reason?: string
|
||||
/**
|
||||
* Whether the rule is IN FORCE right now — `Rule.Enabled` after the active WAN
|
||||
* profile's overrides. This is NOT `GET /api/config`'s `Enabled`: that one is
|
||||
* the desired state the page PUTs back, and on a router with profiles the two
|
||||
* legitimately disagree. Draw rows from this; keep the switch on the other.
|
||||
*/
|
||||
effective_enabled: boolean
|
||||
/** The active profile that CHANGED this rule's state; absent when none did. */
|
||||
overridden_by?: string
|
||||
/** Which way it went. Absent together with `overridden_by`. */
|
||||
override?: 'enabled' | 'disabled'
|
||||
}
|
||||
|
||||
/** GET /api/rules/reachability. `rules` is ALWAYS an array, one entry per rule in
|
||||
@@ -1385,8 +1505,21 @@ export function getGroupsHealth(
|
||||
}
|
||||
|
||||
/**
|
||||
* One group's (or chain's) last test: which member the balancer picked, how fast
|
||||
* it answered, and what the internet saw as the source address.
|
||||
* What the OBSERVATORY measured for one group or chain — not a dial the panel
|
||||
* made.
|
||||
*
|
||||
* This shape used to come from a fresh connection opened on demand, straight at
|
||||
* the target. That was a lie on any router whose proxies are blocked when dialled
|
||||
* directly and work only as a hop behind a tunnel: the card reported dead for a
|
||||
* path that carries traffic all day. The daemon now has exactly one thing that
|
||||
* measures — the background observatory, which probes along the REAL dial path,
|
||||
* per-hop copies and all — and this endpoint reports what it found. There is no
|
||||
* second measurement anywhere, and the panel never opens a connection of its own.
|
||||
*
|
||||
* So read the fields as a READ, not as a test run: `ok` and `delay_ms` are the
|
||||
* observatory's verdict for the path traffic actually takes, and `tested_unix`
|
||||
* (router clock, seconds) is when the OBSERVATORY took that measurement — which
|
||||
* can be a few seconds before the refresh was asked for.
|
||||
*
|
||||
* `ok:true` with an EMPTY `exit_ip`/`exit_country` is a valid, successful result,
|
||||
* not a partial failure: the delay was measured but the exit address could not be
|
||||
@@ -1395,10 +1528,23 @@ export function getGroupsHealth(
|
||||
*
|
||||
* Chains ride the same endpoint. For a chain row, `group` carries the CHAIN's
|
||||
* name and `selected` the node its last group hop picked ('' when the exit hop
|
||||
* isn't a group). Everything else reads the same way.
|
||||
* isn't a group). Per-hop detail is a different read: {@link ChainHopHealth}.
|
||||
*
|
||||
* `ok:false` ⇒ the test failed and `error` carries the human reason; every other
|
||||
* field is meaningless. `tested_unix` is the router's clock, in seconds.
|
||||
* `ok:false` ⇒ there is no usable measurement and `error` carries the human
|
||||
* reason; every other field is meaningless. Four of those reasons are about the
|
||||
* observatory rather than the path, and must not be rendered as "your target is
|
||||
* broken":
|
||||
*
|
||||
* "not routed by any enabled rule, so nothing measures it — the observatory
|
||||
* only probes paths the rules use"
|
||||
* "the observatory has not reached this target yet — it refreshes on the
|
||||
* global probe interval"
|
||||
* "background probing is disabled, so there is nothing to measure this target
|
||||
* with"
|
||||
* "the observatory's probe through this path failed"
|
||||
*
|
||||
* Only the last one is a health finding. The first three say the measurement
|
||||
* does not exist, which is a different thing and a different fix.
|
||||
*/
|
||||
export interface GroupTestResult {
|
||||
group: string // group name — or a chain name for a chain row
|
||||
@@ -1415,7 +1561,11 @@ export interface GroupTestResult {
|
||||
* GET /api/groups/test — progress plus every result so far. `results` is ALWAYS
|
||||
* an array (never null); `done`/`total` count finished vs targeted groups and
|
||||
* chains while `running` is true. Idle reads `{running:false}` with the last
|
||||
* run's results still attached, so a reload after a test still shows what it found.
|
||||
* run's results still attached, so a reload still shows what was last read.
|
||||
*
|
||||
* "Running" means the observatory is working through an out-of-turn refresh pass
|
||||
* over the named targets and this endpoint is collecting what it measures. It is
|
||||
* not the panel dialling anything.
|
||||
*/
|
||||
export interface GroupTestStatus {
|
||||
running: boolean
|
||||
@@ -1448,10 +1598,16 @@ export interface GroupTestStart {
|
||||
}
|
||||
|
||||
/**
|
||||
* POST /api/groups/test — measure a target's delay and exit address. Pass a
|
||||
* group or chain name to test one; pass nothing (or '') to test every group
|
||||
* and every chain. Singleton: a second call while a run is in flight resolves
|
||||
* to `{started:false, reason:'already running'}` rather than failing.
|
||||
* POST /api/groups/test — ask the observatory for an out-of-turn refresh pass,
|
||||
* then report what it measured. Pass a group or chain name to refresh one; pass
|
||||
* nothing (or '') for every group and every chain.
|
||||
*
|
||||
* It does NOT dial. The observatory is the only thing in the daemon that
|
||||
* measures anything, and it measures along the real dial path — so this is the
|
||||
* "don't wait for the next probe interval" button, not a second opinion. The
|
||||
* numbers it returns are the same numbers the cards are already showing, just
|
||||
* fresher. Singleton: a second call while a pass is in flight resolves to
|
||||
* `{started:false, reason:'already running'}` rather than failing.
|
||||
*/
|
||||
export function postGroupsTest(name = ''): Promise<GroupTestStart> {
|
||||
return MOCK
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
/* Buttons — mono, uppercase. .btn is ghost; .btn.primary is solid orange. */
|
||||
/* Buttons — mono, uppercase. .btn is ghost; .btn.primary is solid orange;
|
||||
* .btn.crit is the solid-red destructive commit. */
|
||||
.btn {
|
||||
display: inline-block;
|
||||
padding: 7px 12px;
|
||||
@@ -30,3 +31,16 @@
|
||||
color: #fff;
|
||||
filter: brightness(1.05);
|
||||
}
|
||||
|
||||
/* Destructive commit. The fill is crit stepped a little toward black so white
|
||||
* label text clears 4.5:1 in BOTH themes — the raw --crit is bright enough in
|
||||
* dark mode to fall under it. Red here always means "this removes something". */
|
||||
.btn.crit {
|
||||
border-color: transparent;
|
||||
background: color-mix(in srgb, var(--crit) 88%, #000);
|
||||
color: #fff;
|
||||
}
|
||||
.btn.crit:hover {
|
||||
color: #fff;
|
||||
filter: brightness(1.08);
|
||||
}
|
||||
|
||||
@@ -1,19 +1,27 @@
|
||||
import './Button.css'
|
||||
import { forwardRef } from 'react'
|
||||
import type { ButtonHTMLAttributes } from 'react'
|
||||
|
||||
export interface ButtonProps extends ButtonHTMLAttributes<HTMLButtonElement> {
|
||||
/** `primary` is the solid-orange call to action; `ghost` is the default. */
|
||||
variant?: 'ghost' | 'primary'
|
||||
/**
|
||||
* `primary` is the solid-orange call to action; `crit` is the solid-red
|
||||
* destructive commit (delete, remove) — semantic crit, never the accent;
|
||||
* `ghost` is the default.
|
||||
*/
|
||||
variant?: 'ghost' | 'primary' | 'crit'
|
||||
}
|
||||
|
||||
export function Button({ variant = 'ghost', className, type, ...rest }: ButtonProps) {
|
||||
/** Ref-forwarding so a dialog can park focus on a specific button. */
|
||||
export const Button = forwardRef<HTMLButtonElement, ButtonProps>(function Button(
|
||||
{ variant = 'ghost', className, type, ...rest },
|
||||
ref,
|
||||
) {
|
||||
return (
|
||||
<button
|
||||
ref={ref}
|
||||
type={type ?? 'button'}
|
||||
className={['btn', variant === 'primary' ? 'primary' : '', className]
|
||||
.filter(Boolean)
|
||||
.join(' ')}
|
||||
className={['btn', variant === 'ghost' ? '' : variant, className].filter(Boolean).join(' ')}
|
||||
{...rest}
|
||||
/>
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
@@ -0,0 +1,162 @@
|
||||
/* <ConfirmDialog> — the safety interlock plate.
|
||||
*
|
||||
* This replaces the browser's native confirm dialog, which a browser can mute for
|
||||
* good ("prevent this page from creating additional dialogs"): after that it
|
||||
* returns false with no dialog at all, so every delete button in the panel goes
|
||||
* dead and silent with no way to recover short of a page reload. We draw the
|
||||
* plate ourselves, so nothing can suppress it.
|
||||
*
|
||||
* Faceplate language: a small rack module lifted off the panel — corner screws
|
||||
* (reused from Faceplate.css), an engraved label, a groove above the actions.
|
||||
* Destructive intent is carried by the crit semantic, never by the orange accent:
|
||||
* accent means "this control is active", crit means "this destroys something".
|
||||
*/
|
||||
|
||||
/* The veil is a fixed dark wash in both themes — a light scrim over a light
|
||||
* panel would not read as "the panel is out of reach". Follows the tokens.css
|
||||
* pattern: light base, dark via media query, data-theme overrides win both ways. */
|
||||
.cfm-scrim {
|
||||
--cfm-veil: rgba(33, 29, 21, 0.52);
|
||||
}
|
||||
@media (prefers-color-scheme: dark) {
|
||||
.cfm-scrim {
|
||||
--cfm-veil: rgba(0, 0, 0, 0.66);
|
||||
}
|
||||
}
|
||||
:root[data-theme='light'] .cfm-scrim {
|
||||
--cfm-veil: rgba(33, 29, 21, 0.52);
|
||||
}
|
||||
:root[data-theme='dark'] .cfm-scrim {
|
||||
--cfm-veil: rgba(0, 0, 0, 0.66);
|
||||
}
|
||||
|
||||
.cfm-scrim {
|
||||
position: fixed;
|
||||
inset: 0;
|
||||
z-index: 200;
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
/* Short viewports: the plate scrolls with the veil instead of being clipped. */
|
||||
overflow-y: auto;
|
||||
padding: calc(var(--u, 8px) * 2);
|
||||
background: var(--cfm-veil);
|
||||
animation: cfm-veil-in 0.14s ease-out;
|
||||
}
|
||||
|
||||
.cfm-card {
|
||||
position: relative;
|
||||
width: min(32rem, 100%);
|
||||
max-height: calc(100dvh - var(--u, 8px) * 4);
|
||||
overflow-y: auto;
|
||||
padding: calc(var(--u, 8px) * 3.25);
|
||||
border: 1px solid var(--groove);
|
||||
border-radius: 12px;
|
||||
/* same brushed plate as <Faceplate>, one step brighter so it reads as lifted */
|
||||
background:
|
||||
repeating-linear-gradient(
|
||||
90deg,
|
||||
transparent 0 2px,
|
||||
color-mix(in srgb, var(--edge) 30%, transparent) 2px 3px
|
||||
),
|
||||
linear-gradient(180deg, var(--raised), color-mix(in srgb, var(--raised) 82%, var(--panel)));
|
||||
box-shadow:
|
||||
0 1px 0 var(--edge) inset,
|
||||
0 30px 60px -22px var(--shadow),
|
||||
0 4px 12px var(--shadow);
|
||||
animation: cfm-card-in 0.18s cubic-bezier(0.2, 0.7, 0.3, 1);
|
||||
}
|
||||
.cfm-card:focus {
|
||||
outline: none;
|
||||
}
|
||||
|
||||
/* `still` is set from usePrefersReducedMotion — the plate appears, it never
|
||||
* travels. (The global reduced-motion rule in tokens.css also neutralises the
|
||||
* duration; this keeps the intent explicit at the component.) */
|
||||
.cfm-scrim.still,
|
||||
.cfm-scrim.still .cfm-card {
|
||||
animation: none;
|
||||
}
|
||||
|
||||
@keyframes cfm-veil-in {
|
||||
from {
|
||||
opacity: 0;
|
||||
}
|
||||
to {
|
||||
opacity: 1;
|
||||
}
|
||||
}
|
||||
@keyframes cfm-card-in {
|
||||
from {
|
||||
opacity: 0;
|
||||
transform: translateY(6px) scale(0.99);
|
||||
}
|
||||
to {
|
||||
opacity: 1;
|
||||
transform: none;
|
||||
}
|
||||
}
|
||||
|
||||
/* ---- header: engraved label + state LED ---- */
|
||||
.cfm-hd {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 10px;
|
||||
margin-bottom: calc(var(--u, 8px) * 1.5);
|
||||
}
|
||||
.cfm-label {
|
||||
flex: 1;
|
||||
font-family: var(--font-mono);
|
||||
font-size: 10px;
|
||||
letter-spacing: var(--track-label-wide, 0.24em);
|
||||
color: var(--dim);
|
||||
text-transform: uppercase;
|
||||
}
|
||||
|
||||
/* ---- copy ---- */
|
||||
.cfm-title {
|
||||
margin: 0;
|
||||
font-family: var(--font-mono);
|
||||
font-weight: 700;
|
||||
font-size: 17px;
|
||||
line-height: 1.35;
|
||||
color: var(--ink);
|
||||
/* names can be long and unbroken — wrap rather than push the plate wide */
|
||||
overflow-wrap: anywhere;
|
||||
}
|
||||
.cfm-body {
|
||||
margin: calc(var(--u, 8px) * 1.5) 0 0;
|
||||
max-width: 52ch;
|
||||
font-family: var(--font-sans);
|
||||
font-size: 13.5px;
|
||||
line-height: 1.6;
|
||||
color: var(--dim);
|
||||
overflow-wrap: anywhere;
|
||||
}
|
||||
|
||||
/* ---- action bar ---- */
|
||||
.cfm-actions {
|
||||
display: flex;
|
||||
justify-content: flex-end;
|
||||
gap: calc(var(--u, 8px));
|
||||
margin-top: calc(var(--u, 8px) * 3);
|
||||
padding-top: calc(var(--u, 8px) * 2);
|
||||
border-top: 1px solid var(--groove);
|
||||
}
|
||||
|
||||
@media (max-width: 420px) {
|
||||
.cfm-card {
|
||||
padding: calc(var(--u, 8px) * 2.5);
|
||||
}
|
||||
.cfm-actions {
|
||||
flex-wrap: wrap;
|
||||
}
|
||||
.cfm-actions .btn {
|
||||
flex: 1 1 auto;
|
||||
text-align: center;
|
||||
}
|
||||
/* screws crowd a small plate — drop them rather than collide with the copy */
|
||||
.cfm-card > .screw {
|
||||
display: none;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,281 @@
|
||||
import './ConfirmDialog.css'
|
||||
import {
|
||||
createContext,
|
||||
useCallback,
|
||||
useContext,
|
||||
useEffect,
|
||||
useId,
|
||||
useRef,
|
||||
useState,
|
||||
} from 'react'
|
||||
import type { ReactNode } from 'react'
|
||||
import { createPortal } from 'react-dom'
|
||||
import { Button } from './Button'
|
||||
import { Led } from './Led'
|
||||
import { usePrefersReducedMotion } from './usePrefersReducedMotion'
|
||||
|
||||
/**
|
||||
* How the confirming button is painted.
|
||||
*
|
||||
* crit — the action destroys something. Semantic crit, never the accent.
|
||||
* neutral — the action is a normal commit the operator should read first
|
||||
* (a warning before saving); the accent's call-to-action is correct.
|
||||
*/
|
||||
export type ConfirmTone = 'crit' | 'neutral'
|
||||
|
||||
export interface ConfirmOptions {
|
||||
/** Engraved eyebrow, e.g. "DELETE RULE". Names the operation, not the object. */
|
||||
label?: string
|
||||
/** The question. One line, ends in "?". */
|
||||
title: string
|
||||
/** The consequence — what changes on the router if this goes through. */
|
||||
body?: ReactNode
|
||||
/** Verb on the confirming button. Defaults to "Delete". */
|
||||
confirmLabel?: string
|
||||
/** Verb on the dismissing button. Defaults to "Cancel". */
|
||||
cancelLabel?: string
|
||||
/** Defaults to `crit` — the overwhelmingly common case is a delete. */
|
||||
tone?: ConfirmTone
|
||||
}
|
||||
|
||||
export interface ConfirmDialogProps extends ConfirmOptions {
|
||||
open: boolean
|
||||
/** Called exactly once per dialog, with the operator's answer. */
|
||||
onResolve: (confirmed: boolean) => void
|
||||
}
|
||||
|
||||
const FOCUSABLE =
|
||||
'button:not([disabled]), [href], input:not([disabled]), select:not([disabled]), textarea:not([disabled]), [tabindex]:not([tabindex="-1"])'
|
||||
|
||||
/**
|
||||
* The modal plate itself. Normally reached through `useConfirm()`; exported so a
|
||||
* page that wants to own the open state can render it directly.
|
||||
*
|
||||
* Keyboard contract:
|
||||
* - focus moves to Cancel on open, so a reflex Enter dismisses, never deletes;
|
||||
* - Tab / Shift+Tab cycle inside the plate and cannot reach the page behind it;
|
||||
* - Esc answers "no";
|
||||
* - on close, focus returns to whatever opened the dialog.
|
||||
*/
|
||||
export function ConfirmDialog({
|
||||
open,
|
||||
onResolve,
|
||||
label,
|
||||
title,
|
||||
body,
|
||||
confirmLabel = 'Delete',
|
||||
cancelLabel = 'Cancel',
|
||||
tone = 'crit',
|
||||
}: ConfirmDialogProps) {
|
||||
const titleId = useId()
|
||||
const bodyId = useId()
|
||||
const cardRef = useRef<HTMLDivElement>(null)
|
||||
const cancelRef = useRef<HTMLButtonElement>(null)
|
||||
const openerRef = useRef<HTMLElement | null>(null)
|
||||
const reduced = usePrefersReducedMotion()
|
||||
|
||||
// Take the page out of the tab order, park focus on Cancel, and hand focus
|
||||
// back to the opener when the plate goes away.
|
||||
useEffect(() => {
|
||||
if (!open) return
|
||||
const opener = document.activeElement
|
||||
openerRef.current = opener instanceof HTMLElement ? opener : null
|
||||
|
||||
const prevOverflow = document.body.style.overflow
|
||||
document.body.style.overflow = 'hidden'
|
||||
|
||||
// Cancel is the resting place: an Enter or a Space meant for the page lands
|
||||
// on "no". The destructive button is one Tab away, deliberately.
|
||||
;(cancelRef.current ?? cardRef.current)?.focus()
|
||||
|
||||
return () => {
|
||||
document.body.style.overflow = prevOverflow
|
||||
const back = openerRef.current
|
||||
openerRef.current = null
|
||||
if (back && document.contains(back)) back.focus()
|
||||
}
|
||||
}, [open])
|
||||
|
||||
// Esc answers no; Tab is caged. Capture phase so a page-level key handler
|
||||
// never sees keys aimed at the dialog.
|
||||
useEffect(() => {
|
||||
if (!open) return
|
||||
const onKey = (e: KeyboardEvent) => {
|
||||
if (e.key === 'Escape') {
|
||||
e.preventDefault()
|
||||
e.stopPropagation()
|
||||
onResolve(false)
|
||||
return
|
||||
}
|
||||
if (e.key !== 'Tab') return
|
||||
const card = cardRef.current
|
||||
if (!card) return
|
||||
const list = Array.from(card.querySelectorAll<HTMLElement>(FOCUSABLE))
|
||||
if (list.length === 0) {
|
||||
e.preventDefault()
|
||||
card.focus()
|
||||
return
|
||||
}
|
||||
const first = list[0]
|
||||
const last = list[list.length - 1]
|
||||
const active = document.activeElement as HTMLElement | null
|
||||
if (!active || !card.contains(active)) {
|
||||
e.preventDefault()
|
||||
;(e.shiftKey ? last : first).focus()
|
||||
} else if (e.shiftKey && active === first) {
|
||||
e.preventDefault()
|
||||
last.focus()
|
||||
} else if (!e.shiftKey && active === last) {
|
||||
e.preventDefault()
|
||||
first.focus()
|
||||
}
|
||||
}
|
||||
document.addEventListener('keydown', onKey, true)
|
||||
return () => document.removeEventListener('keydown', onKey, true)
|
||||
}, [open, onResolve])
|
||||
|
||||
if (!open) return null
|
||||
|
||||
return createPortal(
|
||||
<div
|
||||
className={['cfm-scrim', reduced ? 'still' : ''].filter(Boolean).join(' ')}
|
||||
// A click on the field around the plate means "not now". Mousedown (not
|
||||
// click) so a text selection dragged out of the plate can't dismiss it.
|
||||
onMouseDown={(e) => {
|
||||
if (e.target === e.currentTarget) onResolve(false)
|
||||
}}
|
||||
>
|
||||
<div
|
||||
className={`cfm-card tone-${tone}`}
|
||||
ref={cardRef}
|
||||
tabIndex={-1}
|
||||
role="alertdialog"
|
||||
aria-modal="true"
|
||||
aria-labelledby={titleId}
|
||||
aria-describedby={body != null ? bodyId : undefined}
|
||||
>
|
||||
<i className="screw tl" aria-hidden="true" />
|
||||
<i className="screw tr" aria-hidden="true" />
|
||||
<i className="screw bl" aria-hidden="true" />
|
||||
<i className="screw br" aria-hidden="true" />
|
||||
|
||||
{/* Lamp first, then the engraved label — the way a real panel reads, and
|
||||
it keeps the LED off the corner screw. */}
|
||||
<div className="cfm-hd">
|
||||
<Led variant={tone === 'crit' ? 'crit' : 'amber'} />
|
||||
<span className="cfm-label">{label ?? (tone === 'crit' ? 'Confirm delete' : 'Confirm')}</span>
|
||||
</div>
|
||||
|
||||
<h2 className="cfm-title" id={titleId}>
|
||||
{title}
|
||||
</h2>
|
||||
{body != null && (
|
||||
<p className="cfm-body" id={bodyId}>
|
||||
{body}
|
||||
</p>
|
||||
)}
|
||||
|
||||
<div className="cfm-actions">
|
||||
<Button ref={cancelRef} onClick={() => onResolve(false)}>
|
||||
{cancelLabel}
|
||||
</Button>
|
||||
<Button variant={tone === 'crit' ? 'crit' : 'primary'} onClick={() => onResolve(true)}>
|
||||
{confirmLabel}
|
||||
</Button>
|
||||
</div>
|
||||
</div>
|
||||
</div>,
|
||||
document.body,
|
||||
)
|
||||
}
|
||||
|
||||
// ---- provider + hook --------------------------------------------------------
|
||||
|
||||
interface Request extends ConfirmOptions {
|
||||
id: number
|
||||
resolve: (v: boolean) => void
|
||||
}
|
||||
|
||||
const ConfirmCtx = createContext<((o: ConfirmOptions) => Promise<boolean>) | null>(null)
|
||||
|
||||
/**
|
||||
* Mount once at the app root. Everything below can then ask a question and await
|
||||
* the answer.
|
||||
*/
|
||||
export function ConfirmProvider({ children }: { children: ReactNode }) {
|
||||
const [req, setReq] = useState<Request | null>(null)
|
||||
const pending = useRef<Request | null>(null)
|
||||
const seq = useRef(0)
|
||||
|
||||
const confirm = useCallback(
|
||||
(opts: ConfirmOptions) =>
|
||||
new Promise<boolean>((resolve) => {
|
||||
// A second question while one is open answers the first with "no" rather
|
||||
// than leaving its promise — and its caller — hanging forever.
|
||||
pending.current?.resolve(false)
|
||||
seq.current += 1
|
||||
const next: Request = { ...opts, id: seq.current, resolve }
|
||||
pending.current = next
|
||||
setReq(next)
|
||||
}),
|
||||
[],
|
||||
)
|
||||
|
||||
const settle = useCallback((confirmed: boolean) => {
|
||||
const open = pending.current
|
||||
pending.current = null
|
||||
setReq(null)
|
||||
open?.resolve(confirmed)
|
||||
}, [])
|
||||
|
||||
// Teardown must not strand a caller mid-await.
|
||||
useEffect(
|
||||
() => () => {
|
||||
pending.current?.resolve(false)
|
||||
pending.current = null
|
||||
},
|
||||
[],
|
||||
)
|
||||
|
||||
// A question belongs to the page that asked it. The provider outlives the
|
||||
// hash router, so a navigation would otherwise leave a stale plate floating
|
||||
// over a page it has nothing to do with — answer it "no" and clear it.
|
||||
useEffect(() => {
|
||||
const onNav = () => {
|
||||
if (pending.current) settle(false)
|
||||
}
|
||||
window.addEventListener('hashchange', onNav)
|
||||
return () => window.removeEventListener('hashchange', onNav)
|
||||
}, [settle])
|
||||
|
||||
return (
|
||||
<ConfirmCtx.Provider value={confirm}>
|
||||
{children}
|
||||
{req !== null && <ConfirmDialog key={req.id} open onResolve={settle} {...req} />}
|
||||
</ConfirmCtx.Provider>
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* Ask the operator, get a definite answer:
|
||||
*
|
||||
* const confirm = useConfirm()
|
||||
* if (!(await confirm({ title: 'Delete rule "x"?', body: '…' }))) return
|
||||
*
|
||||
* The returned function is stable, so it is safe in a useCallback dep list. It
|
||||
* always settles — cancel, Esc, click-outside and teardown all resolve `false`;
|
||||
* only the confirming button resolves `true`.
|
||||
*
|
||||
* Name it `confirm` at the call site on purpose: the local binding shadows the
|
||||
* global one inside that component, so an accidental bare `confirm(...)` cannot
|
||||
* reach the suppressible native dialog.
|
||||
*/
|
||||
export function useConfirm(): (o: ConfirmOptions) => Promise<boolean> {
|
||||
const ctx = useContext(ConfirmCtx)
|
||||
if (!ctx) {
|
||||
// Loud on purpose. A fallback that quietly resolved false would rebuild the
|
||||
// exact bug this component exists to kill.
|
||||
throw new Error('useConfirm() needs <ConfirmProvider> above it (mounted in main.tsx)')
|
||||
}
|
||||
return ctx
|
||||
}
|
||||
@@ -17,6 +17,8 @@ export { Button } from './Button'
|
||||
export type { ButtonProps } from './Button'
|
||||
export { Select } from './Select'
|
||||
export type { SelectProps, SelectOption } from './Select'
|
||||
export { ConfirmDialog, ConfirmProvider, useConfirm } from './ConfirmDialog'
|
||||
export type { ConfirmDialogProps, ConfirmOptions, ConfirmTone } from './ConfirmDialog'
|
||||
export { Clock } from './Clock'
|
||||
export { CatSuggest } from './CatSuggest'
|
||||
export { SrcPicker } from './SrcPicker'
|
||||
|
||||
+6
-1
@@ -2,12 +2,17 @@ import { StrictMode } from 'react'
|
||||
import { createRoot } from 'react-dom/client'
|
||||
import './tokens.css'
|
||||
import { App } from './App'
|
||||
import { ConfirmProvider } from './components'
|
||||
|
||||
const rootEl = document.getElementById('root')
|
||||
if (!rootEl) throw new Error('#root not found')
|
||||
|
||||
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
|
||||
// unauth / no-link plates) — useConfirm() can never find itself without a host.
|
||||
createRoot(rootEl).render(
|
||||
<StrictMode>
|
||||
<App />
|
||||
<ConfirmProvider>
|
||||
<App />
|
||||
</ConfirmProvider>
|
||||
</StrictMode>,
|
||||
)
|
||||
|
||||
+200
-33
@@ -6,7 +6,7 @@
|
||||
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
|
||||
//
|
||||
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
|
||||
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
|
||||
import type { ApplyResult, ChainHealth, ChainHopHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, Profile, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning, Traffic } from './api'
|
||||
|
||||
let armed = false // a pending commit-confirm auto-rollback
|
||||
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
|
||||
@@ -129,9 +129,25 @@ const CONFIG: Model = {
|
||||
{ Name: 'via-tunnel', Source: 'subscription', Subscription: 'primary', Strategy: 'leastping', Egress: 'awg' },
|
||||
{ Name: 'fallback', Source: 'subscription', Subscription: 'backup', Strategy: 'roundrobin', Egress: '' },
|
||||
],
|
||||
// One multi-hop chain so `?mock` exercises the chain card's Test button and
|
||||
// its result readout: enters through the awg tunnel, exits via the auto group.
|
||||
Chains: [{ Name: 'relay', Hops: ['egress:awg', 'group:auto'] }],
|
||||
// Three chains, one per state the hop readout has to render.
|
||||
Chains: [
|
||||
// The owner's real production shape: leave through a WAN interface, cross an
|
||||
// AmneziaWG node, then three subscription groups in series. The leading
|
||||
// `egress:` is NOT a numbered hop — the daemon lifts it into hop 1's entry
|
||||
// detour — so this reports FOUR hops, and hop 3 is dead while its neighbours
|
||||
// answer. That single red notch in the middle of a live path is the entire
|
||||
// reason per-hop health exists, so `?mock` must show it at a glance.
|
||||
{
|
||||
Name: 'ewan-wg-subs',
|
||||
Hops: ['egress:wan', 'node:home-wg', 'group:auto', 'group:stealth', 'group:via-tunnel'],
|
||||
},
|
||||
// Used, but the observatory hasn't come round yet — every hop untested. Not
|
||||
// dead and not healthy: the state the panel most easily renders as a fault.
|
||||
{ Name: 'sub-fresh', Hops: ['node:home-wg', 'group:fallback'] },
|
||||
// No enabled rule targets it, so the observatory skips it entirely and the
|
||||
// daemon never materialises its hops: `used:false` and NO `hops` key.
|
||||
{ Name: 'relay', Hops: ['egress:awg', 'group:auto'] },
|
||||
],
|
||||
Egresses: [
|
||||
{ Name: 'wan', Type: 'interface', Interface: 'wan' },
|
||||
// An AmneziaWG tunnel — the whole point of a group-level egress binding.
|
||||
@@ -145,6 +161,13 @@ const CONFIG: Model = {
|
||||
Rules: [
|
||||
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
|
||||
{ Name: 'ru-bypass', Enabled: true, Order: 20, DstRuleset: ['ru-inside'], Target: 'direct' },
|
||||
// These two are what make the chains USED — the observatory probes only the
|
||||
// paths an enabled rule can reach, so without them every chain card would
|
||||
// read "not routed" and the hop rail would never appear in `?mock`. Kept
|
||||
// ABOVE the condition-less rule at Order 40, which would otherwise swallow
|
||||
// everything below it and mark them "never applies".
|
||||
{ Name: 'media-via-chain', Enabled: true, Order: 22, DstRuleset: ['yt-geosite'], Target: 'chain:ewan-wg-subs' },
|
||||
{ Name: 'spare-via-chain', Enabled: true, Order: 24, DstRuleset: ['ad-hosts'], Target: 'chain:sub-fresh' },
|
||||
{ Name: 'private-direct', Enabled: true, Order: 30, DstRuleset: ['private-nets'], Target: 'direct' },
|
||||
// A SECOND condition-less rule, above the real default. It reads like a working
|
||||
// rule and does nothing: a rule with no conditions becomes the router's default,
|
||||
@@ -305,16 +328,61 @@ const RULESET_STATUS: RulesetStatus[] = [
|
||||
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
|
||||
* a rule with no conditions is the router's default, and the LAST such rule by
|
||||
* Order wins — every earlier one can never apply. It reads the live CONFIG so
|
||||
* edits made in `?mock` keep the badge honest. */
|
||||
* edits made in `?mock` keep the badge honest.
|
||||
*
|
||||
* It also mirrors model.ResolveActiveProfile + ApplyProfileRuleOverrides, because
|
||||
* `effective_enabled` is the whole point of the endpoint: CONFIG's `mobile-uplink`
|
||||
* is active and both enables and disables rules, so `?mock` shows the same
|
||||
* desired-vs-effective split the field config does. */
|
||||
export async function getRulesReachability(): Promise<RulesReachability> {
|
||||
await wait(60)
|
||||
const rules = CONFIG.Rules ?? []
|
||||
|
||||
// A pin naming an existing, ENABLED profile wins outright. Otherwise auto-select:
|
||||
// highest Priority among enabled profiles, ties by Name, skipping any with an
|
||||
// iface condition (the WAN watcher owns those and expresses its verdict as the pin).
|
||||
const profiles = CONFIG.Profiles ?? []
|
||||
const pinned = String(CONFIG.Globals?.ActiveProfile ?? '').trim()
|
||||
let prof: Profile | null = profiles.find((p) => p.Enabled && p.Name === pinned) ?? null
|
||||
if (!prof) {
|
||||
for (const p of profiles) {
|
||||
if (!p.Enabled || (p.MatchIface ?? []).length > 0) continue
|
||||
const pp = p.Priority ?? 0
|
||||
const bp = prof?.Priority ?? 0
|
||||
if (!prof || pp > bp || (pp === bp && p.Name < prof.Name)) prof = p
|
||||
}
|
||||
}
|
||||
// Enable first, then Disable, so a name in both ends up disabled (Disable wins).
|
||||
const effective = rules.map((r) => Boolean(r.Enabled))
|
||||
if (prof) {
|
||||
const force = (names: string[] | null | undefined, on: boolean) => {
|
||||
for (const raw of names ?? []) {
|
||||
const n = raw.trim()
|
||||
rules.forEach((r, i) => {
|
||||
if (r.Name === n) effective[i] = on
|
||||
})
|
||||
}
|
||||
}
|
||||
force(prof.EnableRules, true)
|
||||
force(prof.DisableRules, false)
|
||||
}
|
||||
const activeProfile = prof
|
||||
|
||||
const out: RuleReach[] = rules.map((r, index) => ({
|
||||
index,
|
||||
name: String(r.Name ?? ''),
|
||||
order: Number(r.Order ?? 0),
|
||||
unreachable: false,
|
||||
shadowed_by_index: -1,
|
||||
effective_enabled: effective[index],
|
||||
// Annotate only where the profile actually FLIPPED the outcome — a profile that
|
||||
// disables an already-off rule has overridden nothing the operator can see.
|
||||
...(activeProfile && effective[index] !== Boolean(r.Enabled)
|
||||
? {
|
||||
overridden_by: activeProfile.Name,
|
||||
override: effective[index] ? ('enabled' as const) : ('disabled' as const),
|
||||
}
|
||||
: {}),
|
||||
}))
|
||||
const conditionless = (r: (typeof rules)[number]): boolean =>
|
||||
!(r.Src ?? []).length &&
|
||||
@@ -325,7 +393,10 @@ export async function getRulesReachability(): Promise<RulesReachability> {
|
||||
String(r.Target ?? '').trim() || (r.Egress ? `egress:${String(r.Egress).trim()}` : '')
|
||||
const defaults = rules
|
||||
.map((r, index) => ({ r, index }))
|
||||
.filter(({ r }) => r.Enabled && conditionless(r) && target(r))
|
||||
// The EFFECTIVE flag, not the configured one: a rule the active profile
|
||||
// switched off is not in force and cannot retire anything (model's
|
||||
// RuleReachability runs over the effective set for the same reason).
|
||||
.filter(({ r, index }) => effective[index] && conditionless(r) && target(r))
|
||||
.sort((a, b) => Number(a.r.Order ?? 0) - Number(b.r.Order ?? 0) || a.index - b.index)
|
||||
const winner = defaults[defaults.length - 1]
|
||||
if (winner) {
|
||||
@@ -405,6 +476,11 @@ export async function getRulesetCategories(source: string): Promise<RulesetCateg
|
||||
// ?mock&warn=1 → a full warning set (critical + warning + info) on top
|
||||
// ?mock&ks=open → healthy plane but a FAIL-OPEN kill-switch, which is what
|
||||
// makes the untunnelable policy inert (F8 case 4)
|
||||
// ?mock&traffic=… → with the plane FULL, where the traffic actually ends up:
|
||||
// split | direct | blocked | blackout | unknown. `direct` is
|
||||
// the field case the readout used to call "Protected" (one
|
||||
// rule, `default → direct`); `unknown` is a daemon too old to
|
||||
// report. Default: tunnel.
|
||||
function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSwitch: string } {
|
||||
const q = typeof location === 'undefined' ? '' : location.search
|
||||
const params = new URLSearchParams(q)
|
||||
@@ -416,6 +492,30 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
|
||||
return { plane: 'full', engine: true, killSwitch }
|
||||
}
|
||||
|
||||
// The daemon's verdict on where traffic goes (apply.Status.traffic). Only
|
||||
// meaningful with the plane installed: with the engine down there is no running
|
||||
// config to judge, and the daemon reports the unknown/zero value — so do the same
|
||||
// here rather than leaving a stale "tunnel" behind a dead engine.
|
||||
function mockTraffic(plane: 'full' | 'hold' | 'none'): Traffic | undefined {
|
||||
if (plane !== 'full') return { verdict: '', default: '', tunnel_rules: 0 }
|
||||
const params = new URLSearchParams(typeof location === 'undefined' ? '' : location.search)
|
||||
switch (params.get('traffic')) {
|
||||
case 'split':
|
||||
return { verdict: 'split', default: 'direct', tunnel_rules: 3 }
|
||||
case 'direct':
|
||||
return { verdict: 'direct', default: 'direct', tunnel_rules: 0 }
|
||||
case 'blocked':
|
||||
return { verdict: 'blocked', default: 'block', tunnel_rules: 2 }
|
||||
case 'blackout':
|
||||
return { verdict: 'blocked', default: 'block', tunnel_rules: 0 }
|
||||
case 'unknown':
|
||||
// A daemon that predates the field sends no `traffic` at all.
|
||||
return undefined
|
||||
default:
|
||||
return { verdict: 'tunnel', default: 'auto', tunnel_rules: 1 }
|
||||
}
|
||||
}
|
||||
|
||||
const MOCK_WARNINGS: StatusWarning[] = [
|
||||
{
|
||||
severity: 'critical',
|
||||
@@ -518,6 +618,7 @@ export async function getStatus(): Promise<Status> {
|
||||
can_rollback: armed || hasLastGood,
|
||||
engine_running: engine,
|
||||
plane,
|
||||
traffic: mockTraffic(plane),
|
||||
warnings: mockWarnings(killSwitch),
|
||||
// Process uptime. Anchored to when this tab loaded plus a fixed head start, so
|
||||
// the reading ticks forward across polls exactly like the real daemon's does.
|
||||
@@ -1094,14 +1195,47 @@ function healthList(): GroupHealth[] {
|
||||
return (CONFIG.Groups ?? []).map((g) => summarise(g.Name, GROUP_MEMBERS.get(g.Name) ?? []))
|
||||
}
|
||||
|
||||
/** Per-chain reachability for the Targets page's "unused" badge (plan §5.E) — the
|
||||
* chain analogue of healthList's `used` field. The mock's single chain `relay` is
|
||||
* NOT referenced by any rule in CONFIG.Rules (they target group:auto / block /
|
||||
* direct), so it reads used=false and its card renders "unused" — exactly the case
|
||||
* the badge exists to surface. A stopped engine reports no chains. */
|
||||
/**
|
||||
* Per-hop health, keyed by chain name — what the observatory measured at each
|
||||
* position of the path, in WIRE order.
|
||||
*
|
||||
* `ewan-wg-subs` is the fixture that matters. Its hop 1 is the AmneziaWG node and
|
||||
* answers; hops 2 and 4 are subscription groups that answer THROUGH it; hop 3 is
|
||||
* a group whose members all time out at that position. Note that hop 4 rolls up
|
||||
* `via-tunnel`, the same group whose standalone card reads 0 of 6 alive — alive
|
||||
* as a hop, dead on its own, both true, because they measure different dial
|
||||
* paths. That contradiction is the whole point of measuring per hop.
|
||||
*
|
||||
* `sub-fresh` is used but never yet reached: every hop untested, nothing dead.
|
||||
* `relay` is absent from this map on purpose — an unused chain is never
|
||||
* materialised, so the daemon sends no `hops` key at all, which is "nothing
|
||||
* measured", not "no hops".
|
||||
*/
|
||||
const CHAIN_HOPS: Record<string, ChainHopHealth[]> = {
|
||||
'ewan-wg-subs': [
|
||||
{ index: 1, tag: 'chain-ewan-wg-subs-h1', kind: 'node', exit: false, state: 'alive', delay_ms: 41, age_seconds: 22, selected: '', total: 1, tested: 1, alive: 1, dead: 0, untested: 0 },
|
||||
{ index: 2, tag: 'chain-ewan-wg-subs-h2', kind: 'group', exit: false, state: 'alive', delay_ms: 96, age_seconds: 18, selected: '🇳🇱 Amsterdam-01', total: 298, tested: 122, alive: 119, dead: 3, untested: 176 },
|
||||
{ index: 3, tag: 'chain-ewan-wg-subs-h3', kind: 'group', exit: false, state: 'dead', delay_ms: 0, age_seconds: 15, selected: '', total: 24, tested: 24, alive: 0, dead: 24, untested: 0 },
|
||||
{ index: 4, tag: 'chain-ewan-wg-subs-h4', kind: 'group', exit: true, state: 'alive', delay_ms: 148, age_seconds: 19, selected: '🇸🇬 Singapore-09', total: 6, tested: 6, alive: 6, dead: 0, untested: 0 },
|
||||
],
|
||||
'sub-fresh': [
|
||||
{ index: 1, tag: 'chain-sub-fresh-h1', kind: 'node', exit: false, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '', total: 1, tested: 0, alive: 0, dead: 0, untested: 1 },
|
||||
{ index: 2, tag: 'chain-sub-fresh-h2', kind: 'group', exit: true, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '', total: 24, tested: 0, alive: 0, dead: 0, untested: 24 },
|
||||
],
|
||||
}
|
||||
|
||||
/** Per-chain reachability plus per-hop health for the Targets page. `used` is the
|
||||
* chain analogue of healthList's field; `hops` is OMITTED (never null, never []),
|
||||
* exactly like the daemon, for a chain the engine never materialised. A stopped
|
||||
* engine reports no chains at all. */
|
||||
function chainHealthList(): ChainHealth[] {
|
||||
if (!mockPlane().engine) return []
|
||||
return (CONFIG.Chains ?? []).map((c) => ({ name: c.Name, used: chainUsed(c.Name) }))
|
||||
return (CONFIG.Chains ?? []).map((c) => {
|
||||
const hops = CHAIN_HOPS[c.Name]
|
||||
const h: ChainHealth = { name: c.Name, used: chainUsed(c.Name) }
|
||||
if (hops) h.hops = hops.map((x) => ({ ...x }))
|
||||
return h
|
||||
})
|
||||
}
|
||||
|
||||
/** A chain is "used" when some enabled routing rule (or Final, or a DNS detour)
|
||||
@@ -1159,15 +1293,23 @@ class ApiErrorLike extends Error {
|
||||
}
|
||||
}
|
||||
|
||||
// Mock group/chain test. Deliberately covers every state the UI has to render,
|
||||
// one per target, so a single offline run exercises all of them:
|
||||
// auto → ok WITH an exit address
|
||||
// stealth → ok WITHOUT one (delay measured, address undeterminable) — a
|
||||
// SUCCESS, and the case the UI most easily gets wrong
|
||||
// relay → the chain: same wire shape, `group` carries the CHAIN's name and
|
||||
// `selected` the node its exit group picked
|
||||
// fallback → a failure carrying a human reason
|
||||
// Results land one per GET poll, so the running/progress state is visible too.
|
||||
// Mock refresh results. The endpoint no longer dials anything: it asks the
|
||||
// observatory to measure out of turn and reports what the observatory found, so
|
||||
// every row here is a READ of a background measurement. Deliberately covers every
|
||||
// state the UI has to render, one per target, so a single offline run exercises
|
||||
// all of them:
|
||||
// auto → ok WITH an exit address
|
||||
// stealth → ok WITHOUT one (delay measured, address undeterminable) — a
|
||||
// SUCCESS, and the case the UI most easily gets wrong
|
||||
// ewan-wg-subs → the chain: same wire shape, `group` carries the CHAIN's name
|
||||
// and `selected` the node its exit hop picked
|
||||
// via-tunnel → the one honest health FAILURE: a probe that ran and failed
|
||||
// fallback,
|
||||
// relay → not routed at all, so no measurement exists to report
|
||||
// sub-fresh → routed, but the observatory hasn't come round yet
|
||||
// The last three are absence of measurement, not a broken target, and the copy
|
||||
// has to keep them apart. Results land one per GET poll, so the running/progress
|
||||
// state is visible too.
|
||||
const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_unix'>> = {
|
||||
auto: {
|
||||
selected: 'nl-reality-2',
|
||||
@@ -1185,32 +1327,57 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
|
||||
ok: true,
|
||||
error: '',
|
||||
},
|
||||
// The chain — Selected is the node the chain's exit group (auto) picked.
|
||||
relay: {
|
||||
selected: 'nl-reality-2',
|
||||
delay_ms: 61,
|
||||
exit_ip: '185.12.34.56',
|
||||
exit_country: 'NL',
|
||||
ok: true,
|
||||
error: '',
|
||||
// The chain, and the pairing that makes the whole feature worth building. A
|
||||
// chain is one series path, so with hop 3 dead the end-to-end probe CANNOT
|
||||
// succeed — this row and the hop rail have to tell one story, not two. The
|
||||
// row says the path is down; the rail says which of the four hops did it,
|
||||
// which is the part nobody could see before.
|
||||
'ewan-wg-subs': {
|
||||
selected: '',
|
||||
delay_ms: 0,
|
||||
exit_ip: '',
|
||||
exit_country: '',
|
||||
ok: false,
|
||||
error: 'the observatory’s probe through this path failed',
|
||||
},
|
||||
// Dead through its tunnel, exactly as its membership health says — the exit
|
||||
// test and the member health tell the same story about the same group.
|
||||
// The one real health failure in the fixture: the observatory's probe ran along
|
||||
// this path and did not come back.
|
||||
'via-tunnel': {
|
||||
selected: '',
|
||||
delay_ms: 0,
|
||||
exit_ip: '',
|
||||
exit_country: '',
|
||||
ok: false,
|
||||
error: 'no member answered through egress awg (6 of 6 timed out)',
|
||||
error: 'the observatory’s probe through this path failed',
|
||||
},
|
||||
// Not a health verdict — nothing routes here, so no measurement of it exists.
|
||||
fallback: {
|
||||
selected: '',
|
||||
delay_ms: 0,
|
||||
exit_ip: '',
|
||||
exit_country: '',
|
||||
ok: false,
|
||||
error: 'no reachable node in the group (all 3 members timed out)',
|
||||
error:
|
||||
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
|
||||
},
|
||||
relay: {
|
||||
selected: '',
|
||||
delay_ms: 0,
|
||||
exit_ip: '',
|
||||
exit_country: '',
|
||||
ok: false,
|
||||
error:
|
||||
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
|
||||
},
|
||||
// Routed, materialised, simply not reached yet. Untested is not dead.
|
||||
'sub-fresh': {
|
||||
selected: '',
|
||||
delay_ms: 0,
|
||||
exit_ip: '',
|
||||
exit_country: '',
|
||||
ok: false,
|
||||
error:
|
||||
'the observatory has not reached this target yet — it refreshes on the global probe interval',
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
+41
-19
@@ -1,6 +1,6 @@
|
||||
import './DNS.css'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import { Button, CatSuggest, Led, SrcPicker, Toggle } from '../components'
|
||||
import { Button, CatSuggest, Led, SrcPicker, Toggle, useConfirm } from '../components'
|
||||
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
|
||||
import type { Alert, DNSRule, Model, Resolver } from '../api'
|
||||
|
||||
@@ -220,6 +220,7 @@ function describeDetour(
|
||||
// ---- page ------------------------------------------------------------------
|
||||
|
||||
export default function DNS() {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<DNSModel | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
|
||||
@@ -422,15 +423,19 @@ export default function DNS() {
|
||||
)
|
||||
|
||||
const removeBlocklist = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = blocklists[idx]
|
||||
if (!window.confirm(`Delete blocklist “${target.Name}”? This removes it from the config.`))
|
||||
return
|
||||
const ok = await confirm({
|
||||
label: 'Delete blocklist',
|
||||
title: `Delete blocklist “${target.Name}”?`,
|
||||
body: 'This removes it from the config.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = blocklists.filter((_, i) => i !== idx)
|
||||
void save({ ...config, Blocklists: next }, `Deleted ${target.Name}`)
|
||||
},
|
||||
[config, blocklists, save],
|
||||
[config, blocklists, save, confirm],
|
||||
)
|
||||
|
||||
// ---- allowlist mutations --------------------------------------------------
|
||||
@@ -457,15 +462,19 @@ export default function DNS() {
|
||||
)
|
||||
|
||||
const removeAllowlist = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = allowlists[idx]
|
||||
if (!window.confirm(`Delete allowlist “${target.Name}”? This removes it from the config.`))
|
||||
return
|
||||
const ok = await confirm({
|
||||
label: 'Delete allowlist',
|
||||
title: `Delete allowlist “${target.Name}”?`,
|
||||
body: 'This removes it from the config.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = allowlists.filter((_, i) => i !== idx)
|
||||
void save({ ...config, Allowlists: next }, `Deleted ${target.Name}`)
|
||||
},
|
||||
[config, allowlists, save],
|
||||
[config, allowlists, save, confirm],
|
||||
)
|
||||
|
||||
// ---- resolver mutations ---------------------------------------------------
|
||||
@@ -492,11 +501,15 @@ export default function DNS() {
|
||||
)
|
||||
|
||||
const removeResolver = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = resolvers[idx]
|
||||
if (!window.confirm(`Delete resolver “${target.Name}”? This removes it from the config.`))
|
||||
return
|
||||
const ok = await confirm({
|
||||
label: 'Delete resolver',
|
||||
title: `Delete resolver “${target.Name}”?`,
|
||||
body: 'This removes it from the config.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = resolvers.filter((_, i) => i !== idx)
|
||||
// Don't leave default/fallback pointing at a resolver that no longer exists.
|
||||
const g = { ...config.Globals }
|
||||
@@ -578,15 +591,19 @@ export default function DNS() {
|
||||
)
|
||||
|
||||
const removeDNSRule = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = dnsRules[idx]
|
||||
if (!window.confirm(`Delete this DNS rule? Matching queries fall back to the default resolver.`))
|
||||
return
|
||||
const ok = await confirm({
|
||||
label: 'Delete DNS rule',
|
||||
title: 'Delete this DNS rule?',
|
||||
body: 'Matching queries fall back to the default resolver.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = dnsRules.filter((_, i) => i !== idx)
|
||||
void save({ ...config, DNSRules: next }, `Deleted DNS rule → ${target.Resolver}`)
|
||||
},
|
||||
[config, dnsRules, save],
|
||||
[config, dnsRules, save, confirm],
|
||||
)
|
||||
|
||||
// ---- alert mutations ------------------------------------------------------
|
||||
@@ -610,14 +627,19 @@ export default function DNS() {
|
||||
)
|
||||
|
||||
const removeAlert = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = alerts[idx]
|
||||
if (!window.confirm(`Delete alert “${target.Name}”? This removes it from the config.`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete alert',
|
||||
title: `Delete alert “${target.Name}”?`,
|
||||
body: 'This removes it from the config.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = alerts.filter((_, i) => i !== idx)
|
||||
void save({ ...config, Alerts: next }, `Deleted ${target.Name}`)
|
||||
},
|
||||
[config, alerts, save],
|
||||
[config, alerts, save, confirm],
|
||||
)
|
||||
|
||||
const setAlertVia = useCallback(
|
||||
|
||||
@@ -138,57 +138,9 @@
|
||||
|
||||
/* inline rename: a quiet pencil affordance beside the name, and the mono input
|
||||
it swaps to — in the same sink/groove tone as the domain editors. */
|
||||
.dev-rename {
|
||||
flex: none;
|
||||
display: inline-flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
width: 22px;
|
||||
height: 22px;
|
||||
padding: 0;
|
||||
border: 1px solid transparent;
|
||||
border-radius: 5px;
|
||||
background: none;
|
||||
color: var(--faint);
|
||||
font-size: 12px;
|
||||
line-height: 1;
|
||||
cursor: pointer;
|
||||
transition: color 0.15s, background 0.15s, border-color 0.15s;
|
||||
}
|
||||
.dev-rename:hover:not(:disabled) {
|
||||
color: var(--accent);
|
||||
background: color-mix(in srgb, var(--accent) 12%, transparent);
|
||||
}
|
||||
.dev-rename:focus-visible {
|
||||
color: var(--accent);
|
||||
border-color: var(--accent);
|
||||
outline: 2px solid var(--accent);
|
||||
outline-offset: 1px;
|
||||
}
|
||||
.dev-rename:disabled {
|
||||
opacity: 0.5;
|
||||
cursor: default;
|
||||
}
|
||||
.dev-name-input {
|
||||
min-width: 0;
|
||||
max-width: 24ch;
|
||||
padding: 4px 8px;
|
||||
border: 1px solid var(--accent);
|
||||
border-radius: 6px;
|
||||
background: var(--sink);
|
||||
color: var(--ink);
|
||||
font-size: 13px;
|
||||
font-weight: 600;
|
||||
letter-spacing: 0.01em;
|
||||
box-shadow: 0 1px 2px var(--shadow) inset;
|
||||
}
|
||||
.dev-name-input:focus-visible {
|
||||
outline: 2px solid var(--accent);
|
||||
outline-offset: 1px;
|
||||
}
|
||||
.dev-name-input:disabled {
|
||||
opacity: 0.55;
|
||||
}
|
||||
/* The pencil button and the name input now live in App.css as .inline-rename /
|
||||
.inline-rename-input — Nodes grew the same affordance and the two pages must
|
||||
not drift. */
|
||||
.dev-id-l2 {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
import './Devices.css'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import { Button, Led, Module, Toggle } from '../components'
|
||||
import { Button, Led, Module, Toggle, useConfirm } from '../components'
|
||||
import type { LedVariant } from '../components'
|
||||
import { apply as apiApply, getConfig, getDevices, putConfig, ApiError } from '../api'
|
||||
import type { Device, DiscoveredDevice, Model } from '../api'
|
||||
@@ -81,6 +81,7 @@ function networkLabel(row: DeviceRow): string {
|
||||
// ---- page ------------------------------------------------------------------
|
||||
|
||||
export default function Devices() {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<Model | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
|
||||
@@ -251,17 +252,22 @@ export default function Devices() {
|
||||
const nameOf = (row: DeviceRow) => row.cfg?.Name || row.hostname || row.ip || 'device'
|
||||
|
||||
const removeControl = useCallback(
|
||||
(row: DeviceRow) => {
|
||||
async (row: DeviceRow) => {
|
||||
if (!config) return
|
||||
const devs = asArray(config.Devices)
|
||||
const idx = matchDevice(devs, row.mac, row.ip)
|
||||
if (idx < 0) return
|
||||
const nm = devs[idx].Name || nameOf(row)
|
||||
if (!window.confirm(`Stop managing “${nm}”? Its per-device rules are removed; it falls back to network defaults.`))
|
||||
return
|
||||
const ok = await confirm({
|
||||
label: 'Stop managing device',
|
||||
title: `Stop managing “${nm}”?`,
|
||||
body: 'Its per-device rules are removed; it falls back to network defaults.',
|
||||
confirmLabel: 'Stop managing',
|
||||
})
|
||||
if (!ok) return
|
||||
void save({ ...config, Devices: devs.filter((_, i) => i !== idx) }, `Removed control for ${nm}`)
|
||||
},
|
||||
[config, save],
|
||||
[config, save, confirm],
|
||||
)
|
||||
|
||||
const loading = config === null && loadError === null && devices === null && devError === null
|
||||
@@ -471,7 +477,7 @@ function DeviceCard({
|
||||
{renaming ? (
|
||||
<input
|
||||
ref={nameInput}
|
||||
className="dev-name-input mono"
|
||||
className="inline-rename-input mono"
|
||||
type="text"
|
||||
spellCheck={false}
|
||||
autoComplete="off"
|
||||
@@ -497,7 +503,7 @@ function DeviceCard({
|
||||
</span>
|
||||
<button
|
||||
type="button"
|
||||
className="dev-rename"
|
||||
className="inline-rename"
|
||||
onClick={beginRename}
|
||||
disabled={busy}
|
||||
aria-label={`Rename ${name}`}
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
import './Networks.css'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import { Button, Led, Select, Toggle } from '../components'
|
||||
import { Button, Led, Select, Toggle, useConfirm } from '../components'
|
||||
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
|
||||
import type { Inbound, Interface, Model, Status } from '../api'
|
||||
import { isLanNetwork, isWanNetwork, useInterfaces } from '../srcOptions'
|
||||
@@ -242,6 +242,7 @@ function computeWarnings(inbounds: Inbound[], ifaces: Interface[]): Warning[] {
|
||||
// ---- page ------------------------------------------------------------------
|
||||
|
||||
export default function Networks({ status }: { status?: Status | null }) {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<Model | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
const ifaces = useInterfaces()
|
||||
@@ -392,23 +393,21 @@ export default function Networks({ status }: { status?: Status | null }) {
|
||||
)
|
||||
|
||||
const removeInbound = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = inbounds[idx]
|
||||
if (
|
||||
!window.confirm(
|
||||
`Delete inbound “${target.Name}”?${
|
||||
intercepts(target)
|
||||
? ` ${target.Network || 'Its network'} stops going through the tunnel.`
|
||||
: ''
|
||||
}`,
|
||||
)
|
||||
)
|
||||
return
|
||||
const ok = await confirm({
|
||||
label: 'Delete inbound',
|
||||
title: `Delete inbound “${target.Name}”?`,
|
||||
body: intercepts(target)
|
||||
? `${target.Network || 'Its network'} stops going through the tunnel.`
|
||||
: undefined,
|
||||
})
|
||||
if (!ok) return
|
||||
const next = inbounds.filter((_, i) => i !== idx)
|
||||
void save({ ...config, Inbounds: next }, `Deleted ${target.Name}`)
|
||||
},
|
||||
[config, inbounds, save],
|
||||
[config, inbounds, save, confirm],
|
||||
)
|
||||
|
||||
return (
|
||||
|
||||
@@ -727,3 +727,53 @@ select.fp-input {
|
||||
width: 9rem;
|
||||
}
|
||||
}
|
||||
|
||||
/* ---- inline node rename ----
|
||||
The pencil / input pair itself is shared (.inline-rename[-input] in App.css);
|
||||
only the row-local sizing and the refusal message live here. A node name is
|
||||
longer than a device name (it carries a protocol and a host), so the field is
|
||||
given more room than the shared 24ch default. */
|
||||
.node-name-input {
|
||||
max-width: 32ch;
|
||||
font-family: var(--font-mono);
|
||||
font-size: 12.5px;
|
||||
}
|
||||
/* Why a rename was refused, pinned under the row it was typed in. Semantic crit:
|
||||
the name did not change, and that must not be mistaken for a saved edit. */
|
||||
.row-err {
|
||||
margin: 2px 0 0;
|
||||
font-size: 11.5px;
|
||||
line-height: 1.45;
|
||||
color: var(--crit);
|
||||
max-width: 68ch;
|
||||
}
|
||||
/* Stated once per subscription bucket: the same rule the locked control in every
|
||||
row carries, so the absent rename is explained before it is looked for. */
|
||||
.group-note {
|
||||
margin: 0;
|
||||
padding: 8px 12px;
|
||||
border: 1px solid var(--groove);
|
||||
border-top: 0;
|
||||
background: color-mix(in srgb, var(--sink) 25%, transparent);
|
||||
font-size: 11.5px;
|
||||
line-height: 1.5;
|
||||
color: var(--faint);
|
||||
}
|
||||
/* The optional name sits beside the link input on a wide row and drops onto its
|
||||
own line when the row can no longer hold both. */
|
||||
.add-name {
|
||||
flex: 0 1 22ch;
|
||||
min-width: 12ch;
|
||||
}
|
||||
.add-row--conf .add-name {
|
||||
flex: none;
|
||||
align-self: stretch;
|
||||
}
|
||||
@media (max-width: 640px) {
|
||||
.add-row {
|
||||
flex-wrap: wrap;
|
||||
}
|
||||
.add-name {
|
||||
flex: 1 1 100%;
|
||||
}
|
||||
}
|
||||
|
||||
+486
-15
@@ -2,7 +2,7 @@ import './Nodes.css'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import type { ReactNode } from 'react'
|
||||
import type { LedVariant } from '../components'
|
||||
import { Button, Led, Toggle } from '../components'
|
||||
import { Button, Led, Toggle, useConfirm } from '../components'
|
||||
import {
|
||||
apply as apiApply,
|
||||
getConfig,
|
||||
@@ -158,6 +158,219 @@ function uniqueName(base: string, taken: Set<string>): string {
|
||||
return `${seed}-${i}`
|
||||
}
|
||||
|
||||
// ---- node names are identity, not a caption --------------------------------
|
||||
//
|
||||
// A node's Name IS its sing-box outbound tag and the only thing every reference
|
||||
// to it spells: a rule target `node:<name>`, a chain hop, a manual group's member
|
||||
// list, a resolver detour, an alert delivery, a subscription fetch detour. Rename
|
||||
// the node alone and every one of those points at nothing — and an unresolved
|
||||
// target does NOT fall back to the default route, the daemon BLOCKS that traffic.
|
||||
// So the rename either carries every reference with it, or it is refused.
|
||||
|
||||
/** Reserved outbound tags. A node called this is skipped by the generator entirely. */
|
||||
const RESERVED_TAGS = ['direct', 'block']
|
||||
|
||||
/**
|
||||
* Prefixes that `model.SplitTarget` reads as a KIND, not as part of a name. A
|
||||
* name starting with one of them makes every bare reference to it ambiguous with
|
||||
* a real `kind:name` reference, so it is refused rather than half-supported.
|
||||
*/
|
||||
const KIND_PREFIXES = ['node', 'group', 'egress', 'chain', 'direct', 'block']
|
||||
|
||||
/** Names are rendered into a line-oriented `uci export`; control chars are stripped there. */
|
||||
function hasControlChar(s: string): boolean {
|
||||
for (let i = 0; i < s.length; i++) {
|
||||
const c = s.charCodeAt(i)
|
||||
if (c < 0x20 || c === 0x7f) return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
/**
|
||||
* Why `name` cannot be a node name here, or null if it can.
|
||||
*
|
||||
* Every rule mirrors something the daemon actually does with the name, not a
|
||||
* house style: reserved tags make generate skip the node; a duplicate makes two
|
||||
* outbounds share a tag and the manager silently keeps the last one; a group of
|
||||
* the same name is dropped by buildGroups ("rename the group"); an
|
||||
* `egress-<name>` collision takes over a real egress outbound; and a control
|
||||
* character is rewritten to a space by sanitizeUCIValue on write, so the saved
|
||||
* name would not be the one you typed.
|
||||
*/
|
||||
function nodeNameError(
|
||||
raw: string,
|
||||
m: Model | null,
|
||||
self: string | null,
|
||||
): string | null {
|
||||
const name = raw.trim()
|
||||
if (!name) return 'A node needs a name.'
|
||||
if (hasControlChar(name))
|
||||
return 'Names can’t contain line breaks or control characters — they’re stripped when the config is written.'
|
||||
if (RESERVED_TAGS.some((t) => t.toLowerCase() === name.toLowerCase()))
|
||||
return `“${name}” is a reserved target name — a node called that is skipped by the engine. Pick another.`
|
||||
const head = name.includes(':') ? name.slice(0, name.indexOf(':')).toLowerCase() : ''
|
||||
if (head && KIND_PREFIXES.includes(head))
|
||||
return `A name starting with “${head}:” reads as a ${head} reference everywhere it’s used. Pick another.`
|
||||
if (!m) return null
|
||||
|
||||
const clash = asArray(m.Nodes).find((n) => n.Name === name && n.Name !== self)
|
||||
if (clash)
|
||||
return clash.FromSub
|
||||
? `“${name}” is already a node from subscription “${clash.FromSub}”. Two nodes with one name share a single outbound — pick another.`
|
||||
: `“${name}” is already another node. Pick another.`
|
||||
if (asArray(m.Groups).some((g) => g.Name === name))
|
||||
return `A group is already named “${name}”. The engine drops the group when a node takes its name — pick another.`
|
||||
const egressClash = asArray(m.Egresses).find((e) => `egress-${e.Name}` === name)
|
||||
if (egressClash)
|
||||
return `“${name}” is the outbound tag of egress “${egressClash.Name}”. Pick another.`
|
||||
return null
|
||||
}
|
||||
|
||||
/** One place a node name is written, as a short label for the rename summary. */
|
||||
interface NodeRefSite {
|
||||
/** Which section — drives the "N rules, M groups" count. */
|
||||
kind: 'rule' | 'group' | 'chain' | 'resolver' | 'alert' | 'subscription' | 'egress'
|
||||
label: string
|
||||
}
|
||||
|
||||
/** A target/detour string naming this node in its prefixed form (`node:<name>`). */
|
||||
const isNodeRef = (v: string | undefined | null, name: string): boolean =>
|
||||
(v ?? '') === `node:${name}`
|
||||
|
||||
/** …or in the bare form the engine also resolves (a group member, a bare hop/target). */
|
||||
const isBareRef = (v: string | undefined | null, name: string): boolean => (v ?? '') === name
|
||||
|
||||
/**
|
||||
* Every place `name` is written outside the node itself. Both spellings count:
|
||||
* `resolveTarget` falls through to a bare node lookup, and a manual group's
|
||||
* member list is bare by contract.
|
||||
*/
|
||||
function findNodeReferences(m: Model, name: string): NodeRefSite[] {
|
||||
const out: NodeRefSite[] = []
|
||||
for (const r of asArray(m.Rules)) {
|
||||
if (isNodeRef(r.Target, name) || isBareRef(r.Target, name))
|
||||
out.push({ kind: 'rule', label: `rule “${r.Name}” target` })
|
||||
}
|
||||
for (const g of asArray(m.Groups)) {
|
||||
if (asArray(g.Nodes).some((n) => n === name))
|
||||
out.push({ kind: 'group', label: `group “${g.Name}” member` })
|
||||
}
|
||||
for (const c of asArray(m.Chains)) {
|
||||
if (asArray(c.Hops).some((h) => isNodeRef(h, name) || isBareRef(h, name)))
|
||||
out.push({ kind: 'chain', label: `chain “${c.Name}” hop` })
|
||||
}
|
||||
for (const r of asArray(m.Resolvers)) {
|
||||
if (isNodeRef(r.Detour, name)) out.push({ kind: 'resolver', label: `resolver “${r.Name}” DNS path` })
|
||||
}
|
||||
for (const a of asArray(m.Alerts)) {
|
||||
if (isNodeRef(a.Via, name)) out.push({ kind: 'alert', label: `alert “${a.Name}” delivery` })
|
||||
}
|
||||
for (const s of asArray(m.Subscriptions)) {
|
||||
if (isNodeRef(s.FetchDetour, name))
|
||||
out.push({ kind: 'subscription', label: `subscription “${s.Name}” fetch` })
|
||||
}
|
||||
for (const e of asArray(m.Egresses)) {
|
||||
if (isNodeRef(e.Target, name)) out.push({ kind: 'egress', label: `egress “${e.Name}” target` })
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
/**
|
||||
* What makes a rename impossible to carry rather than merely wide.
|
||||
*
|
||||
* A BARE reference is just a name; the engine resolves it node-first, then group.
|
||||
* If something else already answers to the old name, we cannot tell which object
|
||||
* a bare reference meant, and rewriting it would move a reference the operator
|
||||
* never pointed at this node. That is a half-done cascade, so the rename is
|
||||
* refused instead — with the collision named, so it can be fixed.
|
||||
*/
|
||||
function bareAmbiguity(m: Model, name: string): string | null {
|
||||
const group = asArray(m.Groups).find((g) => g.Name === name)
|
||||
if (!group) return null
|
||||
const bare = [
|
||||
...asArray(m.Rules)
|
||||
.filter((r) => isBareRef(r.Target, name))
|
||||
.map((r) => `rule “${r.Name}”`),
|
||||
...asArray(m.Chains)
|
||||
.filter((c) => asArray(c.Hops).some((h) => isBareRef(h, name)))
|
||||
.map((c) => `chain “${c.Name}”`),
|
||||
]
|
||||
if (bare.length === 0) return null
|
||||
return `A group is also named “${name}”, and ${bare.join(', ')} point${bare.length === 1 ? 's' : ''} at that bare name — there is no way to tell which of the two is meant. Rename the group first, then this node.`
|
||||
}
|
||||
|
||||
/**
|
||||
* Rewrite every reference from `from` to `to`. Returns a NEW Model with only the
|
||||
* touched sections replaced; the Nodes section is the caller's business.
|
||||
*
|
||||
* Bare references are rewritten too — that is the whole point for a manual
|
||||
* group's member list — which is safe only because `bareAmbiguity` has already
|
||||
* refused the one case where a bare name could mean something else.
|
||||
*/
|
||||
function renameNodeReferences(m: Model, from: string, to: string): Model {
|
||||
if (from === to) return m
|
||||
/** Prefixed-only sites (a detour is never spelled bare). */
|
||||
const pfx = (v: string | undefined) => (isNodeRef(v, from) ? `node:${to}` : v)
|
||||
/** Sites that accept either spelling — each is rewritten in the spelling it already uses. */
|
||||
const either = (v: string | undefined) => {
|
||||
if (isNodeRef(v, from)) return `node:${to}`
|
||||
if (isBareRef(v, from)) return to
|
||||
return v
|
||||
}
|
||||
|
||||
const next: Model = { ...m }
|
||||
if (m.Rules) next.Rules = m.Rules.map((r) => ({ ...r, Target: either(r.Target) }))
|
||||
if (m.Groups)
|
||||
next.Groups = m.Groups.map((g) => ({
|
||||
...g,
|
||||
Nodes: g.Nodes ? g.Nodes.map((n) => (n === from ? to : n)) : g.Nodes,
|
||||
}))
|
||||
if (m.Chains)
|
||||
next.Chains = m.Chains.map((c) => ({
|
||||
...c,
|
||||
Hops: c.Hops ? c.Hops.map((h) => either(h) ?? h) : c.Hops,
|
||||
}))
|
||||
if (m.Resolvers) next.Resolvers = m.Resolvers.map((r) => ({ ...r, Detour: pfx(r.Detour) }))
|
||||
if (m.Alerts) next.Alerts = m.Alerts.map((a) => ({ ...a, Via: pfx(a.Via) }))
|
||||
if (m.Subscriptions)
|
||||
next.Subscriptions = m.Subscriptions.map((s) => ({ ...s, FetchDetour: pfx(s.FetchDetour) }))
|
||||
if (m.Egresses) next.Egresses = m.Egresses.map((e) => ({ ...e, Target: pfx(e.Target) }))
|
||||
return next
|
||||
}
|
||||
|
||||
/** "3 rules, 1 group and 2 chains" — what the rename is about to rewrite. */
|
||||
function refSummary(refs: NodeRefSite[]): string {
|
||||
const plural: Record<NodeRefSite['kind'], [string, string]> = {
|
||||
rule: ['rule', 'rules'],
|
||||
group: ['group', 'groups'],
|
||||
chain: ['chain', 'chains'],
|
||||
resolver: ['resolver', 'resolvers'],
|
||||
alert: ['alert', 'alerts'],
|
||||
subscription: ['subscription', 'subscriptions'],
|
||||
egress: ['egress', 'egresses'],
|
||||
}
|
||||
const order: NodeRefSite['kind'][] = [
|
||||
'rule', 'group', 'chain', 'resolver', 'alert', 'subscription', 'egress',
|
||||
]
|
||||
const parts = order
|
||||
.map((k) => [k, refs.filter((r) => r.kind === k).length] as const)
|
||||
.filter(([, n]) => n > 0)
|
||||
.map(([k, n]) => `${n} ${plural[k][n === 1 ? 0 : 1]}`)
|
||||
if (parts.length === 1) return parts[0]
|
||||
return `${parts.slice(0, -1).join(', ')} and ${parts[parts.length - 1]}`
|
||||
}
|
||||
|
||||
/**
|
||||
* The name a rename just committed to, waiting for its row to come back.
|
||||
*
|
||||
* A row is keyed by the node's NAME, so committing a rename unmounts the row and
|
||||
* mounts a different one — carrying the focused element away with it. This baton
|
||||
* survives that remount: the row that reappears under the new name claims it and
|
||||
* puts the keyboard back on its own rename button, instead of dropping the user
|
||||
* on <body> halfway down a list of 300 nodes.
|
||||
*/
|
||||
let pendingRenameFocus: string | null = null
|
||||
|
||||
// A subscription with more than this many nodes starts collapsed so the list
|
||||
// doesn't become one endless scroll; an active search overrides it.
|
||||
const LARGE_GROUP = 20
|
||||
@@ -289,6 +502,7 @@ function DetourSelect({
|
||||
// ---- page ------------------------------------------------------------------
|
||||
|
||||
export default function Nodes() {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<Model | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
|
||||
@@ -385,6 +599,9 @@ export default function Nodes() {
|
||||
const [nodeInput, setNodeInput] = useState('')
|
||||
const [nodeErr, setNodeErr] = useState<string | null>(null)
|
||||
const [addMode, setAddMode] = useState<'link' | 'conf'>('link')
|
||||
// Optional. Empty keeps the old behaviour (a name derived from the server
|
||||
// address), so "paste a link, press Add" stays a two-step path.
|
||||
const [nodeName, setNodeName] = useState('')
|
||||
const [importing, setImporting] = useState(false)
|
||||
|
||||
// ---- node search + collapsible grouping -----------------------------------
|
||||
@@ -445,8 +662,12 @@ export default function Nodes() {
|
||||
try {
|
||||
const { uri, name } = await importWg(conf)
|
||||
const taken = new Set(nodes.map((n) => n.Name))
|
||||
// A typed name is used AS TYPED — uniqueName would silently turn a
|
||||
// collision into "name-2", which is the confusion this field exists to
|
||||
// end. It is validated instead, and a clash is refused out loud above.
|
||||
const wanted = nodeName.trim()
|
||||
const node: NodeCfg = {
|
||||
Name: uniqueName(name || 'wireguard', taken),
|
||||
Name: wanted || uniqueName(name || 'wireguard', taken),
|
||||
Enabled: true,
|
||||
URI: uri,
|
||||
FromSub: '',
|
||||
@@ -455,6 +676,7 @@ export default function Nodes() {
|
||||
const ok = await save({ ...config, Nodes: [...nodes, node] }, `Added ${node.Name}`)
|
||||
if (ok) {
|
||||
setNodeInput('')
|
||||
setNodeName('')
|
||||
setAddMode('link')
|
||||
}
|
||||
} catch (e) {
|
||||
@@ -463,11 +685,20 @@ export default function Nodes() {
|
||||
setImporting(false)
|
||||
}
|
||||
},
|
||||
[config, nodes, save, flash],
|
||||
[config, nodes, nodeName, save, flash],
|
||||
)
|
||||
|
||||
const addNode = useCallback(async () => {
|
||||
if (!config) return
|
||||
// The name is checked BEFORE the import round-trip, so a bad name costs
|
||||
// nothing and the message lands in the form next to the field.
|
||||
if (nodeName.trim()) {
|
||||
const bad = nodeNameError(nodeName, config, null)
|
||||
if (bad) {
|
||||
setNodeErr(bad)
|
||||
return
|
||||
}
|
||||
}
|
||||
// Auto-detect a pasted config, whichever input it landed in.
|
||||
if (nodeInput.includes(WG_MARKER)) {
|
||||
await addWgConf(nodeInput)
|
||||
@@ -486,10 +717,20 @@ export default function Nodes() {
|
||||
const parsed = parseShareLink(uri)
|
||||
const taken = new Set(nodes.map((n) => n.Name))
|
||||
const base = parsed.suggested || `${parsed.proto.toLowerCase()}-${parsed.host}`.replace(/[^\w.:-]+/g, '-')
|
||||
const node: NodeCfg = { Name: uniqueName(base, taken), Enabled: true, URI: uri, FromSub: '', Egress: '' }
|
||||
const wanted = nodeName.trim()
|
||||
const node: NodeCfg = {
|
||||
Name: wanted || uniqueName(base, taken),
|
||||
Enabled: true,
|
||||
URI: uri,
|
||||
FromSub: '',
|
||||
Egress: '',
|
||||
}
|
||||
const ok = await save({ ...config, Nodes: [...nodes, node] }, `Added ${node.Name}`)
|
||||
if (ok) setNodeInput('')
|
||||
}, [config, nodeInput, nodes, save, addMode, addWgConf])
|
||||
if (ok) {
|
||||
setNodeInput('')
|
||||
setNodeName('')
|
||||
}
|
||||
}, [config, nodeInput, nodeName, nodes, save, addMode, addWgConf])
|
||||
|
||||
const toggleNode = useCallback(
|
||||
(idx: number, on: boolean) => {
|
||||
@@ -501,14 +742,88 @@ export default function Nodes() {
|
||||
)
|
||||
|
||||
const removeNode = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = nodes[idx]
|
||||
if (!window.confirm(`Delete node “${target.Name}”? This removes it from the config.`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete node',
|
||||
title: `Delete node “${target.Name}”?`,
|
||||
body: 'This removes it from the config.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = nodes.filter((_, i) => i !== idx)
|
||||
void save({ ...config, Nodes: next }, `Deleted ${target.Name}`)
|
||||
},
|
||||
[config, nodes, save],
|
||||
[config, nodes, save, confirm],
|
||||
)
|
||||
|
||||
/**
|
||||
* Rename a manual node, carrying every reference with it.
|
||||
*
|
||||
* The name is this node's identity: its outbound tag, and the exact string a
|
||||
* rule target, a chain hop, a manual group's member list, a resolver detour, an
|
||||
* alert delivery and a subscription fetch detour all spell. So the rename is one
|
||||
* atomic save of the Nodes section AND every referencing section, or it does not
|
||||
* happen at all:
|
||||
*
|
||||
* - an invalid or colliding name is refused with the reason (`nodeNameError`);
|
||||
* - a name a GROUP also answers to, with bare references pointing at it, is
|
||||
* refused too — there is no way to know which object those meant, and
|
||||
* guessing would move a reference the operator never pointed here;
|
||||
* - anything else is shown exactly what it will rewrite, and only then saved.
|
||||
*
|
||||
* Errors surface through `onError` so they land in the row that was edited.
|
||||
*/
|
||||
const renameNode = useCallback(
|
||||
async (idx: number, raw: string, onError: (msg: string) => void): Promise<boolean> => {
|
||||
if (!config) return false
|
||||
const target = nodes[idx]
|
||||
const from = target.Name
|
||||
const to = raw.trim()
|
||||
if (to === from) return true
|
||||
// Subscription names come back from the feed on the next update; renaming
|
||||
// one would be undone without warning, so this path is manual-only.
|
||||
if (target.FromSub) {
|
||||
onError(`“${from}” is named by subscription “${target.FromSub}” — the feed rewrites it on the next update.`)
|
||||
return false
|
||||
}
|
||||
const bad = nodeNameError(to, config, from)
|
||||
if (bad) {
|
||||
onError(bad)
|
||||
return false
|
||||
}
|
||||
const blocked = bareAmbiguity(config, from)
|
||||
if (blocked) {
|
||||
onError(blocked)
|
||||
return false
|
||||
}
|
||||
|
||||
const refs = findNodeReferences(config, from)
|
||||
if (refs.length > 0) {
|
||||
const shown = refs.slice(0, 4).map((r) => r.label)
|
||||
const more = refs.length - shown.length
|
||||
const ok = await confirm({
|
||||
tone: 'neutral',
|
||||
label: 'Rename node',
|
||||
title: `Rename “${from}” to “${to}”?`,
|
||||
body: `This also updates ${refSummary(refs)} that point at it — ${shown.join(', ')}${more > 0 ? `, and ${more} more` : ''}. They are saved together, so nothing is left pointing at the old name.`,
|
||||
confirmLabel: 'Rename',
|
||||
})
|
||||
if (!ok) return false
|
||||
}
|
||||
|
||||
// One PUT: the node and every reference move in the same write, so no
|
||||
// intermediate state exists where a reference dangles.
|
||||
const carried = renameNodeReferences(config, from, to)
|
||||
const next = asArray(carried.Nodes).map((n, i) => (i === idx ? { ...n, Name: to } : n))
|
||||
return save(
|
||||
{ ...carried, Nodes: next },
|
||||
refs.length > 0
|
||||
? `Renamed to ${to} — updated ${refs.length} reference${refs.length === 1 ? '' : 's'}`
|
||||
: `Renamed to ${to}`,
|
||||
)
|
||||
},
|
||||
[config, nodes, save, confirm],
|
||||
)
|
||||
|
||||
// Pin (or clear) one node's dial egress. Same optimistic save→apply path as
|
||||
@@ -563,16 +878,20 @@ export default function Nodes() {
|
||||
)
|
||||
|
||||
const removeSub = useCallback(
|
||||
(idx: number) => {
|
||||
async (idx: number) => {
|
||||
if (!config) return
|
||||
const target = subs[idx]
|
||||
const hasCache = nodes.some((n) => n.FromSub === target.Name)
|
||||
const extra = hasCache ? ' Its cached nodes stay until you next apply.' : ''
|
||||
if (!window.confirm(`Delete subscription “${target.Name}”?${extra}`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete subscription',
|
||||
title: `Delete subscription “${target.Name}”?`,
|
||||
body: hasCache ? 'Its cached nodes stay until you next apply.' : undefined,
|
||||
})
|
||||
if (!ok) return
|
||||
const next = subs.filter((_, i) => i !== idx)
|
||||
void save({ ...config, Subscriptions: next }, `Deleted ${target.Name}`)
|
||||
},
|
||||
[config, subs, nodes, save],
|
||||
[config, subs, nodes, save, confirm],
|
||||
)
|
||||
|
||||
// Commit an options edit for one subscription. The editor hands back a fully
|
||||
@@ -736,6 +1055,20 @@ export default function Nodes() {
|
||||
disabled={busy || importing || !config}
|
||||
/>
|
||||
)}
|
||||
<input
|
||||
className="fp-input add-name"
|
||||
type="text"
|
||||
spellCheck={false}
|
||||
autoComplete="off"
|
||||
placeholder="Name (optional)"
|
||||
aria-label="Node name — optional"
|
||||
value={nodeName}
|
||||
onChange={(e) => {
|
||||
setNodeName(e.target.value)
|
||||
if (nodeErr) setNodeErr(null)
|
||||
}}
|
||||
disabled={busy || importing || !config}
|
||||
/>
|
||||
<Button type="submit" variant="primary" disabled={busy || importing || !config}>
|
||||
{importing ? 'Importing…' : saving ? 'Saving…' : addMode === 'conf' ? 'Import' : 'Add node'}
|
||||
</Button>
|
||||
@@ -743,7 +1076,8 @@ export default function Nodes() {
|
||||
<p className="add-hint">
|
||||
{addMode === 'conf'
|
||||
? 'Paste a wg-quick / AmneziaWG .conf — it starts with [Interface].'
|
||||
: 'vless://, ss://, trojan://, hysteria2://… A pasted [Interface] config is imported automatically.'}
|
||||
: 'vless://, ss://, trojan://, hysteria2://… A pasted [Interface] config is imported automatically.'}{' '}
|
||||
Leave the name empty and it’s taken from the server address; you can rename it later.
|
||||
</p>
|
||||
</div>
|
||||
{nodeErr && (
|
||||
@@ -800,6 +1134,7 @@ export default function Nodes() {
|
||||
onToggle={() => toggleGroup(g)}
|
||||
onToggleNode={toggleNode}
|
||||
onRemoveNode={removeNode}
|
||||
onRenameNode={renameNode}
|
||||
onSetEgress={setNodeEgress}
|
||||
/>
|
||||
))}
|
||||
@@ -911,6 +1246,7 @@ function NodeGroup({
|
||||
onToggle,
|
||||
onToggleNode,
|
||||
onRemoveNode,
|
||||
onRenameNode,
|
||||
onSetEgress,
|
||||
}: {
|
||||
group: NodeGroupData
|
||||
@@ -920,6 +1256,7 @@ function NodeGroup({
|
||||
onToggle: () => void
|
||||
onToggleNode: (idx: number, on: boolean) => void
|
||||
onRemoveNode: (idx: number) => void
|
||||
onRenameNode: (idx: number, name: string, onError: (msg: string) => void) => Promise<boolean>
|
||||
onSetEgress: (idx: number, egress: string) => Promise<boolean>
|
||||
}) {
|
||||
const panelId = `node-group-${group.key || 'manual'}`
|
||||
@@ -940,6 +1277,12 @@ function NodeGroup({
|
||||
<span className="group-count mono">{count}</span>
|
||||
</button>
|
||||
</h3>
|
||||
{open && group.key !== '' && (
|
||||
<p className="group-note">
|
||||
Names come from the subscription feed and are rewritten on every update, so nodes in this
|
||||
list can’t be renamed here.
|
||||
</p>
|
||||
)}
|
||||
{open && (
|
||||
<ul id={panelId} className="rows-list group-rows">
|
||||
{group.items.map(({ node, idx }) => (
|
||||
@@ -950,6 +1293,7 @@ function NodeGroup({
|
||||
egressNames={egressNames}
|
||||
onToggle={(on) => onToggleNode(idx, on)}
|
||||
onDelete={() => onRemoveNode(idx)}
|
||||
onRename={(name, onError) => onRenameNode(idx, name, onError)}
|
||||
onSetEgress={(egress) => onSetEgress(idx, egress)}
|
||||
/>
|
||||
))}
|
||||
@@ -965,6 +1309,7 @@ function NodeRow({
|
||||
egressNames,
|
||||
onToggle,
|
||||
onDelete,
|
||||
onRename,
|
||||
onSetEgress,
|
||||
}: {
|
||||
node: NodeCfg
|
||||
@@ -972,6 +1317,7 @@ function NodeRow({
|
||||
egressNames: string[]
|
||||
onToggle: (on: boolean) => void
|
||||
onDelete: () => void
|
||||
onRename: (name: string, onError: (msg: string) => void) => Promise<boolean>
|
||||
onSetEgress: (egress: string) => Promise<boolean>
|
||||
}) {
|
||||
const { proto, host, hasCreds } = useMemo(() => parseShareLink(node.URI), [node.URI])
|
||||
@@ -982,6 +1328,68 @@ function NodeRow({
|
||||
const [open, setOpen] = useState(false)
|
||||
const panelId = `node-egress-${node.FromSub || 'manual'}-${node.Name}`
|
||||
|
||||
// ---- inline rename (same interaction as a device row) ---------------------
|
||||
// Enter commits, Esc cancels, blur commits; a ref-guard keeps Esc-then-blur
|
||||
// from committing twice. Unlike a device, the commit can be REFUSED (a name
|
||||
// collision, or references that can't be carried), so the input stays open
|
||||
// with the reason under it instead of closing on a change that never happened.
|
||||
const [renaming, setRenaming] = useState(false)
|
||||
const [draft, setDraft] = useState(node.Name)
|
||||
const [renameErr, setRenameErr] = useState<string | null>(null)
|
||||
const nameInput = useRef<HTMLInputElement>(null)
|
||||
const renameBtn = useRef<HTMLButtonElement>(null)
|
||||
const finished = useRef(false)
|
||||
|
||||
const beginRename = () => {
|
||||
setDraft(node.Name)
|
||||
setRenameErr(null)
|
||||
finished.current = false
|
||||
setRenaming(true)
|
||||
}
|
||||
const finishRename = async (commit: boolean) => {
|
||||
if (finished.current) return
|
||||
finished.current = true
|
||||
const nm = draft.trim()
|
||||
if (!commit || !nm || nm === node.Name) {
|
||||
setRenaming(false)
|
||||
setRenameErr(null)
|
||||
return
|
||||
}
|
||||
// Armed BEFORE the save: the renamed row remounts the moment the config
|
||||
// state lands, which is before this await resolves. Arming afterwards would
|
||||
// always miss it.
|
||||
pendingRenameFocus = nm
|
||||
const ok = await onRename(nm, (msg) => setRenameErr(msg))
|
||||
if (ok) {
|
||||
setRenaming(false)
|
||||
setRenameErr(null)
|
||||
} else {
|
||||
if (pendingRenameFocus === nm) pendingRenameFocus = null
|
||||
// Refused — hold the field open on the rejected text so it can be fixed.
|
||||
finished.current = false
|
||||
nameInput.current?.focus()
|
||||
}
|
||||
}
|
||||
|
||||
useEffect(() => {
|
||||
if (renaming) {
|
||||
nameInput.current?.focus()
|
||||
nameInput.current?.select()
|
||||
}
|
||||
}, [renaming])
|
||||
|
||||
// Claim the baton if this row is the one the rename produced. The row remounts
|
||||
// while the PUT is still in flight, so on that first pass the button is still
|
||||
// disabled and focus() would be a silent no-op — the baton is held until the
|
||||
// save settles and this effect re-runs with a focusable button.
|
||||
useEffect(() => {
|
||||
if (pendingRenameFocus !== node.Name) return
|
||||
const btn = renameBtn.current
|
||||
if (!btn || btn.disabled) return
|
||||
pendingRenameFocus = null
|
||||
btn.focus()
|
||||
}, [node.Name, busy])
|
||||
|
||||
return (
|
||||
<li className={`row-item node-row${open ? ' node-row--open' : ''}`}>
|
||||
<div className="row-head">
|
||||
@@ -993,10 +1401,73 @@ function NodeRow({
|
||||
/>
|
||||
<div className="row-main">
|
||||
<div className="row-line1">
|
||||
<span className="row-name">{node.Name}</span>
|
||||
{renaming ? (
|
||||
<input
|
||||
ref={nameInput}
|
||||
className="inline-rename-input node-name-input mono"
|
||||
type="text"
|
||||
spellCheck={false}
|
||||
autoComplete="off"
|
||||
value={draft}
|
||||
aria-label={`Rename node ${node.Name}`}
|
||||
aria-invalid={renameErr ? true : undefined}
|
||||
onChange={(e) => {
|
||||
setDraft(e.target.value)
|
||||
if (renameErr) setRenameErr(null)
|
||||
}}
|
||||
onBlur={() => void finishRename(true)}
|
||||
onKeyDown={(e) => {
|
||||
if (e.key === 'Enter') {
|
||||
e.preventDefault()
|
||||
void finishRename(true)
|
||||
} else if (e.key === 'Escape') {
|
||||
e.preventDefault()
|
||||
void finishRename(false)
|
||||
}
|
||||
}}
|
||||
disabled={busy}
|
||||
/>
|
||||
) : (
|
||||
<>
|
||||
<span className="row-name" title={node.Name}>
|
||||
{node.Name}
|
||||
</span>
|
||||
{managed ? (
|
||||
// Not hidden — withheld, with the reason attached. A control
|
||||
// that quietly isn't there reads as a bug; this one states the
|
||||
// rule, and the same sentence is on the group header above.
|
||||
<button
|
||||
type="button"
|
||||
className="inline-rename inline-rename--locked"
|
||||
disabled
|
||||
aria-label={`Can’t rename ${node.Name} — its name comes from subscription “${node.FromSub}” and is rewritten on the next update`}
|
||||
title={`Named by subscription “${node.FromSub}” — the feed rewrites this name on the next update. Rename it in the subscription, or add the node manually.`}
|
||||
>
|
||||
🔒
|
||||
</button>
|
||||
) : (
|
||||
<button
|
||||
ref={renameBtn}
|
||||
type="button"
|
||||
className="inline-rename"
|
||||
onClick={beginRename}
|
||||
disabled={busy}
|
||||
aria-label={`Rename node ${node.Name}`}
|
||||
title="Rename"
|
||||
>
|
||||
✎
|
||||
</button>
|
||||
)}
|
||||
</>
|
||||
)}
|
||||
<span className="badge">{proto}</span>
|
||||
{node.Stale && <span className="badge badge--warn">stale</span>}
|
||||
</div>
|
||||
{renameErr && (
|
||||
<p className="row-err" role="alert">
|
||||
{renameErr}
|
||||
</p>
|
||||
)}
|
||||
<div className="row-line2 mono">
|
||||
<span className="row-host">{host}</span>
|
||||
{hasCreds && (
|
||||
|
||||
@@ -341,7 +341,7 @@ export function Overview({
|
||||
led={{ variant: len(config?.Rules) ? 'on' : 'amber' }}
|
||||
rows={[
|
||||
{ k: 'egresses', v: String(len(config?.Egresses)) },
|
||||
{ k: 'default', v: defaultTarget(config), hot: true },
|
||||
{ k: 'default', v: defaultTarget(status, config), hot: true },
|
||||
]}
|
||||
/>
|
||||
|
||||
@@ -563,12 +563,20 @@ const NAV_LABEL: Record<Route, string> = {
|
||||
// for the apply/rollback flow, where the individual flags are the actual
|
||||
// subject of the page.)
|
||||
|
||||
function defaultTarget(config: Model | null): string {
|
||||
const rules = config?.Rules ?? []
|
||||
if (rules.length === 0) return '—'
|
||||
// The highest Order enabled rule is the effective catch-all.
|
||||
const enabled = rules.filter((r) => r.Enabled)
|
||||
if (enabled.length === 0) return 'none'
|
||||
const last = enabled.reduce((a, b) => (b.Order >= a.Order ? b : a))
|
||||
return last.Target || last.Egress || last.Name
|
||||
/** Where everything not matched by a rule goes — the engine's route `final`.
|
||||
*
|
||||
* Taken from the daemon (status.traffic.default), which reads it off the config
|
||||
* it is running. The guess this replaced was "the highest-Order enabled rule",
|
||||
* and that is not what the default is: a rule only becomes the default by having
|
||||
* NO conditions at all, whatever its Order (model.IsCatchAll), so a specific
|
||||
* high-Order rule was routinely printed here as the router's default. It also
|
||||
* described the config on disk rather than the one running, and could not see a
|
||||
* target that failed to resolve and fell back.
|
||||
*
|
||||
* Falls back to the rule count only when the daemon has not reported — never to
|
||||
* a guess about where traffic goes. */
|
||||
function defaultTarget(status: Status | null, config: Model | null): string {
|
||||
const d = status?.traffic?.default
|
||||
if (d) return d
|
||||
return len(config?.Rules) === 0 ? '—' : 'not reported'
|
||||
}
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
import './Profiles.css'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import { Button, Led, Toggle } from '../components'
|
||||
import { Button, Led, Toggle, useConfirm } from '../components'
|
||||
import { apply as apiApply, getConfig, getInterfaces, putConfig, ApiError } from '../api'
|
||||
import type { Interface, Model, Profile } from '../api'
|
||||
|
||||
@@ -36,6 +36,7 @@ function namesOf(v: unknown): string[] {
|
||||
// ---- page ------------------------------------------------------------------
|
||||
|
||||
export default function Profiles() {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<Model | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
|
||||
@@ -192,9 +193,14 @@ export default function Profiles() {
|
||||
)
|
||||
|
||||
const deleteProfile = useCallback(
|
||||
(name: string) => {
|
||||
async (name: string) => {
|
||||
if (!config) return
|
||||
if (!window.confirm(`Delete profile “${name}”? Its overrides stop applying.`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete profile',
|
||||
title: `Delete profile “${name}”?`,
|
||||
body: 'Its overrides stop applying.',
|
||||
})
|
||||
if (!ok) return
|
||||
const next = profiles.filter((p) => p.Name !== name)
|
||||
const g =
|
||||
config.Globals.ActiveProfile === name
|
||||
@@ -202,7 +208,7 @@ export default function Profiles() {
|
||||
: config.Globals
|
||||
void save({ ...config, Profiles: next, Globals: g }, `Deleted ${name}`)
|
||||
},
|
||||
[config, profiles, save],
|
||||
[config, profiles, save, confirm],
|
||||
)
|
||||
|
||||
// ---- expansion (only one profile editor open at a time) -------------------
|
||||
|
||||
@@ -58,6 +58,45 @@
|
||||
color: var(--ink);
|
||||
}
|
||||
|
||||
/* ---- active-profile banner ----
|
||||
*
|
||||
* Deliberately NOT the accent plate the save→apply bar wears above. Orange is
|
||||
* "there is something for you to do" on this faceplate, and an active WAN profile
|
||||
* is a standing condition, not a pending action. A quiet plate with an amber tag
|
||||
* reads as "note the state" — and it is the SAME amber the overridden rows below
|
||||
* carry, so the banner and its rows are visibly one story rather than two
|
||||
* unrelated oddities. */
|
||||
.rt-prof-banner {
|
||||
display: flex;
|
||||
align-items: flex-start;
|
||||
gap: 10px;
|
||||
margin: 0 0 calc(var(--u, 8px) * 2.5);
|
||||
padding: 10px 14px;
|
||||
border: 1px solid color-mix(in srgb, var(--amber) 35%, var(--groove));
|
||||
border-radius: 8px;
|
||||
background: color-mix(in srgb, var(--amber) 7%, transparent);
|
||||
font-family: var(--font-sans);
|
||||
font-size: 12px;
|
||||
line-height: 1.55;
|
||||
color: var(--dim);
|
||||
}
|
||||
.rt-prof-banner strong {
|
||||
color: var(--ink);
|
||||
font-weight: 600;
|
||||
}
|
||||
.rt-prof-tag {
|
||||
flex: none;
|
||||
margin-top: 1px;
|
||||
padding: 2px 7px;
|
||||
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
|
||||
border-radius: 999px;
|
||||
background: color-mix(in srgb, var(--amber) 12%, transparent);
|
||||
font-size: 9px;
|
||||
letter-spacing: var(--track-label);
|
||||
text-transform: uppercase;
|
||||
color: var(--amber);
|
||||
}
|
||||
|
||||
/* ---- empty state ---- */
|
||||
.rt-empty {
|
||||
padding: calc(var(--u, 8px) * 4) 0 calc(var(--u, 8px) * 3);
|
||||
@@ -289,6 +328,51 @@
|
||||
opacity: 0.62;
|
||||
}
|
||||
|
||||
/* ---- a rule the active WAN profile overrides ----
|
||||
*
|
||||
* The row itself needs no new paint: an overridden-off rule already wears `.off`
|
||||
* (it is off, whatever its switch says) and an overridden-on rule wears nothing
|
||||
* (it is on). What was missing was never colour — it was the sentence naming who
|
||||
* decided. So this is the per-row twin of the banner and borrows .rt-dead-note's
|
||||
* type wholesale: same voice, same size, one <p> margin to reset. */
|
||||
/* The same pill as .rt-badge.dead, so the two override states read as one pair,
|
||||
* but in accent — a rule the profile forces ON is active, and active is orange on
|
||||
* this faceplate. The pill is also what keeps it from running into the plain
|
||||
* "default route · final" badge beside it, where "final on · by profile" read as
|
||||
* one phrase. */
|
||||
.rt-badge.prof-on {
|
||||
padding: 1px 7px;
|
||||
border: 1px solid var(--accent-soft);
|
||||
border-radius: 999px;
|
||||
background: color-mix(in srgb, var(--accent) 10%, transparent);
|
||||
}
|
||||
|
||||
.rt-prof-note {
|
||||
margin: 0;
|
||||
}
|
||||
.rt-prof-note strong {
|
||||
color: var(--ink);
|
||||
font-weight: 600;
|
||||
}
|
||||
|
||||
/* Switch + its legend. The caption shows ONLY while a profile overrides the rule,
|
||||
* and it is what keeps the control honest: the plate says what the router is
|
||||
* doing, this says the switch is about the saved setting. A legend under the
|
||||
* control it names is the faceplate's own idiom. */
|
||||
.rt-switch {
|
||||
display: inline-flex;
|
||||
flex-direction: column;
|
||||
align-items: center;
|
||||
gap: 3px;
|
||||
}
|
||||
.rt-switch-note {
|
||||
font-family: var(--font-mono);
|
||||
font-size: 8.5px;
|
||||
letter-spacing: var(--track-label);
|
||||
text-transform: uppercase;
|
||||
color: var(--faint);
|
||||
}
|
||||
|
||||
/* ---- target chip (styled like the artifact's group:auto mono chips) ---- */
|
||||
.rt-target {
|
||||
display: inline-flex;
|
||||
|
||||
+162
-35
@@ -1,7 +1,7 @@
|
||||
import './Routing.css'
|
||||
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import type { FormEvent, ReactNode } from 'react'
|
||||
import { Button, CatSuggest, SrcPicker, Toggle } from '../components'
|
||||
import { Button, CatSuggest, SrcPicker, Toggle, useConfirm } from '../components'
|
||||
import {
|
||||
apply as apiApply,
|
||||
getConfig,
|
||||
@@ -43,6 +43,21 @@ type RRule = Rule & {
|
||||
LegacyDst?: string[] | null
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether a rule is in force, kept strictly apart from whether it is switched on.
|
||||
*
|
||||
* `on` is the EFFECTIVE state — what the router is actually doing — and every mark
|
||||
* on the row is drawn from it. `profile`/`dir` are set only when the active WAN
|
||||
* profile is the reason the two differ, so the row can name who overrode the
|
||||
* saved setting instead of leaving the operator to guess why a switch that reads
|
||||
* "on" routes nothing.
|
||||
*/
|
||||
type RuleForce = {
|
||||
on: boolean
|
||||
profile: string | null
|
||||
dir: 'enabled' | 'disabled' | null
|
||||
}
|
||||
|
||||
/**
|
||||
* Everything `Proto` can match, and nothing else. The engine understands two
|
||||
* transports and exactly ten application protocols its sniffers can name
|
||||
@@ -366,6 +381,7 @@ const browserTZName = (): string => {
|
||||
}
|
||||
|
||||
export default function Routing() {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<Model | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
const [actionError, setActionError] = useState<string | null>(null)
|
||||
@@ -483,7 +499,7 @@ export default function Routing() {
|
||||
}, [config])
|
||||
|
||||
/**
|
||||
* The verdict for one rule, or null when it can fire.
|
||||
* The daemon's verdict for one rule, or null when we have none that describes it.
|
||||
*
|
||||
* Verdicts are fetched separately from the config, so between an optimistic edit
|
||||
* and the refetch they can describe the PREVIOUS rule list. Re-checking the
|
||||
@@ -491,18 +507,53 @@ export default function Routing() {
|
||||
* badge on a working rule: a mismatch means the verdict is not about this row,
|
||||
* and no badge is the honest answer.
|
||||
*/
|
||||
const shadowOf = useCallback(
|
||||
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
|
||||
const verdictOf = useCallback(
|
||||
(r: RRule): RuleReach | null => {
|
||||
const i = modelIndex.get(r)
|
||||
if (i === undefined) return null
|
||||
const v = reach.get(i)
|
||||
if (!v || !v.unreachable || !v.shadowed_by) return null
|
||||
if (v.name !== r.Name || v.order !== r.Order) return null
|
||||
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
|
||||
if (!v || v.name !== r.Name || v.order !== r.Order) return null
|
||||
return v
|
||||
},
|
||||
[modelIndex, reach],
|
||||
)
|
||||
|
||||
const shadowOf = useCallback(
|
||||
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
|
||||
const v = verdictOf(r)
|
||||
if (!v || !v.unreachable || !v.shadowed_by) return null
|
||||
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
|
||||
},
|
||||
[verdictOf],
|
||||
)
|
||||
|
||||
/**
|
||||
* Whether a rule is IN FORCE, and who decided that — the two states this page
|
||||
* used to conflate.
|
||||
*
|
||||
* `Rule.Enabled` from /api/config is the DESIRED state: what the operator saved,
|
||||
* what the switch edits, what gets PUT back. The active WAN profile can override
|
||||
* it in either direction, and then the desired state is no longer what the router
|
||||
* is doing. Drawing the row from `Enabled` is what let a config with two rules
|
||||
* `enabled '1'` show two live switches while the engine ran one chain.
|
||||
*
|
||||
* With no verdict — an older daemon, a stopped one, or one still describing the
|
||||
* previous config — the desired state is all we know, so the row falls back to it
|
||||
* and claims no profile rather than inventing one. The `typeof` guard is for the
|
||||
* older daemon specifically: it answers without `effective_enabled` at all, and
|
||||
* reading `undefined` as false would gray out every rule on the page.
|
||||
*/
|
||||
const forceOf = useCallback(
|
||||
(r: RRule): RuleForce => {
|
||||
const v = verdictOf(r)
|
||||
if (!v || typeof v.effective_enabled !== 'boolean') {
|
||||
return { on: !!r.Enabled, profile: null, dir: null }
|
||||
}
|
||||
return { on: v.effective_enabled, profile: v.overridden_by ?? null, dir: v.override ?? null }
|
||||
},
|
||||
[verdictOf],
|
||||
)
|
||||
|
||||
// Rulesets are named domain/IP lists rules match against (rule.DstRuleset).
|
||||
const rulesets = useMemo<Ruleset[]>(
|
||||
() => [...((config?.Rulesets as Ruleset[] | null | undefined) ?? [])],
|
||||
@@ -603,13 +654,17 @@ export default function Routing() {
|
||||
// Delete the ruleset AND strip its name from any rule that referenced it, so no
|
||||
// rule is left pointing at a matcher that no longer exists (one atomic persist).
|
||||
const deleteRuleset = useCallback(
|
||||
(name: string) => {
|
||||
async (name: string) => {
|
||||
if (!config) return
|
||||
const used = rulesetUsage.get(name) ?? 0
|
||||
const warn = used
|
||||
? `Delete ruleset "${name}"? It'll be removed from ${used} rule${used === 1 ? '' : 's'} that match it.`
|
||||
: `Delete ruleset "${name}"?`
|
||||
if (!window.confirm(warn)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete ruleset',
|
||||
title: `Delete ruleset "${name}"?`,
|
||||
body: used
|
||||
? `It'll be removed from ${used} rule${used === 1 ? '' : 's'} that match it.`
|
||||
: undefined,
|
||||
})
|
||||
if (!ok) return
|
||||
const nextRulesets = rulesets.filter((r) => r.Name !== name)
|
||||
const nextRules = rules.map((r) => {
|
||||
const cur = r.DstRuleset ?? []
|
||||
@@ -620,7 +675,7 @@ export default function Routing() {
|
||||
`ruleset ${name} deleted`,
|
||||
)
|
||||
},
|
||||
[config, rules, rulesets, rulesetUsage, persist],
|
||||
[config, rules, rulesets, rulesetUsage, persist, confirm],
|
||||
)
|
||||
|
||||
const onToggle = useCallback(
|
||||
@@ -660,14 +715,19 @@ export default function Routing() {
|
||||
)
|
||||
|
||||
const onDelete = useCallback(
|
||||
(name: string) => {
|
||||
if (!window.confirm(`Delete rule "${name}"? Traffic it matched will fall through to the next rule.`)) return
|
||||
async (name: string) => {
|
||||
const ok = await confirm({
|
||||
label: 'Delete rule',
|
||||
title: `Delete rule "${name}"?`,
|
||||
body: 'Traffic it matched will fall through to the next rule.',
|
||||
})
|
||||
if (!ok) return
|
||||
commitRules(
|
||||
rules.filter((r) => r.Name !== name),
|
||||
`${name} deleted`,
|
||||
)
|
||||
},
|
||||
[rules, commitRules],
|
||||
[rules, commitRules, confirm],
|
||||
)
|
||||
|
||||
// Insert a new rule just above the catch-all (so a specific rule can actually match).
|
||||
@@ -805,7 +865,14 @@ export default function Routing() {
|
||||
)
|
||||
}
|
||||
|
||||
const enabledCount = rules.filter((r) => r.Enabled).length
|
||||
// One force verdict per displayed rule, computed once and handed down — the row,
|
||||
// the counter and the banner must all be reading the SAME answer.
|
||||
const force = rules.map((r) => forceOf(r))
|
||||
// EFFECTIVE, not configured. A counter that added up saved switches said "2 / 2
|
||||
// active" for a config the router was running one rule of.
|
||||
const enabledCount = force.filter((f) => f.on).length
|
||||
const overridden = force.filter((f) => f.profile !== null)
|
||||
const overrideProfile = overridden[0]?.profile ?? null
|
||||
|
||||
return (
|
||||
<section className="page" aria-label="Routing rules">
|
||||
@@ -814,12 +881,28 @@ export default function Routing() {
|
||||
Rules run top to bottom on the bus — the <strong>first match wins</strong>. Traffic that
|
||||
reaches the bottom follows the default route.
|
||||
</p>
|
||||
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules active`}>
|
||||
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules in force`}>
|
||||
{enabledCount}
|
||||
<small> / {rules.length} active</small>
|
||||
<small> / {rules.length} in force</small>
|
||||
</span>
|
||||
</div>
|
||||
|
||||
{/* Said once at the top, so the per-row badges below read as consequences of
|
||||
one thing rather than as N unrelated oddities. Only shown when a profile
|
||||
actually changed something: a router that uses no profiles, or one whose
|
||||
profile agrees with every saved switch, gets no banner at all. */}
|
||||
{overrideProfile && (
|
||||
<p className="rt-prof-banner" role="status">
|
||||
<span className="rt-prof-tag mono">profile</span>
|
||||
<span>
|
||||
<strong className="mono">{overrideProfile}</strong> is the active WAN profile and is
|
||||
overriding {overridden.length === 1 ? '1 rule' : `${overridden.length} rules`} below. The
|
||||
switches keep showing what you saved; the rows show what the router is running. Change
|
||||
which rules a profile forces on the <strong>Profiles</strong> page.
|
||||
</span>
|
||||
</p>
|
||||
)}
|
||||
|
||||
{actionError && (
|
||||
<p className="page-error" role="alert">
|
||||
{actionError}
|
||||
@@ -867,6 +950,7 @@ export default function Routing() {
|
||||
busy={saving}
|
||||
editingOther={editingRule !== null}
|
||||
shadow={shadowOf(r)}
|
||||
force={force[i]}
|
||||
onEdit={onEditRule}
|
||||
onToggle={onToggle}
|
||||
onMove={onMove}
|
||||
@@ -917,6 +1001,7 @@ function RuleRow({
|
||||
busy,
|
||||
editingOther,
|
||||
shadow,
|
||||
force,
|
||||
onEdit,
|
||||
onToggle,
|
||||
onMove,
|
||||
@@ -929,6 +1014,9 @@ function RuleRow({
|
||||
editingOther: boolean
|
||||
/** Set when the daemon reports this rule can never fire; null when it can. */
|
||||
shadow: { by: string; byOrder: number; reason: string } | null
|
||||
/** Whether the rule is IN FORCE, and which profile decided that (see RuleForce).
|
||||
* Every mark on this row comes from here; `rule.Enabled` drives only the switch. */
|
||||
force: RuleForce
|
||||
onEdit: (name: string) => void
|
||||
onToggle: (name: string) => void
|
||||
onMove: (name: string, dir: 'up' | 'down') => void
|
||||
@@ -956,14 +1044,14 @@ function RuleRow({
|
||||
const isDefault = isCatchAll(rule) && !dead
|
||||
const target = effectiveTarget(rule)
|
||||
const tone = targetTone(target)
|
||||
const cls = [
|
||||
'rt-rule',
|
||||
rule.Enabled ? '' : 'off',
|
||||
isDefault ? 'final' : '',
|
||||
inert ? 'dead' : '',
|
||||
]
|
||||
// Dimmed by the EFFECTIVE state, never by the saved one. A rule the active
|
||||
// profile switched off is not in force, and the row has to read that way even
|
||||
// though its switch — which edits the saved setting — is still on.
|
||||
const cls = ['rt-rule', force.on ? '' : 'off', isDefault ? 'final' : '', inert ? 'dead' : '']
|
||||
.filter(Boolean)
|
||||
.join(' ')
|
||||
// What the switch says, spelled out, for the moment the two disagree.
|
||||
const savedState = rule.Enabled ? 'on' : 'off'
|
||||
|
||||
return (
|
||||
<li className={cls}>
|
||||
@@ -997,6 +1085,11 @@ function RuleRow({
|
||||
{isDefault && <span className="rt-badge">default route · final</span>}
|
||||
{unmigrated && <span className="rt-badge dead">held off · not migrated</span>}
|
||||
{dead && <span className="rt-badge dead">never applies</span>}
|
||||
{/* Amber for the rule the profile switched OFF (warn semantics: wired but
|
||||
not connected), accent for the one it switched ON — orange is the
|
||||
faceplate's active state, and a force-enabled rule is exactly that. */}
|
||||
{force.dir === 'disabled' && <span className="rt-badge dead">off · by profile</span>}
|
||||
{force.dir === 'enabled' && <span className="rt-badge prof-on">on · by profile</span>}
|
||||
</div>
|
||||
<div className="rt-match">
|
||||
{unmigrated ? (
|
||||
@@ -1024,6 +1117,22 @@ function RuleRow({
|
||||
<Matchers rule={rule} />
|
||||
)}
|
||||
</div>
|
||||
{/* Added BELOW the matchers, not instead of them: the rule's conditions are
|
||||
still worth reading — the operator is deciding whether to change the
|
||||
profile or the rule.
|
||||
Two clauses only. The banner at the top of the page already carries the
|
||||
general explanation and the way to change it, and a profile that
|
||||
overrides several rules would otherwise repeat that paragraph on every
|
||||
one of them. What is left is the part only this row can say: whether it
|
||||
is in force, and what its own switch is showing instead. */}
|
||||
{force.profile && (
|
||||
<p className="rt-dead-note rt-prof-note">
|
||||
{force.dir === 'disabled' ? 'Not in force' : 'In force'} — profile{' '}
|
||||
<strong className="mono">{force.profile}</strong> switches this rule{' '}
|
||||
{force.dir === 'disabled' ? 'off' : 'on'}. The switch still reads{' '}
|
||||
<strong>{savedState}</strong>: that is the saved setting.
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
|
||||
<div className={`rt-target ${tone}`} title={`target: ${target}`}>
|
||||
@@ -1049,16 +1158,34 @@ function RuleRow({
|
||||
emits neither), so enabling it here would save a live rule with no
|
||||
destination left at all — the catch-all this tripwire exists to
|
||||
prevent. `shaterd migrate` clears LegacyDst and the switch comes back. */}
|
||||
<Toggle
|
||||
pressed={rule.Enabled}
|
||||
onChange={() => onToggle(rule.Name)}
|
||||
label={
|
||||
unmigrated
|
||||
? `Rule ${rule.Name} is held disabled until the config is migrated`
|
||||
: `${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`
|
||||
}
|
||||
disabled={frozen || unmigrated}
|
||||
/>
|
||||
{/* The switch edits the SAVED setting and nothing else, so it keeps showing
|
||||
rule.Enabled even while the active profile forces the opposite. Mirroring
|
||||
the effective state here would be worse than the bug it replaces: the
|
||||
operator would flip a switch that was never theirs, and the PUT would
|
||||
write the profile's decision into UCI as if they had chosen it. The row
|
||||
above says what the router is doing; the "saved" caption says what this
|
||||
control is for. */}
|
||||
<span className="rt-switch">
|
||||
<Toggle
|
||||
pressed={rule.Enabled}
|
||||
onChange={() => onToggle(rule.Name)}
|
||||
label={
|
||||
unmigrated
|
||||
? `Rule ${rule.Name} is held disabled until the config is migrated`
|
||||
: force.profile
|
||||
? `Saved setting for rule ${rule.Name} is ${savedState}; profile ${force.profile} is forcing it ${
|
||||
force.dir === 'disabled' ? 'off' : 'on'
|
||||
}. This switch changes the saved setting only.`
|
||||
: `${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`
|
||||
}
|
||||
disabled={frozen || unmigrated}
|
||||
/>
|
||||
{force.profile && (
|
||||
<span className="rt-switch-note" aria-hidden="true">
|
||||
saved
|
||||
</span>
|
||||
)}
|
||||
</span>
|
||||
<button
|
||||
type="button"
|
||||
className="rt-del"
|
||||
|
||||
@@ -704,6 +704,13 @@
|
||||
.tg-test--bad .tg-test-msg {
|
||||
color: var(--crit);
|
||||
}
|
||||
/* "Nothing measured this" is not a failure and must never be dressed as one: an
|
||||
unlit lamp and the faintest text on the card, the same register the group
|
||||
readout uses for its unmeasured state. */
|
||||
.tg-test--none .tg-test-msg {
|
||||
font-family: var(--font-sans);
|
||||
color: var(--faint);
|
||||
}
|
||||
.tg-test--wait .tg-test-msg {
|
||||
color: var(--amber);
|
||||
}
|
||||
@@ -1038,6 +1045,221 @@
|
||||
}
|
||||
|
||||
/* ---- responsive ---- */
|
||||
/* ---- chain hop rail ----
|
||||
* The chain section's signature, and the one place this card spends any
|
||||
* boldness: the path is drawn as a CONDUCTOR with a numbered lamp at each hop,
|
||||
* and the conductor is SEVERED below the first hop that was probed and did not
|
||||
* answer. A chain is a single series path, so the question is never "how many
|
||||
* hops are green", it is "where does my traffic stop" — and a broken line answers
|
||||
* that before a word has been read.
|
||||
*
|
||||
* The two marks carry two different facts and must not be conflated:
|
||||
* - the LAMP is that hop's own measurement (good / warn / crit / unlit), and a
|
||||
* hop below the break that answered keeps its green, because it did answer;
|
||||
* - the CONDUCTOR is reachability through the path, which really does stop.
|
||||
*
|
||||
* Orange is untouched here. Semantics carry every colour, and everything that is
|
||||
* not a lamp is groove-grey. No transitions and no animation anywhere in the
|
||||
* rail, so there is nothing for reduced-motion to switch off. */
|
||||
.ch-rail {
|
||||
gap: 8px;
|
||||
}
|
||||
.ch-eyebrow {
|
||||
font-family: var(--font-mono);
|
||||
font-size: 9px;
|
||||
letter-spacing: var(--track-label);
|
||||
text-transform: uppercase;
|
||||
color: var(--faint);
|
||||
}
|
||||
.ch-hops {
|
||||
--ch-num: 1.8ch; /* the engraved hop number's gutter */
|
||||
--ch-gap: 8px;
|
||||
--ch-led: 10px; /* must match .led's width */
|
||||
--ch-lampy: 14px; /* row top → lamp centre; the conductor's anchor */
|
||||
/* x of the conductor: the number gutter, one gap, then the lamp's centre */
|
||||
--ch-spine: calc(var(--ch-num) + var(--ch-gap) + var(--ch-led) / 2);
|
||||
list-style: none;
|
||||
margin: 0;
|
||||
padding: 0;
|
||||
}
|
||||
.ch-hop {
|
||||
position: relative;
|
||||
display: grid;
|
||||
grid-template-columns: var(--ch-num) var(--ch-led) minmax(0, 1fr);
|
||||
column-gap: var(--ch-gap);
|
||||
align-items: start;
|
||||
}
|
||||
|
||||
/* the conductor — two halves per row, so a break lands on one link only */
|
||||
.ch-hop::before,
|
||||
.ch-hop::after {
|
||||
content: '';
|
||||
position: absolute;
|
||||
left: var(--ch-spine);
|
||||
width: 2px;
|
||||
margin-left: -1px;
|
||||
/* Brighter than a plain groove: this line IS the readout, and at groove
|
||||
strength it disappeared into the panel and took the whole idea with it. */
|
||||
background: color-mix(in srgb, var(--dim) 55%, var(--groove));
|
||||
}
|
||||
.ch-hop::before {
|
||||
top: 0;
|
||||
height: calc(var(--ch-lampy) - var(--ch-led) / 2 - 3px);
|
||||
}
|
||||
.ch-hop::after {
|
||||
top: calc(var(--ch-lampy) + var(--ch-led) / 2 + 3px);
|
||||
bottom: 0;
|
||||
}
|
||||
/* Nothing feeds hop 1 from above, and nothing leaves the exit downward — the
|
||||
path starts and ends inside this rail. */
|
||||
.ch-hop--first::before {
|
||||
display: none;
|
||||
}
|
||||
.ch-hop--exit::after {
|
||||
bottom: auto;
|
||||
height: 9px;
|
||||
}
|
||||
/* …the exit ends on a crossbar instead of trailing off: end of line. */
|
||||
.ch-hop--exit .ch-socket {
|
||||
position: relative;
|
||||
}
|
||||
.ch-hop--exit .ch-socket::after {
|
||||
content: '';
|
||||
position: absolute;
|
||||
left: 50%;
|
||||
transform: translateX(-50%);
|
||||
top: calc(var(--ch-lampy) + var(--ch-led) / 2 + 12px);
|
||||
width: 11px;
|
||||
height: 2px;
|
||||
background: color-mix(in srgb, var(--dim) 55%, var(--groove));
|
||||
}
|
||||
|
||||
/* THE SEVER. Everything from the dead hop's outgoing link downward is drawn as a
|
||||
broken conductor: unmistakably not-a-line at a glance, and unmistakably not a
|
||||
colour, because a colour here would compete with the lamps that carry health. */
|
||||
.ch-hop--dead::after,
|
||||
.ch-hop--severed::before,
|
||||
.ch-hop--severed::after {
|
||||
background: repeating-linear-gradient(
|
||||
to bottom,
|
||||
color-mix(in srgb, var(--dim) 45%, var(--groove)) 0 3px,
|
||||
transparent 3px 7px
|
||||
);
|
||||
}
|
||||
|
||||
.ch-num {
|
||||
font-size: 10px;
|
||||
line-height: calc(var(--ch-lampy) * 2);
|
||||
text-align: right;
|
||||
color: var(--faint);
|
||||
}
|
||||
.ch-socket {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
height: calc(var(--ch-lampy) * 2);
|
||||
}
|
||||
.ch-body {
|
||||
min-width: 0;
|
||||
/* Separates one hop from the next. The conductor runs through this space, so
|
||||
too little of it and two hops read as one wrapped row. */
|
||||
padding-bottom: 8px;
|
||||
}
|
||||
.ch-l1 {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
flex-wrap: wrap;
|
||||
gap: 8px;
|
||||
min-height: calc(var(--ch-lampy) * 2);
|
||||
}
|
||||
.ch-name {
|
||||
font-size: 12px;
|
||||
color: var(--ink);
|
||||
overflow-wrap: anywhere;
|
||||
}
|
||||
.ch-hop--untested .ch-name {
|
||||
color: var(--dim);
|
||||
}
|
||||
/* The exit marker is NEUTRAL on purpose. The config path above this rail tags its
|
||||
exit green, which is free there — but in here green means "answering", and a
|
||||
green badge on the last hop would read as a health claim about it. */
|
||||
.ch-tag {
|
||||
font-family: var(--font-mono);
|
||||
font-size: 8.5px;
|
||||
letter-spacing: 0.14em;
|
||||
text-transform: uppercase;
|
||||
color: var(--faint);
|
||||
}
|
||||
.ch-delay {
|
||||
font-size: 11.5px;
|
||||
font-weight: 700;
|
||||
color: var(--ink);
|
||||
}
|
||||
.ch-quiet {
|
||||
font-family: var(--font-sans);
|
||||
font-size: 12px;
|
||||
color: var(--faint);
|
||||
}
|
||||
.ch-age {
|
||||
margin-left: auto;
|
||||
font-size: 10.5px;
|
||||
color: var(--faint);
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
/* The counters read exactly as they do on a group card — alive out of TESTED,
|
||||
with the untested remainder as a quiet aside only when there is one. Same
|
||||
register, same weights, deliberately not a second dialect. */
|
||||
.ch-l2 {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
flex-wrap: wrap;
|
||||
gap: 8px;
|
||||
margin-top: 1px;
|
||||
font-size: 11.5px;
|
||||
letter-spacing: 0.02em;
|
||||
color: var(--dim);
|
||||
}
|
||||
.ch-count {
|
||||
font-size: 12px;
|
||||
color: var(--dim);
|
||||
white-space: nowrap;
|
||||
}
|
||||
.ch-count b {
|
||||
font-size: 14px;
|
||||
font-weight: 700;
|
||||
color: var(--ink);
|
||||
}
|
||||
.ch-hop--dead .ch-count b {
|
||||
color: var(--crit);
|
||||
}
|
||||
.ch-word {
|
||||
font-size: 10.5px;
|
||||
letter-spacing: 0.12em;
|
||||
text-transform: uppercase;
|
||||
color: var(--faint);
|
||||
}
|
||||
.ch-dead {
|
||||
padding: 1px 6px;
|
||||
border-radius: 4px;
|
||||
background: color-mix(in srgb, var(--crit) 12%, transparent);
|
||||
font-size: 10.5px;
|
||||
color: var(--crit);
|
||||
white-space: nowrap;
|
||||
cursor: help;
|
||||
}
|
||||
.ch-rest {
|
||||
font-size: 10.5px;
|
||||
color: var(--faint);
|
||||
}
|
||||
.ch-now {
|
||||
max-width: 28ch;
|
||||
overflow: hidden;
|
||||
text-overflow: ellipsis;
|
||||
white-space: nowrap;
|
||||
font-size: 10.5px;
|
||||
color: var(--dim);
|
||||
}
|
||||
|
||||
@media (max-width: 640px) {
|
||||
.tg-sec-hd {
|
||||
flex-wrap: wrap;
|
||||
@@ -1060,6 +1282,20 @@
|
||||
.gh-now {
|
||||
max-width: 100%;
|
||||
}
|
||||
/* On a phone the age stamp stops being pushed to a lonely right edge and just
|
||||
joins the end of the hop's line; the selected node gets the full width
|
||||
instead of an ellipsis it doesn't need. */
|
||||
.ch-age {
|
||||
margin-left: 0;
|
||||
}
|
||||
.ch-now {
|
||||
max-width: 100%;
|
||||
}
|
||||
/* Every field of a hop wraps onto its own line at this width, so the gap
|
||||
between hops has to grow with them or the rail reads as one block of text. */
|
||||
.ch-body {
|
||||
padding-bottom: 12px;
|
||||
}
|
||||
.gh-mems {
|
||||
max-height: 260px;
|
||||
}
|
||||
|
||||
+407
-99
@@ -1,6 +1,6 @@
|
||||
import './Targets.css'
|
||||
import { Fragment, useCallback, useEffect, useMemo, useRef, useState } from 'react'
|
||||
import { Button, Led, Toggle } from '../components'
|
||||
import { Button, Led, Toggle, useConfirm } from '../components'
|
||||
import type { LedVariant } from '../components'
|
||||
import {
|
||||
apply as apiApply,
|
||||
@@ -22,6 +22,8 @@ import type {
|
||||
GroupTestResult,
|
||||
GroupTestStatus,
|
||||
Chain,
|
||||
ChainHealth,
|
||||
ChainHopHealth,
|
||||
Egress,
|
||||
Interface,
|
||||
Node,
|
||||
@@ -161,18 +163,18 @@ function renameReferences(m: Model, kind: RefKind, from: string, to: string): Mo
|
||||
}
|
||||
|
||||
/**
|
||||
* The sentence a delete confirmation appends: what still points at this target,
|
||||
* and what happens to it. Empty list ⇒ an explicit "nothing references it", so
|
||||
* the operator can delete a stray with confidence instead of guessing.
|
||||
* The body of a delete confirmation: what still points at this target, and what
|
||||
* happens to it. Empty list ⇒ an explicit "nothing references it", so the
|
||||
* operator can delete a stray with confidence instead of guessing.
|
||||
*/
|
||||
function refWarning(refs: RefSite[]): string {
|
||||
if (refs.length === 0) return ' Nothing references it.'
|
||||
if (refs.length === 0) return 'Nothing references it.'
|
||||
const shown = refs.slice(0, 4).map((r) => r.label)
|
||||
const more = refs.length - shown.length
|
||||
const list = `${shown.join(', ')}${more > 0 ? `, and ${more} more` : ''}`
|
||||
return refs.length === 1
|
||||
? ` It is referenced by ${list}, whose traffic will be blocked (an unresolved target never falls through to the default route).`
|
||||
: ` It is referenced by ${refs.length} places — ${list} — whose traffic will be blocked (an unresolved target never falls through to the default route).`
|
||||
? `It is referenced by ${list}, whose traffic will be blocked (an unresolved target never falls through to the default route).`
|
||||
: `It is referenced by ${refs.length} places — ${list} — whose traffic will be blocked (an unresolved target never falls through to the default route).`
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -436,14 +438,14 @@ const normalizeTest = (st: GroupTestStatus): GroupTestStatus => ({
|
||||
})
|
||||
|
||||
/**
|
||||
* How the header names the reach of a running exit test. The name matters more
|
||||
* How the header names the reach of a running refresh pass. The name matters more
|
||||
* than the number when there is only one: "auto" tells the operator which button
|
||||
* they pressed; "1 target" tells them nothing they didn't already know.
|
||||
*/
|
||||
function scopeLabel(scope: string[], targetCount: number): string {
|
||||
if (scope.length === 1) return scope[0]
|
||||
if (scope.length === 0) return 'exits' // pre-scope daemon — say nothing false
|
||||
return scope.length >= targetCount ? 'every exit' : `${scope.length} exits`
|
||||
if (scope.length === 0) return 'targets' // pre-scope daemon — say nothing false
|
||||
return scope.length >= targetCount ? 'every target' : `${scope.length} targets`
|
||||
}
|
||||
|
||||
/** Which editor (add or edit-by-name) is open within a section. */
|
||||
@@ -457,6 +459,7 @@ interface Opt {
|
||||
// ---- page ------------------------------------------------------------------
|
||||
|
||||
export default function Targets() {
|
||||
const confirm = useConfirm()
|
||||
const [config, setConfig] = useState<Model | null>(null)
|
||||
const [loadError, setLoadError] = useState<string | null>(null)
|
||||
|
||||
@@ -589,10 +592,12 @@ export default function Targets() {
|
||||
[health],
|
||||
)
|
||||
|
||||
// ---- group/chain exit test: how fast, through which node, out which address --
|
||||
// The POST only kicks a run off, and a 2 s poll of the GET carries progress
|
||||
// plus every result so far. One endpoint covers groups and chains alike:
|
||||
// POST with a group or chain name tests that one; an empty name tests them all.
|
||||
// ---- out-of-turn refresh: how fast, through which node, out which address ----
|
||||
// The POST does NOT dial. It asks the observatory — the only thing in the daemon
|
||||
// that measures anything, and it measures along the real dial path — to come
|
||||
// round out of turn; a 2 s poll of the GET carries progress plus every reading
|
||||
// so far. One endpoint covers groups and chains alike: POST with a name refreshes
|
||||
// that one, an empty name refreshes them all.
|
||||
const [gtest, setGtest] = useState<GroupTestStatus>(IDLE_TEST)
|
||||
const [gtestErr, setGtestErr] = useState<string | null>(null)
|
||||
const [polling, setPolling] = useState(false)
|
||||
@@ -635,7 +640,7 @@ export default function Targets() {
|
||||
void readTest().then((st) => {
|
||||
if (!alive || !st || st.running) return
|
||||
setPolling(false)
|
||||
flash('Group test complete')
|
||||
flash('Readings refreshed')
|
||||
})
|
||||
}, 2000)
|
||||
return () => {
|
||||
@@ -651,16 +656,16 @@ export default function Targets() {
|
||||
if (r.started) {
|
||||
setGtestErr(null)
|
||||
setPolling(true)
|
||||
flash(name ? `Testing ${name}…` : 'Testing every exit…')
|
||||
flash(name ? `Refreshing ${name}…` : 'Refreshing every reading…')
|
||||
void readTest()
|
||||
} else if (r.reason === 'already running') {
|
||||
setPolling(true) // pick up the run someone else started
|
||||
flash('A group test is already running')
|
||||
setPolling(true) // pick up the pass someone else started
|
||||
flash('The prober is already refreshing')
|
||||
} else {
|
||||
flash(`Couldn’t start the test — ${r.reason || 'the daemon refused it'}`)
|
||||
flash(`Couldn’t ask for a refresh — ${r.reason || 'the daemon refused it'}`)
|
||||
}
|
||||
} catch (e) {
|
||||
flash(`Couldn’t start the test — ${errText(e)}`)
|
||||
flash(`Couldn’t ask for a refresh — ${errText(e)}`)
|
||||
}
|
||||
},
|
||||
[flash, readTest],
|
||||
@@ -787,13 +792,18 @@ export default function Targets() {
|
||||
)
|
||||
|
||||
const removeGroup = useCallback(
|
||||
(name: string) => {
|
||||
async (name: string) => {
|
||||
if (!config) return
|
||||
const refs = findReferences(config, 'group', name)
|
||||
if (!window.confirm(`Delete group “${name}”?${refWarning(refs)}`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete group',
|
||||
title: `Delete group “${name}”?`,
|
||||
body: refWarning(refs),
|
||||
})
|
||||
if (!ok) return
|
||||
void save({ ...config, Groups: groups.filter((g) => g.Name !== name) }, `Deleted ${name}`)
|
||||
},
|
||||
[config, groups, save],
|
||||
[config, groups, save, confirm],
|
||||
)
|
||||
|
||||
// ---- chain mutations ------------------------------------------------------
|
||||
@@ -820,13 +830,18 @@ export default function Targets() {
|
||||
)
|
||||
|
||||
const removeChain = useCallback(
|
||||
(name: string) => {
|
||||
async (name: string) => {
|
||||
if (!config) return
|
||||
const refs = findReferences(config, 'chain', name)
|
||||
if (!window.confirm(`Delete chain “${name}”?${refWarning(refs)}`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete chain',
|
||||
title: `Delete chain “${name}”?`,
|
||||
body: refWarning(refs),
|
||||
})
|
||||
if (!ok) return
|
||||
void save({ ...config, Chains: chains.filter((c) => c.Name !== name) }, `Deleted ${name}`)
|
||||
},
|
||||
[config, chains, save],
|
||||
[config, chains, save, confirm],
|
||||
)
|
||||
|
||||
// ---- egress mutations -----------------------------------------------------
|
||||
@@ -853,13 +868,18 @@ export default function Targets() {
|
||||
)
|
||||
|
||||
const removeEgress = useCallback(
|
||||
(name: string) => {
|
||||
async (name: string) => {
|
||||
if (!config) return
|
||||
const refs = findReferences(config, 'egress', name)
|
||||
if (!window.confirm(`Delete egress “${name}”?${refWarning(refs)}`)) return
|
||||
const ok = await confirm({
|
||||
label: 'Delete egress',
|
||||
title: `Delete egress “${name}”?`,
|
||||
body: refWarning(refs),
|
||||
})
|
||||
if (!ok) return
|
||||
void save({ ...config, Egresses: egresses.filter((e) => e.Name !== name) }, `Deleted ${name}`)
|
||||
},
|
||||
[config, egresses, save],
|
||||
[config, egresses, save, confirm],
|
||||
)
|
||||
|
||||
const busy = saving || applying
|
||||
@@ -894,10 +914,11 @@ export default function Targets() {
|
||||
<h2 className="tg-sec-title">Groups</h2>
|
||||
<span className="tg-sec-count mono">{groups.length} configured</span>
|
||||
{/* The observatory's background probing is invisible by design — it
|
||||
keeps every used group's and chain's numbers fresh on its own. The
|
||||
one manual run left is the exit test: it is scoped to the groups
|
||||
and chains it names, so its progress says WHICH, and its badge
|
||||
lands only on those cards. */}
|
||||
keeps every used group's and chain's numbers fresh on its own, along
|
||||
the path traffic actually takes. The one manual control left does
|
||||
not measure anything itself: it asks that prober to come round out
|
||||
of turn. It is scoped to the groups and chains it names, so its
|
||||
progress says WHICH, and its badge lands only on those cards. */}
|
||||
<div className="tg-sec-ctl">
|
||||
{groupHealthOn && (
|
||||
<>
|
||||
@@ -905,11 +926,11 @@ export default function Targets() {
|
||||
<span
|
||||
className="tg-run tg-run--exit"
|
||||
role="status"
|
||||
title="An exit test sends one connection through each group or chain it covers and reports the delay and the address the internet sees."
|
||||
title="The background prober is measuring the targets this refresh covers, along the path each one's traffic really takes."
|
||||
>
|
||||
<Led variant="amber" pulse />
|
||||
<span className="tg-run-what">
|
||||
exit test · {scopeLabel(asArray(gtest.scope), groups.length + chains.length)}
|
||||
refreshing · {scopeLabel(asArray(gtest.scope), groups.length + chains.length)}
|
||||
</span>
|
||||
<span className="tg-run-n mono">
|
||||
{gtest.done}/{gtest.total}
|
||||
@@ -919,9 +940,9 @@ export default function Targets() {
|
||||
<Button
|
||||
onClick={() => void runTest()}
|
||||
disabled={busy || !config || (groups.length === 0 && chains.length === 0) || gtest.running}
|
||||
title="Send one connection through each group and chain and report the delay and the exit address the internet sees"
|
||||
title="Ask the background prober to measure every group and chain out of turn, then show what it measured. The panel opens no connection of its own."
|
||||
>
|
||||
{gtest.running ? 'Testing…' : 'Test every exit'}
|
||||
{gtest.running ? 'Refreshing…' : 'Refresh every reading'}
|
||||
</Button>
|
||||
</>
|
||||
)}
|
||||
@@ -941,6 +962,13 @@ export default function Targets() {
|
||||
dials out through a tunnel measures them through that tunnel, so the same node can be alive
|
||||
in one group and dead in another.
|
||||
</p>
|
||||
<p className="tg-sec-note">
|
||||
One thing measures, and the panel is not it. A background prober walks every path your
|
||||
rules use — hop by hop, exactly as traffic goes — and every number on this page is a read
|
||||
of what it found. <strong>Refresh every reading</strong> asks it to come round out of turn
|
||||
instead of waiting for the next pass; it opens no connection of its own, so a target no
|
||||
rule routes through has nothing to report and says so.
|
||||
</p>
|
||||
|
||||
{groupHealthOn && healthErr && (
|
||||
<p className="tg-test-err" role="alert">
|
||||
@@ -951,7 +979,7 @@ export default function Targets() {
|
||||
|
||||
{groupHealthOn && gtestErr && (
|
||||
<p className="tg-test-err" role="alert">
|
||||
Couldn’t read the test results — {gtestErr}.{' '}
|
||||
Couldn’t read the refreshed numbers — {gtestErr}.{' '}
|
||||
<button className="linkish" onClick={() => void readTest()}>
|
||||
Retry
|
||||
</button>
|
||||
@@ -1094,7 +1122,9 @@ export default function Targets() {
|
||||
chain={c}
|
||||
busy={busy}
|
||||
showHealth={groupHealthOn}
|
||||
used={healthByChain.get(c.Name)?.used}
|
||||
// The whole chain health record, not just `.used` — the card
|
||||
// renders the observatory's per-hop measurements from it.
|
||||
health={healthByChain.get(c.Name)}
|
||||
test={testByGroup.get(c.Name)}
|
||||
// The badge is this card's business only when the run names it.
|
||||
testing={gtest.running && testScope.has(c.Name)}
|
||||
@@ -1216,7 +1246,7 @@ function GroupRow({
|
||||
group: Group
|
||||
busy: boolean
|
||||
/** Group health checks are on (Settings). When false, the card drops its health
|
||||
* readout, its exit-test readout and its Test button — it is config only. */
|
||||
* readout, its end-to-end reading and its Refresh button — it is config only. */
|
||||
showHealth: boolean
|
||||
/** This group's membership health, or undefined when the engine hasn't built
|
||||
* it (not applied yet, or dropped for having no usable members). */
|
||||
@@ -1226,14 +1256,14 @@ function GroupRow({
|
||||
healthKnown: boolean
|
||||
test?: GroupTestResult
|
||||
/**
|
||||
* A group exit test covering THIS group is in flight.
|
||||
* A refresh pass covering THIS group is in flight.
|
||||
*
|
||||
* Deliberately not "a test is running": the caller resolves it against the run's
|
||||
* scope. There is no per-card equivalent for the health run — that one measures
|
||||
* every group at once and is reported once, in the section header.
|
||||
*/
|
||||
testing: boolean
|
||||
/** Any exit test is in flight; the daemon runs one at a time. */
|
||||
/** Any refresh pass is in flight; the daemon runs one at a time. */
|
||||
testBusy: boolean
|
||||
onTest: () => void
|
||||
onEdit: () => void
|
||||
@@ -1293,7 +1323,11 @@ function GroupRow({
|
||||
health={health}
|
||||
healthKnown={healthKnown}
|
||||
/>
|
||||
<GroupTestReadout test={test} pending={testing && !test} />
|
||||
<GroupTestReadout
|
||||
test={test}
|
||||
pending={testing && !test}
|
||||
hideAbsence={health?.used === false}
|
||||
/>
|
||||
</>
|
||||
)}
|
||||
</div>
|
||||
@@ -1304,7 +1338,7 @@ function GroupRow({
|
||||
editLabel={`Edit group ${group.Name}`}
|
||||
deleteLabel={`Delete group ${group.Name}`}
|
||||
onTest={showHealth ? onTest : undefined}
|
||||
testLabel={showHealth ? `Test the exit of group ${group.Name}` : undefined}
|
||||
testLabel={showHealth ? `Refresh the reading for group ${group.Name}` : undefined}
|
||||
testDisabled={testBusy}
|
||||
/>
|
||||
</li>
|
||||
@@ -1363,21 +1397,7 @@ function GroupHealthReadout({
|
||||
// its members would stay "untested" forever. That is a fact about the ROUTING
|
||||
// CONFIG, not about the members — so instead of counters that could only ever
|
||||
// read as a permanent unknown, the card says so, quietly: unused, not unwell.
|
||||
if (!health.used) {
|
||||
return (
|
||||
<div className="gh gh--unused">
|
||||
<div className="gh-line">
|
||||
<span
|
||||
className="gh-unused"
|
||||
title="No enabled rule routes through this group, so its members are not probed. Add it to a rule to see health."
|
||||
>
|
||||
unused
|
||||
</span>
|
||||
<span className="gh-quiet">not probed — no enabled rule routes through this group</span>
|
||||
</div>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
if (!health.used) return <NotRoutedNote kind="group" />
|
||||
|
||||
const v = verdictOf(health)
|
||||
const { total, tested, alive, dead, untested } = health
|
||||
@@ -1540,6 +1560,56 @@ function GroupHealthReadout({
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* The card's answer when NOTHING ROUTES THROUGH THIS TARGET. Shared by the group
|
||||
* card and the chain card, because it is the same misunderstanding on both.
|
||||
*
|
||||
* It has to carry two statements, and the old one-liner ("not probed — no enabled
|
||||
* rule routes through this group") only carried the first. Read fast it still
|
||||
* landed as a verdict: a card that normally shows health and today shows a grey
|
||||
* pill reads as "the health is bad". So the two meanings are now separated, on
|
||||
* purpose and in this order:
|
||||
*
|
||||
* 1. the ROUTING FACT — nothing routes here, so nothing measures it;
|
||||
* 2. the NON-FACT — this is not a health reading at all. Absent numbers are
|
||||
* absence of measurement, never failure.
|
||||
*
|
||||
* For a GROUP there is a third line, and it is the confusion this whole change
|
||||
* exists to end: a group used only as a hop inside a chain is never routed to
|
||||
* DIRECTLY, so it correctly reads unused here while carrying real traffic as a
|
||||
* hop. Its health is measured at that hop, on the chain's card.
|
||||
*
|
||||
* Unused is neutral — groove-grey, never amber, never crit. It is a state of the
|
||||
* config, and the config is not sick.
|
||||
*/
|
||||
function NotRoutedNote({ kind }: { kind: 'group' | 'chain' }) {
|
||||
return (
|
||||
<div className="gh gh--unused">
|
||||
<div className="gh-line">
|
||||
<span className="gh-unused">unused</span>
|
||||
<span className="gh-quiet">
|
||||
No enabled rule routes through this {kind}, so the observatory never probes it.
|
||||
</span>
|
||||
</div>
|
||||
<p className="gh-say">
|
||||
That is a routing fact, not a health reading. There are no numbers here because nothing
|
||||
measured this {kind} — not because it failed.
|
||||
</p>
|
||||
{kind === 'group' ? (
|
||||
<p className="gh-say">
|
||||
A group used only as a hop inside a chain reads unused here on purpose: the rules point at
|
||||
the chain, not at the group. Its members are measured at that hop, so its real health is on
|
||||
that chain’s card, hop by hop.
|
||||
</p>
|
||||
) : (
|
||||
<p className="gh-say">
|
||||
Point a rule at this chain and the observatory starts measuring every hop within seconds.
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* One group's member rows, fetched on demand.
|
||||
*
|
||||
@@ -1645,7 +1715,9 @@ function MemberRow({ member }: { member: GroupMemberHealth }) {
|
||||
}
|
||||
|
||||
/**
|
||||
* What a group test found, in the four states it actually has.
|
||||
* What the OBSERVATORY measured for this target end to end, in the four states it
|
||||
* actually has. Nothing here was dialled by the panel — it is a read of the
|
||||
* background prober's own measurement along the real path.
|
||||
*
|
||||
* The one worth spelling out: `ok` with an EMPTY `exit_ip` is a SUCCESS. The
|
||||
* delay was measured; only the address lookup came back empty. Rendering that as
|
||||
@@ -1653,15 +1725,45 @@ function MemberRow({ member }: { member: GroupMemberHealth }) {
|
||||
* traffic, so it reads as a result with the address slot marked unknown — dim,
|
||||
* not red, and the LED stays green.
|
||||
*/
|
||||
function GroupTestReadout({ test, pending }: { test?: GroupTestResult; pending: boolean }) {
|
||||
/**
|
||||
* Errors that mean NO MEASUREMENT EXISTS, as opposed to "this target is broken".
|
||||
*
|
||||
* Three of the observatory's four failure reasons are about the observatory, not
|
||||
* about the path: nothing routes here, nothing has reached it yet, or background
|
||||
* probing is switched off. Painting those crit-red — which is what `ok:false`
|
||||
* used to buy you — reports a fault that nobody has found, on a target that may
|
||||
* be carrying traffic perfectly. Only "the observatory's probe through this path
|
||||
* failed" is a health finding, and it is deliberately NOT in this list.
|
||||
*
|
||||
* Matched on a stable fragment rather than the whole sentence, so a daemon that
|
||||
* rewords the tail still classifies. An error we don't recognise stays red: an
|
||||
* unknown failure is likelier to be real than not, and that is the safe default.
|
||||
*/
|
||||
const NO_MEASUREMENT = [
|
||||
'not routed by any enabled rule',
|
||||
'has not reached this target yet',
|
||||
'background probing is disabled',
|
||||
]
|
||||
const isAbsence = (err: string): boolean => NO_MEASUREMENT.some((frag) => err.includes(frag))
|
||||
|
||||
function GroupTestReadout({
|
||||
test,
|
||||
pending,
|
||||
hideAbsence,
|
||||
}: {
|
||||
test?: GroupTestResult
|
||||
pending: boolean
|
||||
/** The card already explains why nothing measures this target (the unused
|
||||
* note), so an absence error here would just say it a second time. */
|
||||
hideAbsence?: boolean
|
||||
}) {
|
||||
if (pending) {
|
||||
// "testing", never "measuring": the health run owns that word and covers every
|
||||
// group at once. Two runs that read the same on a card is how one group's test
|
||||
// came to look like all four were busy.
|
||||
// Names who is working and on what: the prober, on this target. The badge is
|
||||
// scoped to the cards the run covers, so it can say "this one" honestly.
|
||||
return (
|
||||
<div className="tg-test tg-test--wait" role="status">
|
||||
<Led variant="amber" pulse />
|
||||
<span className="tg-test-msg">testing this exit…</span>
|
||||
<span className="tg-test-msg">waiting for the prober to measure this…</span>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
@@ -1670,10 +1772,22 @@ function GroupTestReadout({ test, pending }: { test?: GroupTestResult; pending:
|
||||
const at = test.tested_unix ? fmtClock(test.tested_unix) : ''
|
||||
|
||||
if (!test.ok) {
|
||||
// No measurement exists. Unlit lamp, quiet text: this panel's way of saying
|
||||
// "no verdict", which is precisely the state — never a red one.
|
||||
if (isAbsence(test.error)) {
|
||||
if (hideAbsence) return null
|
||||
return (
|
||||
<div className="tg-test tg-test--none" role="status">
|
||||
<Led variant="off" />
|
||||
<span className="tg-test-msg">{test.error}</span>
|
||||
{at && <span className="tg-test-at mono">{at}</span>}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
return (
|
||||
<div className="tg-test tg-test--bad" role="status">
|
||||
<Led variant="crit" />
|
||||
<span className="tg-test-msg">{test.error || 'the test failed'}</span>
|
||||
<span className="tg-test-msg">{test.error || 'the probe failed'}</span>
|
||||
{at && <span className="tg-test-at mono">{at}</span>}
|
||||
</div>
|
||||
)
|
||||
@@ -2098,7 +2212,7 @@ function ChainRow({
|
||||
chain,
|
||||
busy,
|
||||
showHealth,
|
||||
used,
|
||||
health,
|
||||
test,
|
||||
testing,
|
||||
testBusy,
|
||||
@@ -2109,19 +2223,19 @@ function ChainRow({
|
||||
chain: Chain
|
||||
busy: boolean
|
||||
/** Group health checks are on (Settings). When false, the card drops its
|
||||
* exit-test readout and Test button — it is config only. */
|
||||
* health readout and Refresh button — it is config only. */
|
||||
showHealth: boolean
|
||||
/** This chain's reachability (GroupHealth.Used's chain analogue, plan §5.E).
|
||||
* undefined ⇒ the health endpoint hasn't reported this chain (not applied yet, or
|
||||
* a daemon version without chains): no badge. false ⇒ no enabled rule routes
|
||||
* through the chain, so the observatory never probes it and the card renders
|
||||
* "unused" instead of an exit-test readout. */
|
||||
used?: boolean
|
||||
/** Everything the observatory knows about this chain: whether any enabled rule
|
||||
* routes through it, and the per-hop measurements along it.
|
||||
* undefined ⇒ the health endpoint hasn't reported this chain at all (not
|
||||
* applied yet, or a daemon version without chains): the card says nothing
|
||||
* rather than guessing. */
|
||||
health?: ChainHealth
|
||||
test?: GroupTestResult
|
||||
/** An exit test covering THIS chain is in flight (the caller resolves it
|
||||
/** A refresh pass covering THIS chain is in flight (the caller resolves it
|
||||
* against the run's scope, exactly as for a group card). */
|
||||
testing: boolean
|
||||
/** Any exit test is in flight; the daemon runs one at a time. */
|
||||
/** Any refresh pass is in flight; the daemon runs one at a time. */
|
||||
testBusy: boolean
|
||||
onTest: () => void
|
||||
onEdit: () => void
|
||||
@@ -2165,25 +2279,21 @@ function ChainRow({
|
||||
{showHealth && (
|
||||
<>
|
||||
{/* A chain no enabled rule routes through is never probed (the
|
||||
observatory walks only reachable paths), so instead of an exit-test
|
||||
readout the card says so, quietly — the same "unused" pattern the
|
||||
group card uses (GroupHealthReadout), not a new design. `used` is
|
||||
undefined until the health endpoint reports this chain (or from a
|
||||
daemon version without chains): no badge then. */}
|
||||
{used === false && (
|
||||
<div className="gh gh--unused">
|
||||
<div className="gh-line">
|
||||
<span
|
||||
className="gh-unused"
|
||||
title="No enabled rule routes through this chain, so its exit is not probed. Add it to a rule to see health."
|
||||
>
|
||||
unused
|
||||
</span>
|
||||
<span className="gh-quiet">not probed — no enabled rule routes through this chain</span>
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
<GroupTestReadout test={test} pending={testing && !test} />
|
||||
observatory walks only reachable paths), so instead of a health
|
||||
readout the card says so — the same "unused" note the group card
|
||||
uses, not a new design. `health` is undefined until the endpoint
|
||||
reports this chain (or on a daemon without chains): say nothing
|
||||
then rather than guess. */}
|
||||
{health?.used === false ? (
|
||||
<NotRoutedNote kind="chain" />
|
||||
) : health?.used ? (
|
||||
<ChainHopRail chain={chain.Name} defs={hops} hops={health.hops} />
|
||||
) : null}
|
||||
<GroupTestReadout
|
||||
test={test}
|
||||
pending={testing && !test}
|
||||
hideAbsence={health?.used === false}
|
||||
/>
|
||||
</>
|
||||
)}
|
||||
</div>
|
||||
@@ -2194,13 +2304,210 @@ function ChainRow({
|
||||
editLabel={`Edit chain ${chain.Name}`}
|
||||
deleteLabel={`Delete chain ${chain.Name}`}
|
||||
onTest={showHealth ? onTest : undefined}
|
||||
testLabel={showHealth ? `Test the exit of chain ${chain.Name}` : undefined}
|
||||
testLabel={showHealth ? `Refresh the reading for chain ${chain.Name}` : undefined}
|
||||
testDisabled={testBusy}
|
||||
/>
|
||||
</li>
|
||||
)
|
||||
}
|
||||
|
||||
// ---- chain hop rail --------------------------------------------------------
|
||||
|
||||
/**
|
||||
* The API gives hops an index and no name. The page already knows the names — the
|
||||
* model's own `Hops` strings ("egress:ewan", "node:awgout", "group:sub0") — so
|
||||
* zip the two by POSITION.
|
||||
*
|
||||
* Two things make that safe rather than clever. A LEADING `egress:` is not a
|
||||
* numbered hop: the daemon lifts it into hop 1's entry detour, so it is dropped
|
||||
* before counting. And if the counts still disagree — a chain that splices
|
||||
* sub-chains gets FLATTENED by the daemon, producing more wire hops than the
|
||||
* config lists — every label is dropped. A hop labelled with its neighbour's name
|
||||
* is worse than a hop with no name at all: it would send someone to fix the wrong
|
||||
* target.
|
||||
*/
|
||||
function hopLabels(defs: string[], hops: ChainHopHealth[]): (string | undefined)[] {
|
||||
const numbered = defs.length > 0 && defs[0].startsWith('egress:') ? defs.slice(1) : defs
|
||||
if (numbered.length !== hops.length) return hops.map(() => undefined)
|
||||
return hops.map((h) => (h.index >= 1 && h.index <= numbered.length ? numbered[h.index - 1] : undefined))
|
||||
}
|
||||
|
||||
/**
|
||||
* One hop's lamp.
|
||||
*
|
||||
* `dead` is crit and `untested` is an UNLIT socket — never red, because nothing
|
||||
* has been measured and an unlit lamp is this panel's way of saying "no verdict".
|
||||
* The fourth case is the page's existing house reading, applied here for
|
||||
* consistency rather than invented: a group hop that is carrying traffic but has
|
||||
* confirmed failures on its board is amber. `state` stays the daemon's word for
|
||||
* "can this hop carry traffic"; the amber only qualifies HOW WELL.
|
||||
*/
|
||||
function hopLed(h: ChainHopHealth): LedVariant {
|
||||
if (h.state === 'dead') return 'crit'
|
||||
if (h.state === 'untested') return 'off'
|
||||
return h.dead > 0 ? 'amber' : 'on'
|
||||
}
|
||||
|
||||
/**
|
||||
* What the observatory measured at each position of a chain — the reading the
|
||||
* daemon always took and the panel never showed.
|
||||
*
|
||||
* THE DESIGN RISK, and the one place this card spends any boldness: the rail
|
||||
* draws the CONDUCTOR as well as the lamps, and severs it below the first dead
|
||||
* hop. A chain is a single series path, so the operator's real question is never
|
||||
* "how many hops are green" — it is "where does my traffic stop". Four lamps in a
|
||||
* column answer the first question and leave the second to arithmetic. A broken
|
||||
* conductor answers the second one before you have read a single word, which is
|
||||
* the whole reason this feature exists.
|
||||
*
|
||||
* It stays honest by keeping the two facts on two different marks. Each LAMP is
|
||||
* that hop's own measurement and never changes because of a hop in front of it —
|
||||
* a hop after the break that answers still shows green, because it really did
|
||||
* answer. The CONDUCTOR is reachability through the path, and that genuinely does
|
||||
* stop at the break. Nothing here is re-derived from the daemon's counters; the
|
||||
* only thing the panel adds is what a chain structurally is.
|
||||
*
|
||||
* Everything around the rail is deliberately quiet: no colour but the semantic
|
||||
* lamps, no motion at all, the orange accent untouched.
|
||||
*/
|
||||
function ChainHopRail({
|
||||
chain,
|
||||
defs,
|
||||
hops,
|
||||
}: {
|
||||
chain: string
|
||||
/** The chain's configured hops, straight off the model — the only source of names. */
|
||||
defs: string[]
|
||||
/** Absent ⇒ the engine never materialised per-hop outbounds. NOT "no hops". */
|
||||
hops?: ChainHopHealth[]
|
||||
}) {
|
||||
const ordered = useMemo(() => [...asArray(hops)].sort((a, b) => a.index - b.index), [hops])
|
||||
const labels = useMemo(() => hopLabels(defs, ordered), [defs, ordered])
|
||||
|
||||
// The first hop that was probed and did not answer. Everything after it is
|
||||
// unreachable THROUGH THIS CHAIN, whatever its own lamp says. `untested` is
|
||||
// never a break: nothing was measured, so nothing is known to be severed.
|
||||
const breakAt = ordered.findIndex((h) => h.state === 'dead')
|
||||
|
||||
if (ordered.length === 0) {
|
||||
// Say why, in one line, instead of an empty rail. The daemon collapses a
|
||||
// single-target chain into a plain alias and never builds copies to measure,
|
||||
// so we can tell the two absences apart from the config alone.
|
||||
const numbered = defs.filter((d, i) => !(i === 0 && d.startsWith('egress:')))
|
||||
return (
|
||||
<div className="gh gh--absent">
|
||||
<span className="gh-absent-msg">
|
||||
{numbered.length <= 1
|
||||
? 'This chain has a single hop, so the engine points traffic straight at that target instead of building a path to measure. Its health is on that target’s own card.'
|
||||
: 'The engine hasn’t built this chain’s hops yet, so there is nothing measured per hop. They appear once it is running with this config applied.'}
|
||||
</span>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
return (
|
||||
<div className="gh ch-rail">
|
||||
<span className="ch-eyebrow">measured, hop by hop</span>
|
||||
<ol className="ch-hops">
|
||||
{ordered.map((h, i) => {
|
||||
const label = labels[i]
|
||||
const severed = breakAt >= 0 && i > breakAt
|
||||
const cls = [
|
||||
'ch-hop',
|
||||
`ch-hop--${h.state}`,
|
||||
severed ? 'ch-hop--severed' : '',
|
||||
h.exit ? 'ch-hop--exit' : '',
|
||||
i === 0 ? 'ch-hop--first' : '',
|
||||
]
|
||||
.filter(Boolean)
|
||||
.join(' ')
|
||||
const age = fmtAge(h.age_seconds)
|
||||
return (
|
||||
<li key={h.tag || h.index} className={cls}>
|
||||
<span className="ch-num mono" aria-hidden="true">
|
||||
{h.index}
|
||||
</span>
|
||||
<span className="ch-socket">
|
||||
<Led variant={hopLed(h)} />
|
||||
</span>
|
||||
<div className="ch-body">
|
||||
<div className="ch-l1">
|
||||
<span className="ch-name mono" title={`engine outbound ${h.tag}`}>
|
||||
{label ?? (h.kind === 'group' ? 'a group hop' : 'a node hop')}
|
||||
</span>
|
||||
{h.exit && <span className="ch-tag">exit</span>}
|
||||
{h.state === 'alive' && h.delay_ms > 0 && (
|
||||
<span className="ch-delay mono">{h.delay_ms} ms</span>
|
||||
)}
|
||||
{h.state === 'untested' && <span className="ch-quiet">not measured yet</span>}
|
||||
{age && <span className="ch-age mono">{age}</span>}
|
||||
</div>
|
||||
|
||||
{/* A node hop IS its own measurement (total 1), so counters would
|
||||
only restate the lamp. A group hop rolls up its per-hop member
|
||||
copies, and those read exactly as they do everywhere else in
|
||||
this app: alive out of TESTED, with the untested remainder as a
|
||||
quiet aside only when there is one. */}
|
||||
{h.kind === 'group' && h.total > 0 && (
|
||||
<div className="ch-l2">
|
||||
{h.tested === 0 ? (
|
||||
<span className="ch-rest mono">
|
||||
{h.total} member{h.total === 1 ? '' : 's'}, none measured
|
||||
</span>
|
||||
) : (
|
||||
<>
|
||||
<span className="ch-count mono">
|
||||
<b>{h.alive}</b> / {h.tested}
|
||||
</span>
|
||||
<span className="ch-word">alive</span>
|
||||
{h.dead > 0 && (
|
||||
<span
|
||||
className="ch-dead mono"
|
||||
title={`${h.dead} member${h.dead === 1 ? '' : 's'} were probed at this hop and did not answer`}
|
||||
>
|
||||
{h.dead} not answering
|
||||
</span>
|
||||
)}
|
||||
{h.untested > 0 && (
|
||||
<span className="ch-rest mono">
|
||||
tested {h.tested} of {h.total}
|
||||
</span>
|
||||
)}
|
||||
</>
|
||||
)}
|
||||
{h.selected && (
|
||||
<span
|
||||
className="ch-now mono"
|
||||
title={`Traffic crossing hop ${h.index} of “${chain}” is on ${h.selected}`}
|
||||
>
|
||||
now → {h.selected}
|
||||
</span>
|
||||
)}
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
</li>
|
||||
)
|
||||
})}
|
||||
</ol>
|
||||
|
||||
{/* The sentence the rail's shape implies, written out — because the break is
|
||||
the answer someone came here for, and a graphic alone should never be the
|
||||
only place a finding exists. */}
|
||||
{breakAt >= 0 && (
|
||||
<p className="gh-say gh-say--bad">
|
||||
Hop {ordered[breakAt].index}
|
||||
{labels[breakAt] ? ` (${labels[breakAt]})` : ''} was probed and did not answer. A chain is
|
||||
one path, so traffic stops there
|
||||
{breakAt < ordered.length - 1
|
||||
? ' — the hops after it answer on their own, but nothing reaches them through this chain.'
|
||||
: '.'}
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
function ChainEditor({
|
||||
initial,
|
||||
hopOptions,
|
||||
@@ -2675,8 +2982,8 @@ function RowActions({
|
||||
busy: boolean
|
||||
editLabel: string
|
||||
deleteLabel: string
|
||||
// Only groups and chains can be tested, so the control is optional and absent
|
||||
// everywhere else rather than a disabled stub on every row.
|
||||
// Only groups and chains are probed, so the refresh control is optional and
|
||||
// absent everywhere else rather than a disabled stub on every row.
|
||||
onTest?: () => void
|
||||
testLabel?: string
|
||||
testDisabled?: boolean
|
||||
@@ -2689,8 +2996,9 @@ function RowActions({
|
||||
onClick={onTest}
|
||||
disabled={busy || testDisabled}
|
||||
aria-label={testLabel}
|
||||
title="Ask the background prober to measure this target out of turn. It does not open a connection from the panel."
|
||||
>
|
||||
Test
|
||||
Refresh
|
||||
</Button>
|
||||
)}
|
||||
<Button className="tg-act" onClick={onEdit} disabled={busy} aria-label={editLabel}>
|
||||
|
||||
@@ -0,0 +1,150 @@
|
||||
// protectionState — the one sentence the whole panel shows about "am I protected".
|
||||
//
|
||||
// Run with `npm test` (node's built-in test runner + native TypeScript stripping;
|
||||
// no test dependency is added to the SPA, which ships inside the daemon binary).
|
||||
//
|
||||
// The case this file was written for is "plane full, traffic direct": the exact
|
||||
// state of a live router — one enabled rule, `default → direct`, no groups, no
|
||||
// rule-sets — where every part of the data plane was installed and the readout
|
||||
// therefore said "Protected — traffic from your network is going through the
|
||||
// tunnel", under a green LED, while the whole LAN went out the plain WAN.
|
||||
//
|
||||
// planeState.ts has no runtime imports (both of its imports are `import type`),
|
||||
// so this runs against the real module with nothing stubbed.
|
||||
|
||||
import { test } from 'node:test'
|
||||
import assert from 'node:assert/strict'
|
||||
|
||||
import { protectionState } from './planeState.ts'
|
||||
import type { Status, Traffic } from './api.ts'
|
||||
|
||||
/** A healthy, fully-installed router; `traffic` is what each case varies. */
|
||||
function status(over: Partial<Status> = {}): Status {
|
||||
return {
|
||||
running: true,
|
||||
enabled: true,
|
||||
active: true,
|
||||
table: true,
|
||||
hash: 'abc',
|
||||
version: '1.11.0-shater',
|
||||
kill_switch: 'closed',
|
||||
engine_running: true,
|
||||
plane: 'full',
|
||||
warnings: [],
|
||||
...over,
|
||||
}
|
||||
}
|
||||
|
||||
function withTraffic(traffic: Traffic | undefined): Status {
|
||||
return status({ traffic })
|
||||
}
|
||||
|
||||
// --- the field case ---------------------------------------------------------
|
||||
|
||||
test('plane full + default direct is NOT reported as protected', () => {
|
||||
const s = protectionState(withTraffic({ verdict: 'direct', default: 'direct', tunnel_rules: 0 }))
|
||||
assert.notEqual(s.headline, 'Protected')
|
||||
assert.equal(s.variant, 'crit')
|
||||
assert.equal(s.alarm, true)
|
||||
// The claim that was false must not survive anywhere in the copy.
|
||||
assert.doesNotMatch(s.detail, /going through the tunnel/)
|
||||
// ...and the honest consequence must be stated, not implied.
|
||||
assert.match(s.detail, /real address/)
|
||||
})
|
||||
|
||||
// --- the other verdicts under a full plane ----------------------------------
|
||||
|
||||
test('plane full + default into a tunnel is protected', () => {
|
||||
const s = protectionState(withTraffic({ verdict: 'tunnel', default: 'auto', tunnel_rules: 1 }))
|
||||
assert.equal(s.variant, 'on')
|
||||
assert.equal(s.headline, 'Protected')
|
||||
assert.equal(s.alarm, false)
|
||||
})
|
||||
|
||||
test('plane full + direct default with tunnelling rules is split, not protected', () => {
|
||||
const s = protectionState(withTraffic({ verdict: 'split', default: 'direct', tunnel_rules: 3 }))
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.notEqual(s.headline, 'Protected')
|
||||
// Says how much is protected, and that the default is not.
|
||||
assert.match(s.detail, /3 rules/)
|
||||
assert.match(s.detail, /normal internet connection/)
|
||||
// A working selective setup must not raise a banner on every other page.
|
||||
assert.equal(s.alarm, false)
|
||||
})
|
||||
|
||||
test('split names a single rule in the singular', () => {
|
||||
const s = protectionState(withTraffic({ verdict: 'split', default: 'direct', tunnel_rules: 1 }))
|
||||
assert.match(s.detail, /^One rule sends traffic/)
|
||||
})
|
||||
|
||||
test('plane full + blocked default with rules leaks nothing and is never crit', () => {
|
||||
const s = protectionState(withTraffic({ verdict: 'blocked', default: 'block', tunnel_rules: 2 }))
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.equal(s.alarm, false)
|
||||
assert.match(s.detail, /nothing is leaving unprotected/)
|
||||
})
|
||||
|
||||
test('plane full + blocked default with no rules says the network has no way out', () => {
|
||||
const s = protectionState(withTraffic({ verdict: 'blocked', default: 'block', tunnel_rules: 0 }))
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.equal(s.alarm, true)
|
||||
assert.doesNotMatch(s.detail, /going through the tunnel/)
|
||||
})
|
||||
|
||||
test('plane full with no verdict claims nothing either way', () => {
|
||||
for (const t of [undefined, { verdict: '' as const }]) {
|
||||
const s = protectionState(withTraffic(t))
|
||||
assert.notEqual(s.headline, 'Protected')
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.equal(s.alarm, false)
|
||||
}
|
||||
})
|
||||
|
||||
// --- the branches that were already correct ---------------------------------
|
||||
|
||||
test('no status yet', () => {
|
||||
const s = protectionState(null)
|
||||
assert.equal(s.variant, 'off')
|
||||
assert.equal(s.alarm, false)
|
||||
})
|
||||
|
||||
test('service switched off is a deliberate state, not a fault', () => {
|
||||
const s = protectionState(status({ enabled: false }))
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.equal(s.headline, 'Turned off')
|
||||
assert.equal(s.alarm, false)
|
||||
})
|
||||
|
||||
test('hold: the kill-switch caught it — protected, offline', () => {
|
||||
const s = protectionState(status({ plane: 'hold', engine_running: false, active: false }))
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.equal(s.alarm, true)
|
||||
assert.match(s.headline, /blocked/)
|
||||
})
|
||||
|
||||
test('none + fail-closed is the leak, and it is crit', () => {
|
||||
const s = protectionState(status({ plane: 'none', table: false, engine_running: false }))
|
||||
assert.equal(s.variant, 'crit')
|
||||
assert.equal(s.alarm, true)
|
||||
})
|
||||
|
||||
test('none + fail-open is the operator’s documented choice, stated not alarmed at', () => {
|
||||
const s = protectionState(
|
||||
status({ plane: 'none', table: false, engine_running: false, kill_switch: 'open' }),
|
||||
)
|
||||
assert.equal(s.variant, 'amber')
|
||||
assert.equal(s.alarm, true)
|
||||
})
|
||||
|
||||
test('daemon too old to send `plane` keeps its own fallback', () => {
|
||||
// Nothing here may depend on `traffic`: a daemon with no `plane` has no
|
||||
// `traffic` either, and this branch reads what it can observe instead.
|
||||
const { plane, ...noPlane } = status()
|
||||
void plane
|
||||
assert.equal(protectionState(noPlane as Status).headline, 'Protected')
|
||||
assert.equal(protectionState({ ...noPlane, running: false } as Status).headline, 'Service stopped')
|
||||
assert.equal(
|
||||
protectionState({ ...noPlane, active: false } as Status).headline,
|
||||
'Starting up',
|
||||
)
|
||||
})
|
||||
+104
-7
@@ -11,7 +11,7 @@
|
||||
// same router differently.
|
||||
|
||||
import type { LedVariant } from './components'
|
||||
import type { Status } from './api'
|
||||
import type { Status, Traffic } from './api'
|
||||
|
||||
export interface ProtectionState {
|
||||
variant: LedVariant
|
||||
@@ -31,6 +31,24 @@ export interface ProtectionState {
|
||||
* hold — the kill-switch caught it. Protected, but offline.
|
||||
* none (fail-closed) — there is no protection at all. Online, and exposed.
|
||||
* Collapsing them would erase the only difference that matters.
|
||||
*
|
||||
* PLANE IS NOT THE WHOLE ANSWER, AND THAT USED TO BE A LIE. `plane: 'full'`
|
||||
* returned "Protected — traffic from your network is going through the tunnel",
|
||||
* which is a claim `plane` cannot support: it only says the nft table, the policy
|
||||
* routing and the engine are all installed. Where the diverted packets go once the
|
||||
* engine has them is decided by the engine's default route, and a router in the
|
||||
* field ran with one rule — `default → direct`, no groups, no rule-sets. Fully
|
||||
* installed plane, zero tunnel, whole LAN out the plain WAN with its real address,
|
||||
* green LED, "Protected".
|
||||
*
|
||||
* So `full` now branches on `status.traffic`, the daemon's verdict on the config
|
||||
* it is actually running (see api.ts TrafficVerdict). It is computed on the daemon
|
||||
* because only the daemon knows what was GENERATED and STARTED: the panel's
|
||||
* /api/config is desired state, which diverges from the running one whenever edits
|
||||
* are unapplied or a rollback is pending, and re-deriving the default route from it
|
||||
* would mean a second implementation of the generator's rule loop — schedules,
|
||||
* shadowed catch-alls, targets that failed to resolve and fell back to direct —
|
||||
* drifting against the first.
|
||||
*/
|
||||
export function protectionState(status: Status | null): ProtectionState {
|
||||
if (!status) {
|
||||
@@ -56,12 +74,7 @@ export function protectionState(status: Status | null): ProtectionState {
|
||||
|
||||
switch (status.plane) {
|
||||
case 'full':
|
||||
return {
|
||||
variant: 'on',
|
||||
headline: 'Protected',
|
||||
detail: 'Traffic from your network is going through the tunnel.',
|
||||
alarm: false,
|
||||
}
|
||||
return fullPlaneState(status.traffic)
|
||||
case 'hold':
|
||||
return {
|
||||
variant: 'amber',
|
||||
@@ -112,3 +125,87 @@ export function protectionState(status: Status | null): ProtectionState {
|
||||
alarm: false,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The plane is fully installed — now say where the traffic it carries ends up.
|
||||
*
|
||||
* Only `tunnel` earns "Protected". The other verdicts each describe a real router
|
||||
* someone can be sitting in front of, and they are kept apart because the thing to
|
||||
* DO about them differs:
|
||||
*
|
||||
* split — deliberate for most people who reach it, accidental for the rest
|
||||
* (a default rule that was never pointed anywhere). Amber, but no
|
||||
* alarm: raising a banner on every page of a working selective setup
|
||||
* is how a banner stops being read.
|
||||
* direct — the engine is running and forwarding every connection out the plain
|
||||
* WAN. The traffic outcome is identical to `plane: 'none'` under a
|
||||
* closed kill-switch, so it gets the same weight: crit, and it
|
||||
* interrupts. The wording differs because the fix does — nothing
|
||||
* failed here, the routing simply says "direct".
|
||||
* blocked — the fail-closed default. Nothing is leaking, so this is never crit;
|
||||
* with no tunnelling rules at all it means the network has no way out
|
||||
* and someone should be told why.
|
||||
* unknown — an older daemon, or the seconds between this daemon starting and its
|
||||
* first apply. We do not know, so we do not claim. Saying "Protected"
|
||||
* here is the exact bug being removed.
|
||||
*/
|
||||
function fullPlaneState(traffic: Traffic | undefined): ProtectionState {
|
||||
const tunnelRules = traffic?.tunnel_rules ?? 0
|
||||
|
||||
switch (traffic?.verdict) {
|
||||
case 'tunnel':
|
||||
return {
|
||||
variant: 'on',
|
||||
headline: 'Protected',
|
||||
detail: 'Traffic from your network is going through the tunnel.',
|
||||
alarm: false,
|
||||
}
|
||||
case 'split':
|
||||
return {
|
||||
variant: 'amber',
|
||||
headline: 'Partly protected — the rest goes out directly',
|
||||
detail: `${ruleCount(tunnelRules)} through the tunnel. Everything they don’t match leaves through your normal internet connection, with your real address.`,
|
||||
alarm: false,
|
||||
}
|
||||
case 'direct':
|
||||
return {
|
||||
variant: 'crit',
|
||||
headline: 'Not protected — nothing is going through the tunnel',
|
||||
detail:
|
||||
'The service is running, but your routing sends every connection straight out your normal internet connection, with your real address. On the Routing page, point the default rule at a group or a node.',
|
||||
alarm: true,
|
||||
}
|
||||
case 'blocked':
|
||||
return tunnelRules > 0
|
||||
? {
|
||||
variant: 'amber',
|
||||
headline: 'Partly protected — everything else is blocked',
|
||||
detail: `${ruleCount(tunnelRules)} through the tunnel. Anything they don’t match is blocked instead of being let out, so nothing is leaving unprotected.`,
|
||||
alarm: false,
|
||||
}
|
||||
: {
|
||||
variant: 'amber',
|
||||
headline: 'Nothing is getting out',
|
||||
detail:
|
||||
'No rule sends traffic anywhere, so every connection from your network is being blocked rather than let out unprotected. Add a default rule on the Routing page.',
|
||||
alarm: true,
|
||||
}
|
||||
default:
|
||||
return {
|
||||
variant: 'amber',
|
||||
headline: 'Checking where traffic goes',
|
||||
detail:
|
||||
'The router is up and handling your traffic. It hasn’t reported yet whether that traffic is going through the tunnel.',
|
||||
alarm: false,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** "One rule sends traffic" / "4 rules send traffic", so the detail lines above
|
||||
* can name a number the operator can go and count on the Routing page. Falls
|
||||
* back to the vague form only if the daemon sent a verdict without a count. */
|
||||
function ruleCount(n: number): string {
|
||||
if (n <= 0) return 'Some traffic goes'
|
||||
if (n === 1) return 'One rule sends traffic'
|
||||
return `${n} rules send traffic`
|
||||
}
|
||||
|
||||
+6
-1
@@ -18,5 +18,10 @@
|
||||
"noFallthroughCasesInSwitch": true,
|
||||
"forceConsistentCasingInFileNames": true
|
||||
},
|
||||
"include": ["src", "vite.config.ts"]
|
||||
"include": ["src", "vite.config.ts"],
|
||||
// *.test.ts runs under node's built-in test runner (`npm test`), which strips
|
||||
// types rather than checking them. They are excluded here because they import
|
||||
// node:test / node:assert, and the SPA deliberately carries no @types/node — it
|
||||
// is embedded in the daemon binary, so every devDependency is weight on a router.
|
||||
"exclude": ["src/**/*.test.ts"]
|
||||
}
|
||||
|
||||
@@ -44,6 +44,13 @@ type URLTest struct {
|
||||
group *URLTestGroup
|
||||
interruptExternalConnections bool
|
||||
balancer *balancer // lx: SPEC 019 — nil for least_test (default)
|
||||
// lx: health board §5.C — true when options.SelfCheck == false: the group's
|
||||
// OWN probing schedule (PostStart warm-up + Touch ticker) is stood down and
|
||||
// the observatory is the only thing that measures its members. Stored
|
||||
// INVERTED so the zero value keeps today's behaviour for every construction
|
||||
// path that does not go through NewURLTest (hand-built groups in tests).
|
||||
// See option.URLTestOutboundOptions.SelfCheck for the full reasoning.
|
||||
selfCheckDisabled bool
|
||||
}
|
||||
|
||||
func NewURLTest(ctx context.Context, router adapter.Router, logger log.ContextLogger, tag string, options option.URLTestOutboundOptions) (adapter.Outbound, error) {
|
||||
@@ -71,6 +78,9 @@ func NewURLTest(ctx context.Context, router adapter.Router, logger log.ContextLo
|
||||
idleTimeout: time.Duration(options.IdleTimeout),
|
||||
interruptExternalConnections: options.InterruptExistConnections,
|
||||
balancer: balancer,
|
||||
// nil/absent means true (self-check on) — the documented default, so a
|
||||
// config written before the flag existed behaves exactly as it always has.
|
||||
selfCheckDisabled: options.SelfCheck != nil && !*options.SelfCheck,
|
||||
}
|
||||
if len(outbound.tags) == 0 {
|
||||
return nil, E.New("missing tags")
|
||||
@@ -92,6 +102,9 @@ func (s *URLTest) Start() error {
|
||||
return err
|
||||
}
|
||||
group.balancer = s.balancer // lx: SPEC 019 v2 — health-check drives the pool through it
|
||||
// lx: health board §5.C — carry the stand-down flag onto the group the same
|
||||
// way the balancer travels: set after construction, immutable from then on.
|
||||
group.selfCheckDisabled = s.selfCheckDisabled
|
||||
if s.balancer != nil {
|
||||
// lx: health board §5.B — slot liveness reads through the board verdict, so a
|
||||
// death recorded by any prober or a failed dial takes effect on the next pick,
|
||||
@@ -319,6 +332,12 @@ type URLTestGroup struct {
|
||||
lastActive common.TypedValue[time.Time]
|
||||
lastSelected common.TypedValue[string] // lx: SPEC 019 — Now() in balanced modes
|
||||
balancer *balancer // lx: SPEC 019 v2 — round_robin pool; nil for least_test
|
||||
// lx: health board §5.C — mirrors URLTest.selfCheckDisabled (set by Start,
|
||||
// immutable afterwards, zero value = probing on). Guards ONLY the group's
|
||||
// own schedule: the PostStart warm-up sweep and the Touch ticker. An
|
||||
// explicit CheckOutbounds/URLTest call is untouched — the flag stands down
|
||||
// the schedule, not the capability.
|
||||
selfCheckDisabled bool
|
||||
}
|
||||
|
||||
func NewURLTestGroup(ctx context.Context, outboundManager adapter.OutboundManager, logger log.Logger, outbounds []adapter.Outbound, link string, interval time.Duration, tolerance uint16, idleTimeout time.Duration, interruptExternalConnections bool) (*URLTestGroup, error) {
|
||||
@@ -362,14 +381,35 @@ func (g *URLTestGroup) PostStart() {
|
||||
g.lastActive.Store(time.Now())
|
||||
// lx: SPEC 019 v2 — seed the pool so round_robin can route from the first connection,
|
||||
// before the first health-check completes (history-warm nodes first, else config order).
|
||||
// The seed only READS the board, so it runs even with the self-check stood down.
|
||||
g.seedPool()
|
||||
go g.CheckOutbounds(false)
|
||||
// lx: health board §5.C — the warm-up sweep is the first half of the group's
|
||||
// own probing schedule, and it fires for EVERY group at box start, including
|
||||
// groups no routing rule reaches. For those, the sweep dials every member
|
||||
// directly from the router — a path nothing uses — and records the outcome
|
||||
// under the members' base tags, forging the board reading the observatory
|
||||
// exists to keep honest. A stood-down group therefore skips it entirely; the
|
||||
// observatory (or nothing, for a truly unused group) is what measures its
|
||||
// members.
|
||||
if !g.selfCheckDisabled {
|
||||
go g.CheckOutbounds(false)
|
||||
}
|
||||
}
|
||||
|
||||
func (g *URLTestGroup) Touch() {
|
||||
if !g.started {
|
||||
return
|
||||
}
|
||||
// lx: health board §5.C — Touch's only job is to keep the group's OWN
|
||||
// probing ticker alive while traffic flows. With the self-check stood down
|
||||
// there is deliberately no ticker to start or feed: the observatory owns the
|
||||
// schedule, and a stray dial through an unused group (a stale rule cache, a
|
||||
// manual pin) must not arm 30 minutes of direct probing under the members'
|
||||
// base tags. Checked before the lock because the flag is immutable after
|
||||
// Start, exactly like the started fast-path above.
|
||||
if g.selfCheckDisabled {
|
||||
return
|
||||
}
|
||||
g.access.Lock()
|
||||
defer g.access.Unlock()
|
||||
if g.ticker != nil {
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
package group
|
||||
|
||||
// lx: health board §5.C tests — SelfCheck stands the group's OWN probing
|
||||
// schedule down: no PostStart warm-up sweep, no Touch ticker. The explicit
|
||||
// CheckOutbounds path stays available, and the nil default keeps probing.
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/sagernet/sing-box/common/urltest"
|
||||
"github.com/sagernet/sing-box/log"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
)
|
||||
|
||||
// waitForHistory polls until the store holds an entry for tag or the deadline
|
||||
// passes; reports whether it appeared. PostStart's sweep runs on its own
|
||||
// goroutine, so both directions of the assertion need a bounded wait.
|
||||
func waitForHistory(hist *urltest.HistoryStorage, tag string, deadline time.Duration) bool {
|
||||
stop := time.Now().Add(deadline)
|
||||
for time.Now().Before(stop) {
|
||||
if hist.LoadURLTestHistory(tag) != nil {
|
||||
return true
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// A group with the self-check stood down writes NOTHING to the history storage
|
||||
// on PostStart: the warm-up sweep — which would dial the member directly from
|
||||
// the router and mark the failure under its base tag — must not fire. And
|
||||
// Touch, the other half of the schedule, must not start a ticker either.
|
||||
func TestSelfCheckDisabledPostStartWritesNothing(t *testing.T) {
|
||||
hist := urltest.NewHistoryStorage()
|
||||
a := &healthNode{tag: "a", fail: true}
|
||||
manager := managerOf(a)
|
||||
g := healthTestGroup(hist, manager, a)
|
||||
g.selfCheckDisabled = true
|
||||
|
||||
g.PostStart()
|
||||
// The absence of a write is the assertion, so give the (non-existent) sweep
|
||||
// real time to have happened before declaring victory.
|
||||
if waitForHistory(hist, "a", 150*time.Millisecond) {
|
||||
t.Fatal("a stood-down group's PostStart wrote to the board; the warm-up sweep must not fire")
|
||||
}
|
||||
|
||||
g.Touch()
|
||||
g.access.Lock()
|
||||
ticker := g.ticker
|
||||
g.access.Unlock()
|
||||
if ticker != nil {
|
||||
t.Fatal("Touch armed the probing ticker on a stood-down group")
|
||||
}
|
||||
}
|
||||
|
||||
// The default (SelfCheck nil, i.e. the zero-value field on a hand-built group)
|
||||
// keeps today's behaviour: PostStart's warm-up sweep runs and records the
|
||||
// failing member on the board.
|
||||
func TestSelfCheckDefaultStillProbesOnPostStart(t *testing.T) {
|
||||
hist := urltest.NewHistoryStorage()
|
||||
a := &healthNode{tag: "a", fail: true}
|
||||
manager := managerOf(a)
|
||||
g := healthTestGroup(hist, manager, a)
|
||||
|
||||
g.PostStart()
|
||||
if !waitForHistory(hist, "a", 5*time.Second) {
|
||||
t.Fatal("default group's PostStart never probed; the self-check must stay on unless stood down")
|
||||
}
|
||||
if v := hist.Verdict("a", 10*time.Minute); v != urltest.VerdictDead {
|
||||
t.Fatalf("verdict(a) = %v, want dead from the warm-up sweep", v)
|
||||
}
|
||||
}
|
||||
|
||||
// An EXPLICIT CheckOutbounds still probes a stood-down group: the flag
|
||||
// suppresses the group's own schedule, never a deliberate request (the adapter
|
||||
// interface a human or an API invokes on purpose).
|
||||
func TestSelfCheckDisabledExplicitCheckStillProbes(t *testing.T) {
|
||||
hist := urltest.NewHistoryStorage()
|
||||
a := &healthNode{tag: "a", fail: true}
|
||||
manager := managerOf(a)
|
||||
g := healthTestGroup(hist, manager, a)
|
||||
g.selfCheckDisabled = true
|
||||
|
||||
g.CheckOutbounds(true)
|
||||
if hist.LoadURLTestHistory("a") == nil {
|
||||
t.Fatal("an explicit CheckOutbounds(true) did not probe; the flag must only stand down the schedule")
|
||||
}
|
||||
}
|
||||
|
||||
// The option → outbound plumbing: nil/absent means on, an explicit false means
|
||||
// stood down, an explicit true means on. NewURLTest is the only place the
|
||||
// option is read, so this is where a plumbing regression would hide.
|
||||
func TestSelfCheckOptionPlumbing(t *testing.T) {
|
||||
build := func(selfCheck *bool) *URLTest {
|
||||
t.Helper()
|
||||
opts := option.URLTestOutboundOptions{Outbounds: []string{"a"}}
|
||||
opts.SelfCheck = selfCheck
|
||||
ob, err := NewURLTest(context.Background(), nil, log.NewNOPFactory().Logger(), "t", opts)
|
||||
if err != nil {
|
||||
t.Fatalf("NewURLTest: %v", err)
|
||||
}
|
||||
return ob.(*URLTest)
|
||||
}
|
||||
if build(nil).selfCheckDisabled {
|
||||
t.Fatal("nil SelfCheck must keep the self-check ON (the compatibility default)")
|
||||
}
|
||||
on, off := true, false
|
||||
if build(&on).selfCheckDisabled {
|
||||
t.Fatal("SelfCheck=true must keep the self-check on")
|
||||
}
|
||||
if !build(&off).selfCheckDisabled {
|
||||
t.Fatal("SelfCheck=false must stand the self-check down")
|
||||
}
|
||||
}
|
||||
+27
-13
@@ -18,9 +18,10 @@
|
||||
# OpenWrt package can $(INSTALL_BIN) the arch-matched artifact.
|
||||
# 5. Prints a size table + a per-arch static check (ELF type / no PT_INTERP).
|
||||
#
|
||||
# Router build tag set = D9 (musl-static). We deliberately DROP with_purego and
|
||||
# with_naive_outbound: they pull cronet-go, which forces a glibc PT_INTERP even
|
||||
# with CGO_ENABLED=0, making the binary unusable on musl OpenWrt.
|
||||
# Router build tag set = D9/D23 (musl-static). It is DEFINED IN, and only in,
|
||||
# scripts/router-tags.sh (sourced below) — that file documents every tag and is
|
||||
# machine-checked against the declared feature list by shater/buildtags's test.
|
||||
# Run scripts/check-router-tags.sh after touching it.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/build-shaterd.sh [VERSION] [--fast]
|
||||
@@ -82,16 +83,15 @@ fi
|
||||
[ -n "$VERSION" ] || VERSION="v0.2.0-dev"
|
||||
|
||||
# --- config -----------------------------------------------------------------
|
||||
# D9 router tag set (musl-static). Keep in sync with docs-shater/DECISIONS.md D9.
|
||||
# No with_gvisor: the data plane is tproxy/redirect (netplane), generate never
|
||||
# emits a tun inbound, so the userspace gvisor stack was 3.6 MB of dead weight
|
||||
# (tun would fall back to the system stack anyway).
|
||||
# No with_clash_api: the panel is shater's own; generate never emits a clash_api
|
||||
# service ("the shater generator emits none of those" — shater/engine/engine.go).
|
||||
# No with_dhcp: shater resolvers are udp/tcp/doh/dot/local/fakeip — no "dhcp://"
|
||||
# DNS transport is ever generated, and the slim registry never registers it.
|
||||
ROUTER_TAGS="with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command"
|
||||
LDFLAGS="-X github.com/sagernet/sing-box/constant.Version=${VERSION} -checklinkname=0 -s -w -buildid="
|
||||
# D9/D23 router tag set (musl-static). The set itself lives in ONE place —
|
||||
# scripts/router-tags.sh — because it is also parsed by shater/buildtags's test,
|
||||
# which proves it still covers every feature docs-shater/FEATURES.md declares.
|
||||
# Do not re-inline it here: that split is exactly how `with_gvisor` went missing
|
||||
# while `with_wireguard` stayed (D23).
|
||||
# shellcheck source=router-tags.sh
|
||||
. "$SCRIPT_DIR/router-tags.sh"
|
||||
ROUTER_TAGS="$SHATER_ROUTER_TAGS"
|
||||
LDFLAGS="-X github.com/sagernet/sing-box/constant.Version=${VERSION} ${SHATER_ROUTER_LDFLAGS} -s -w -buildid="
|
||||
|
||||
UPX_BIN="${UPX:-upx}"
|
||||
# UPX itself treats the environment variable UPX as extra command-line options, so
|
||||
@@ -111,6 +111,20 @@ echo " version : $VERSION"
|
||||
echo " tags : $ROUTER_TAGS"
|
||||
echo " upx : $UPX_BIN"
|
||||
echo " go : $(go version)"
|
||||
|
||||
# --- gate: does this tag set still support what we declare? (D23) ------------
|
||||
# Cheap (one tiny tag-less package, no network, ~1 s) and it travels with the
|
||||
# BUILD rather than with a CI config, so an artifact produced by hand on a
|
||||
# developer's machine gets the same guarantee. The heavier half — actually
|
||||
# constructing every declared protocol under these tags — is
|
||||
# scripts/check-router-tags.sh, which CI runs before this script.
|
||||
if ! (cd "$REPO" && go test -count=1 ./shater/buildtags/ >/dev/null); then
|
||||
echo >&2
|
||||
echo " ABORT: the router tag set no longer covers a declared feature." >&2
|
||||
echo " Details: go test ./shater/buildtags/" >&2
|
||||
echo " Full check: scripts/check-router-tags.sh" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo
|
||||
|
||||
# --- step 1: build the SPA --------------------------------------------------
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# check-router-tags.sh — prove the SHIPPED build-tag set still supports every
|
||||
# feature shater declares (D23).
|
||||
#
|
||||
# WHY (2026-07-25): the router tag set is a trimmed subset of upstream's, but the
|
||||
# test suite builds with the FULL upstream set — so the one combination we
|
||||
# actually ship was never exercised. `with_gvisor` got trimmed while
|
||||
# `with_wireguard` stayed, and every shipped binary answered a WireGuard node
|
||||
# with "gVisor is not included in this build". Compiling is not evidence.
|
||||
#
|
||||
# WHAT IT RUNS
|
||||
# 1. shater/buildtags, TAG-LESS — reads scripts/router-tags.sh and fails if a
|
||||
# declared feature (buildtags.Features) lost a build tag it needs. Cheap,
|
||||
# hostable anywhere, catches the trim at the moment it happens.
|
||||
# 2. shater/generate + shater/buildtags, WITH THE SHIPPED TAG SET on linux —
|
||||
# constructs one node of every declared protocol through box.New+Start, and
|
||||
# cross-checks that the tag detectors match the set the compiler was given.
|
||||
# This is the half that catches "the tag is there but insufficient".
|
||||
#
|
||||
# The run is unprivileged (no tproxy inbound is built) and offline apart from Go
|
||||
# module downloads.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/check-router-tags.sh
|
||||
#
|
||||
# Env:
|
||||
# SHATER_GO_IMAGE docker image used to reach linux from a non-linux host
|
||||
# (default golang:1.26 — keep it >= go.mod's toolchain).
|
||||
# SHATER_NO_DOCKER=1 fail instead of falling back to docker.
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
cd "$REPO"
|
||||
|
||||
# shellcheck source=router-tags.sh
|
||||
. "$SCRIPT_DIR/router-tags.sh"
|
||||
|
||||
echo "== router build-tag check =="
|
||||
echo " tags : $SHATER_ROUTER_TAGS"
|
||||
echo " ldflags: $SHATER_ROUTER_LDFLAGS"
|
||||
echo
|
||||
|
||||
# --- 1. static: does the set still cover the declared features? --------------
|
||||
# No tags, no OS constraint: this is the check that would have caught the outage
|
||||
# on the developer's own machine.
|
||||
echo "== [1/2] declared features vs. the shipped tag set (no tags needed) =="
|
||||
go test -count=1 ./shater/buildtags/
|
||||
echo
|
||||
|
||||
# --- 2. behavioural: does the shipped combination actually construct? --------
|
||||
# The protocol-construction test is linux-only (box.New validates the loop-guard
|
||||
# routing_mark only there). From a non-linux host, re-exec this script inside a
|
||||
# golang container rather than silently skipping — a skipped guard is no guard.
|
||||
if [ "$(go env GOOS)" != "linux" ] && [ "${SHATER_TAGCHECK_IN_DOCKER:-0}" != "1" ]; then
|
||||
if [ "${SHATER_NO_DOCKER:-0}" = "1" ] || ! command -v docker >/dev/null 2>&1; then
|
||||
echo " ERROR: step 2 needs linux (GOOS=$(go env GOOS)) and docker is unavailable/disabled." >&2
|
||||
echo " Run this script on the linux CI runner or the OpenWrt VM." >&2
|
||||
exit 1
|
||||
fi
|
||||
image="${SHATER_GO_IMAGE:-golang:1.26}"
|
||||
echo "== [2/2] re-exec on linux via docker ($image) =="
|
||||
host_repo="$REPO"
|
||||
command -v cygpath >/dev/null 2>&1 && host_repo="$(cygpath -w "$REPO")"
|
||||
# Named volumes keep the module/build cache warm between runs; MSYS2_ARG_CONV_EXCL
|
||||
# stops Git Bash from rewriting the container-side paths into windows ones.
|
||||
MSYS2_ARG_CONV_EXCL='*' MSYS_NO_PATHCONV=1 docker run --rm \
|
||||
-v "$host_repo":/src \
|
||||
-v shater-tagcheck-gomod:/go/pkg/mod \
|
||||
-v shater-tagcheck-gocache:/root/.cache/go-build \
|
||||
-w /src \
|
||||
-e SHATER_TAGCHECK_IN_DOCKER=1 \
|
||||
"$image" bash scripts/check-router-tags.sh
|
||||
exit $?
|
||||
fi
|
||||
|
||||
echo "== [2/2] every declared protocol constructs under the SHIPPED tags =="
|
||||
out="$(mktemp)"
|
||||
trap 'rm -f "$out"' EXIT
|
||||
set +e
|
||||
SHATER_ROUTER_TAG_CHECK=1 go test -count=1 -v \
|
||||
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
|
||||
-run 'TestShippedTagSetConstructsDeclaredProtocols|TestEveryTagGatedFeatureIsProbed' \
|
||||
./shater/generate/ >"$out" 2>&1
|
||||
rc_gen=$?
|
||||
set -e
|
||||
sed 's/^/ /' "$out"
|
||||
|
||||
set +e
|
||||
SHATER_ROUTER_TAG_CHECK=1 go test -count=1 -v \
|
||||
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
|
||||
-run 'TestCompiledTagsMatchTheShippedSet' \
|
||||
./shater/buildtags/ >"$out" 2>&1
|
||||
rc_tags=$?
|
||||
set -e
|
||||
sed 's/^/ /' "$out"
|
||||
|
||||
if [ "$rc_gen" -ne 0 ] || [ "$rc_tags" -ne 0 ]; then
|
||||
echo
|
||||
echo " FAILED: the tag set we SHIP cannot do what we declare." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# A guard that silently runs nothing is worse than no guard: prove the tests were
|
||||
# actually compiled in and executed (build tags / file renames could exclude them).
|
||||
for want in TestShippedTagSetConstructsDeclaredProtocols TestCompiledTagsMatchTheShippedSet; do
|
||||
if ! SHATER_ROUTER_TAG_CHECK=1 go test -count=1 -v \
|
||||
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
|
||||
-run "$want" ./shater/generate/ ./shater/buildtags/ 2>&1 | grep -q -- "--- PASS: $want"; then
|
||||
echo " FAILED: $want did not run (build-tag/file-name drift?)" >&2
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
echo
|
||||
echo "== OK: the shipped tag set covers every declared feature, and every =="
|
||||
echo "== declared protocol constructs through box.New under it. =="
|
||||
@@ -0,0 +1,78 @@
|
||||
# shellcheck shell=sh
|
||||
#
|
||||
# router-tags.sh — THE build-tag set of the shipped `shaterd` router binary.
|
||||
#
|
||||
# This file is DATA, not a program: it is `.`-sourced by
|
||||
# - scripts/build-shaterd.sh (the ship build)
|
||||
# - scripts/check-router-tags.sh (the guard that proves the set is complete)
|
||||
# and it is PARSED by shater/buildtags/buildtags_test.go, which asserts that
|
||||
# every feature docs-shater/FEATURES.md declares supported has its build tags
|
||||
# present here. Change the set here and nowhere else.
|
||||
#
|
||||
# WHY A TRIMMED SET AT ALL (D9): upstream's DEFAULT_BUILD_TAGS registers the
|
||||
# whole sing-box zoo. We drop what shater/generate can never emit, because a
|
||||
# router binary pays for every tag twice — flash and (UPX unpacks into anonymous
|
||||
# pages) resident RAM. We do NOT drop what a declared feature needs to run.
|
||||
#
|
||||
# WHY THIS FILE EXISTS (the 2026-07-25 WireGuard outage): the set used to be a
|
||||
# string literal inside build-shaterd.sh, with nothing connecting it to the
|
||||
# feature list. `with_gvisor` was trimmed as "unreachable code" while
|
||||
# `with_wireguard` stayed — so every shipped binary answered a WireGuard node
|
||||
# with "gVisor is not included in this build". No test caught it: the test suite
|
||||
# builds with the FULL upstream tag set, so the SHIPPED combination was never
|
||||
# exercised. One file + one test now hold the two halves together (D23).
|
||||
#
|
||||
# ---- the set -----------------------------------------------------------------
|
||||
# with_gvisor userspace netstack. REQUIRED BY with_wireguard: both
|
||||
# transport/wireguard device constructors
|
||||
# (newStackDevice AND newSystemStackDevice) are stubs
|
||||
# returning tun.ErrGVisorNotIncluded without it, so
|
||||
# box.New dies at "create WireGuard device". Not
|
||||
# optional as long as we ship WireGuard/AmneziaWG.
|
||||
# with_quic hysteria2 + tuic outbounds, QUIC/HTTP3 DNS
|
||||
# transports, and the vless/vmess `type=quic` transport
|
||||
# (shater/registry/registry_quic.go).
|
||||
# with_wireguard registers the wireguard endpoint
|
||||
# (shater/registry/registry_wireguard.go).
|
||||
# with_awg AmneziaWG obfuscation params (jc/jmin/jmax, s1-s4,
|
||||
# h1-h4, i1-i5) actually reach the device
|
||||
# (transport/wireguard/device_awg.go). A driving
|
||||
# product requirement — FEATURES.md marks it [MVP].
|
||||
# with_utls uTLS fingerprints AND REALITY: common/tls/
|
||||
# reality_client.go is itself `//go:build with_utls`.
|
||||
# with_xhttp the XHTTP/SplitHTTP v2ray transport
|
||||
# (shater/parse emits type=xhttp).
|
||||
# badlinkname badtls fast path (common/badtls/*.go are
|
||||
# `go1.25 && badlinkname`). Needs -checklinkname=0 in
|
||||
# SHATER_ROUTER_LDFLAGS below or the LINK step fails.
|
||||
# tfogo_checklinkname0 same deal for tfo-go's linkname use.
|
||||
# with_lx_command lx daemon command server. Inert for shaterd (nothing
|
||||
# under shater/ imports sing-box/daemon or libbox —
|
||||
# `go list -deps ./shater/cmd/shaterd` links neither),
|
||||
# kept only so the router set stays a subset of the lx
|
||||
# desktop set. Costs nothing; safe to drop later.
|
||||
#
|
||||
# Deliberately NOT here (each is unreachable for shater, not merely unused):
|
||||
# with_purego,with_naive_outbound cronet-go forces a glibc PT_INTERP even at
|
||||
# CGO_ENABLED=0 -> will not run on musl (D9).
|
||||
# with_clash_api the panel is shater's own web server;
|
||||
# generate emits no clash_api service.
|
||||
# with_dhcp resolver types are udp/tcp/doh/dot/local/
|
||||
# fakeip; no dhcp:// transport is generated.
|
||||
# with_tailscale,with_acme,with_ech,with_usbip,with_cloudflared,with_ocm,
|
||||
# with_ccm,with_v2ray_api,with_reality_server
|
||||
# nothing in shater/parse or shater/generate
|
||||
# can produce them; hysteria2 `ech=` is
|
||||
# refused with a warning in the parser.
|
||||
#
|
||||
# Adding a tag here is cheap. REMOVING one is a product decision: run
|
||||
# `scripts/check-router-tags.sh` — it fails if the set no longer covers a
|
||||
# declared feature, and it fails if the shipped combination cannot construct
|
||||
# every protocol through box.New.
|
||||
|
||||
SHATER_ROUTER_TAGS="with_gvisor,with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command"
|
||||
|
||||
# Linker flags the tag set REQUIRES (they are not optional trimming: `badlinkname`
|
||||
# without -checklinkname=0 fails at link time with
|
||||
# "invalid reference to crypto/tls.(*Conn).handlePostHandshakeMessage").
|
||||
SHATER_ROUTER_LDFLAGS="-checklinkname=0"
|
||||
+84
-12
@@ -74,6 +74,19 @@ type Applier struct {
|
||||
stateMu sync.RWMutex
|
||||
holding bool
|
||||
|
||||
// traffic is WHERE THE TRAFFIC GOES under the config that is currently running:
|
||||
// tunnelled, split, straight out, or blocked (see generate.TrafficOf). It is
|
||||
// computed from the generated option.Options at the moment they are handed to
|
||||
// the engine, so it describes what runs rather than what is on disk.
|
||||
//
|
||||
// Deliberately separate from `holding` and from Plane, which answer "how much of
|
||||
// the data plane is installed". The panel used to read Plane == "full" as
|
||||
// "protected" and said so under a green LED on a router whose only rule was
|
||||
// `default -> direct`; the plane really was fully installed, and the whole LAN
|
||||
// really was going out the plain WAN. Zero value = not known (no successful
|
||||
// apply in this process yet), which is NOT the same as "tunnel".
|
||||
traffic generate.Traffic // guarded by stateMu
|
||||
|
||||
// lastWarnings is the normalised warning set from the last SUCCESSFUL apply,
|
||||
// published through Status so the panel can show fail-open degradations
|
||||
// (a blocklist that did not load, a DoH host left reachable) instead of
|
||||
@@ -192,11 +205,15 @@ func (a *Applier) configureObservatory(m *model.Model, opts option.Options) {
|
||||
})
|
||||
}
|
||||
|
||||
// TestGroups launches the engine's one-shot exit test (delay + exit address, F2)
|
||||
// of the named groups/chains, returning started=false when a run is already in
|
||||
// flight or the engine is absent. names empty/nil = every group and every chain.
|
||||
// It takes NEITHER the apply mutex nor the flock, so kicking off a test never
|
||||
// blocks behind an Apply.
|
||||
// TestGroups launches the engine's one-shot group/chain test (F2): the engine
|
||||
// asks its observatory for an out-of-turn pass and reports what it measured,
|
||||
// plus the exit address for alive rule-routed targets — it no longer dials
|
||||
// health probes of its own. Returns started=false when a run is already in
|
||||
// flight or the engine is absent. names empty/nil = every group and every
|
||||
// chain. probeURL is passed through for signature stability and IGNORED by the
|
||||
// engine (the probe URL is a global observatory setting now). It takes NEITHER
|
||||
// the apply mutex nor the flock, so kicking off a test never blocks behind an
|
||||
// Apply.
|
||||
func (a *Applier) TestGroups(names []string, probeURL string) (started bool) {
|
||||
if a.eng == nil {
|
||||
return false
|
||||
@@ -514,6 +531,12 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
|
||||
}
|
||||
a.setHolding(false)
|
||||
a.lastGood = m
|
||||
// Publish where this config actually sends traffic, read off the very options
|
||||
// the engine was just handed (a.eng.Apply above). The engine's hash fast-path
|
||||
// may have skipped a swap, in which case these options are hash-equal to what is
|
||||
// already running — either way they are the running config, which is the only
|
||||
// config this verdict may describe.
|
||||
a.setTraffic(generate.TrafficOf(opts))
|
||||
// Publish the warnings of THIS successful apply, in one normalised set, and log
|
||||
// them in one consistent format. Status carries them to the panel so a
|
||||
// fail-open degradation is visible in the UI instead of only in logread.
|
||||
@@ -557,6 +580,11 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
|
||||
// With the kill-switch OPEN nothing is installed — fail-open is the operator's
|
||||
// documented choice and this must not quietly override it.
|
||||
func (a *Applier) holdLocked(m *model.Model, cause error) {
|
||||
// The engine is not carrying anything, so whatever the last running config did
|
||||
// with traffic is no longer true of this router. Forget it either way — a stale
|
||||
// "tunnel" verdict left behind by a config that is no longer running is the same
|
||||
// reassuring lie in a different place.
|
||||
a.setTraffic(generate.Traffic{})
|
||||
if !killSwitchClosed(m.Globals) {
|
||||
a.log.Warn("engine is down and kill_switch=open: LAN traffic is NOT protected (documented fail-open): ", cause)
|
||||
return
|
||||
@@ -640,24 +668,52 @@ func (a *Applier) setHolding(v bool) {
|
||||
a.stateMu.Unlock()
|
||||
}
|
||||
|
||||
// Traffic returns where the traffic of the CURRENTLY RUNNING config goes. The
|
||||
// zero value means no config of this process's is running (nothing applied yet,
|
||||
// or the plane was torn down / put on hold), and callers must render that as
|
||||
// "unknown", never as protected.
|
||||
func (a *Applier) Traffic() generate.Traffic {
|
||||
a.stateMu.RLock()
|
||||
defer a.stateMu.RUnlock()
|
||||
return a.traffic
|
||||
}
|
||||
|
||||
func (a *Applier) setTraffic(t generate.Traffic) {
|
||||
a.stateMu.Lock()
|
||||
a.traffic = t
|
||||
a.stateMu.Unlock()
|
||||
}
|
||||
|
||||
func (a *Applier) setWarnings(ws []Warning) {
|
||||
a.stateMu.Lock()
|
||||
a.lastWarnings = ws
|
||||
a.stateMu.Unlock()
|
||||
}
|
||||
|
||||
// Warnings returns the normalised warning set from the last successful apply.
|
||||
// Never nil: an empty slice means "the last apply was clean", which the panel
|
||||
// must render differently from "no apply has run yet" (Active/Plane cover that).
|
||||
// Warnings returns the normalised warning set from the last successful apply,
|
||||
// PLUS whatever is wrong right now that no apply can describe. Never nil: an
|
||||
// empty slice means "the last apply was clean", which the panel must render
|
||||
// differently from "no apply has run yet" (Active/Plane cover that).
|
||||
//
|
||||
// The live half is currently the engine's abandoned generations
|
||||
// (engineTeardownWarnings). It is computed at READ time rather than folded into
|
||||
// lastWarnings on purpose: a superseded box that will not shut down is a
|
||||
// condition of the process, not a property of a config. Folding it in would make
|
||||
// it appear only after the NEXT successful apply and then stay published long
|
||||
// after the shutdown finally completed — reporting a leak that is over, and
|
||||
// staying silent about one that is not. Read-time means it shows up the instant
|
||||
// it happens and clears itself the instant it resolves.
|
||||
func (a *Applier) Warnings() []Warning {
|
||||
var out []Warning
|
||||
if a.eng != nil {
|
||||
out = engineTeardownWarnings(a.eng.PendingCloses())
|
||||
}
|
||||
a.stateMu.RLock()
|
||||
defer a.stateMu.RUnlock()
|
||||
if a.lastWarnings == nil {
|
||||
if out == nil && a.lastWarnings == nil {
|
||||
return []Warning{}
|
||||
}
|
||||
out := make([]Warning, len(a.lastWarnings))
|
||||
copy(out, a.lastWarnings)
|
||||
return out
|
||||
return append(out, a.lastWarnings...)
|
||||
}
|
||||
|
||||
// Reconcile re-reads UCI and either tears down (disabled) or re-applies (enabled).
|
||||
@@ -723,6 +779,7 @@ func (a *Applier) Teardown() error {
|
||||
a.lastGood = nil
|
||||
a.lastNft = ""
|
||||
a.setHolding(false)
|
||||
a.setTraffic(generate.Traffic{})
|
||||
a.setWarnings(nil)
|
||||
// The plane is gone, so the logged set no longer describes anything. Forget it,
|
||||
// and the next apply re-announces its warnings in full rather than staying
|
||||
@@ -1003,6 +1060,20 @@ type Status struct {
|
||||
// say so rather than looking healthy.
|
||||
Plane string `json:"plane"`
|
||||
|
||||
// Traffic is WHERE THE TRAFFIC GOES under the running config — tunnelled, split,
|
||||
// straight out, or blocked (see generate.Traffic).
|
||||
//
|
||||
// Plane does NOT answer this, and reading it as if it did is the defect this
|
||||
// field exists for. Plane == "full" only means the table, the policy routing and
|
||||
// the engine are all in place; a router whose one rule is `default -> direct` has
|
||||
// all three and sends every packet out the plain WAN with its real address. The
|
||||
// panel showed that as "Protected — traffic is going through the tunnel".
|
||||
//
|
||||
// Zero value (Verdict == "") means unknown: no successful apply has run in this
|
||||
// daemon process yet, or the plane is on hold / torn down. A consumer must render
|
||||
// that as unknown and never as protected.
|
||||
Traffic generate.Traffic `json:"traffic"`
|
||||
|
||||
// Warnings is the normalised warning set from the last successful apply:
|
||||
// everything that was skipped, degraded or left un-applied while the apply
|
||||
// still succeeded. Always non-nil so the panel can map over it unconditionally.
|
||||
@@ -1082,6 +1153,7 @@ func (a *Applier) Status() Status {
|
||||
Hash: a.eng.Hash(),
|
||||
CanRollback: a.canRollback(),
|
||||
EngineRunning: a.eng.Running(),
|
||||
Traffic: a.Traffic(),
|
||||
Warnings: a.Warnings(),
|
||||
}
|
||||
s.StartedUnix, s.UptimeSeconds = processUptime(time.Now())
|
||||
|
||||
@@ -0,0 +1,283 @@
|
||||
package apply
|
||||
|
||||
// Regression cover for the leaked-engine-generation defect.
|
||||
//
|
||||
// Observed on the router: one shaterd process was carrying up to FOUR sing-box
|
||||
// instances at once. sing-box stamps every log line with the elapsed seconds of
|
||||
// ITS OWN instance, so the same process printed `ERROR[2015]` and `ERROR[0129]`
|
||||
// in the same second — two engines half an hour apart in age, both alive, both
|
||||
// dialling, both holding WireGuard devices built from the same private keys. A
|
||||
// full daemon stop+start collapsed it back to one generation, which places the
|
||||
// leak squarely on the config re-apply path rather than on startup.
|
||||
//
|
||||
// The tests below pin the two halves of the fix:
|
||||
//
|
||||
// TestApplySwapsLeaveExactlyOneEngineGeneration — the healthy path really
|
||||
// retires the old instance (its listener is provably gone), N times in a row.
|
||||
// TestStuckEngineCloseDoesNotBlockTheApply — a shutdown that never returns
|
||||
// is bounded, does not stall the apply, and is REPORTED as a critical warning
|
||||
// for exactly as long as it is true.
|
||||
|
||||
import (
|
||||
"io"
|
||||
"net"
|
||||
"net/netip"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
"github.com/sagernet/sing-box/shater/engine"
|
||||
"github.com/sagernet/sing/common/json/badoption"
|
||||
)
|
||||
|
||||
// mixedOn builds a minimal but REAL engine config: a mixed inbound bound to
|
||||
// 127.0.0.1:port plus a direct outbound. Two configs with different ports hash
|
||||
// differently, so each Apply is a genuine swap rather than a hash-gate no-op —
|
||||
// and the bound port is the observable that proves whether the old instance
|
||||
// actually died.
|
||||
func mixedOn(port uint16) option.Options {
|
||||
listen := badoption.Addr(netip.MustParseAddr("127.0.0.1"))
|
||||
return option.Options{
|
||||
Log: &option.LogOptions{Level: "error"},
|
||||
Inbounds: []option.Inbound{{
|
||||
Type: C.TypeMixed,
|
||||
Tag: "mixed-in",
|
||||
Options: &option.HTTPMixedInboundOptions{
|
||||
ListenOptions: option.ListenOptions{Listen: &listen, ListenPort: port},
|
||||
},
|
||||
}},
|
||||
Outbounds: []option.Outbound{{
|
||||
Type: C.TypeDirect,
|
||||
Tag: "direct-out",
|
||||
Options: &option.DirectOutboundOptions{},
|
||||
}},
|
||||
}
|
||||
}
|
||||
|
||||
// portFree reports whether 127.0.0.1:port can be bound right now — i.e. whether
|
||||
// the instance that used to listen there is really gone. Retried briefly because
|
||||
// a listener is released by Close, not by the return of Close's caller.
|
||||
func portFree(port uint16) bool {
|
||||
deadline := time.Now().Add(3 * time.Second)
|
||||
for {
|
||||
ln, err := net.Listen("tcp", net.JoinHostPort("127.0.0.1", strconv.Itoa(int(port))))
|
||||
if err == nil {
|
||||
_ = ln.Close()
|
||||
return true
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
return false
|
||||
}
|
||||
time.Sleep(20 * time.Millisecond)
|
||||
}
|
||||
}
|
||||
|
||||
// TestApplySwapsLeaveExactlyOneEngineGeneration is the core regression: after N
|
||||
// sequential applies the process must be carrying ONE engine, not N.
|
||||
//
|
||||
// "Carrying" is checked two ways on purpose. Generations() is the engine's own
|
||||
// accounting (running instance + every retirement still in flight) and would
|
||||
// catch a retirement that silently never completes. The port check is
|
||||
// independent of that accounting: if the superseded instance were still alive it
|
||||
// would still hold its listener, and the bind would fail. A fix that only
|
||||
// reset a pointer would pass the first check and fail the second.
|
||||
func TestApplySwapsLeaveExactlyOneEngineGeneration(t *testing.T) {
|
||||
const (
|
||||
firstPort = 18801
|
||||
applies = 5
|
||||
)
|
||||
|
||||
a := New(engine.New(), nil)
|
||||
t.Cleanup(func() { _ = a.eng.Close() })
|
||||
|
||||
for i := 0; i < applies; i++ {
|
||||
port := uint16(firstPort + i)
|
||||
changed, err := a.eng.Apply(mixedOn(port))
|
||||
if err != nil {
|
||||
t.Fatalf("apply #%d (port %d): %v", i+1, port, err)
|
||||
}
|
||||
if !changed {
|
||||
t.Fatalf("apply #%d: every config here differs, so the swap must be real", i+1)
|
||||
}
|
||||
if got := a.eng.Generations(); got != 1 {
|
||||
t.Fatalf("after apply #%d the process carries %d engine instances, want exactly 1 — "+
|
||||
"a superseded generation is still alive (this is the four-generations-in-one-process leak)",
|
||||
i+1, got)
|
||||
}
|
||||
if stuck := a.eng.PendingCloses(); len(stuck) != 0 {
|
||||
t.Fatalf("after apply #%d: %d generation(s) abandoned, want none: %+v", i+1, len(stuck), stuck)
|
||||
}
|
||||
if i > 0 {
|
||||
prev := uint16(firstPort + i - 1)
|
||||
if !portFree(prev) {
|
||||
t.Fatalf("after apply #%d the PREVIOUS generation still holds 127.0.0.1:%d — "+
|
||||
"the old engine was replaced in the field but never actually stopped", i+1, prev)
|
||||
}
|
||||
}
|
||||
// The warning set must stay clean while teardown is healthy: a critical
|
||||
// warning that cries wolf on every apply is worse than none.
|
||||
for _, w := range a.Warnings() {
|
||||
if w.Section == "engine" {
|
||||
t.Fatalf("after apply #%d a healthy swap produced an engine warning: %+v", i+1, w)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if err := a.eng.Close(); err != nil {
|
||||
t.Fatalf("close: %v", err)
|
||||
}
|
||||
if got := a.eng.Generations(); got != 0 {
|
||||
t.Fatalf("after Close the process carries %d engine instances, want 0", got)
|
||||
}
|
||||
if !portFree(firstPort + applies - 1) {
|
||||
t.Fatalf("after Close the last generation still holds its listener")
|
||||
}
|
||||
}
|
||||
|
||||
// hangingCloser returns a box closer that BLOCKS the first close it is handed
|
||||
// until release() is called, and performs every later close normally. That is the
|
||||
// shape of the real fault: one subsystem of one generation (a WireGuard endpoint)
|
||||
// refuses to come down, while the rest of the process is fine.
|
||||
func hangingCloser() (closer func(io.Closer) error, release func()) {
|
||||
gate := make(chan struct{})
|
||||
first := make(chan struct{}, 1)
|
||||
first <- struct{}{}
|
||||
return func(c io.Closer) error {
|
||||
select {
|
||||
case <-first:
|
||||
<-gate // the stuck generation: never returns until released
|
||||
return c.Close() // ...and then really does close, so the port frees
|
||||
default:
|
||||
return c.Close()
|
||||
}
|
||||
}, func() {
|
||||
close(gate)
|
||||
}
|
||||
}
|
||||
|
||||
// TestStuckEngineCloseDoesNotBlockTheApply pins all four requirements of the
|
||||
// bounded teardown at once:
|
||||
//
|
||||
// 1. the apply COMPLETES — a shutdown that never returns must not hold the
|
||||
// control plane (and therefore the panel) hostage;
|
||||
// 2. the new engine is running afterwards — fail-closed semantics are unchanged,
|
||||
// the swap succeeded;
|
||||
// 3. the abandoned generation is REPORTED as a critical warning, by name, for as
|
||||
// long as it is still running — it is not silently swallowed;
|
||||
// 4. generations do not stack: a further apply while the leak persists leaves
|
||||
// one live instance plus the one abandoned one, not three.
|
||||
func TestStuckEngineCloseDoesNotBlockTheApply(t *testing.T) {
|
||||
const budget = 200 * time.Millisecond
|
||||
restoreBudget := engine.SetCloseBudget(budget)
|
||||
defer restoreBudget()
|
||||
|
||||
a := New(engine.New(), nil)
|
||||
|
||||
// Generation 1 comes up with the REAL closer still installed.
|
||||
if _, err := a.eng.Apply(mixedOn(18811)); err != nil {
|
||||
t.Fatalf("apply #1: %v", err)
|
||||
}
|
||||
|
||||
closer, release := hangingCloser()
|
||||
restoreCloser := engine.SetBoxCloser(closer)
|
||||
released := false
|
||||
defer func() {
|
||||
if !released {
|
||||
release()
|
||||
}
|
||||
restoreCloser()
|
||||
_ = a.eng.Close()
|
||||
}()
|
||||
|
||||
// (1) Generation 2: the retirement of generation 1 will never return.
|
||||
start := time.Now()
|
||||
changed, err := a.eng.Apply(mixedOn(18812))
|
||||
elapsed := time.Since(start)
|
||||
if err != nil {
|
||||
t.Fatalf("apply #2 must SUCCEED despite the stuck teardown: %v", err)
|
||||
}
|
||||
if !changed {
|
||||
t.Fatalf("apply #2: expected a real swap")
|
||||
}
|
||||
// Generously bounded: the budget plus box.New+Start. The point is that it
|
||||
// returned at all — before the fix this waited on Box.Close forever.
|
||||
if elapsed > budget+20*time.Second {
|
||||
t.Fatalf("apply #2 took %s: a stuck teardown must not stall the apply", elapsed)
|
||||
}
|
||||
|
||||
// (2) fail-closed semantics unchanged: the new engine really is up.
|
||||
if !a.eng.Running() {
|
||||
t.Fatalf("apply #2: the new engine must be running")
|
||||
}
|
||||
|
||||
// (3) the leak is visible, named, and critical.
|
||||
stuck := a.eng.PendingCloses()
|
||||
if len(stuck) != 1 {
|
||||
t.Fatalf("PendingCloses() = %+v, want exactly the one abandoned generation", stuck)
|
||||
}
|
||||
if stuck[0].Generation != 1 {
|
||||
t.Errorf("abandoned generation = %d, want 1", stuck[0].Generation)
|
||||
}
|
||||
ws := a.Status().Warnings // the exact set `shaterd status` and the panel read
|
||||
var found *Warning
|
||||
for i := range ws {
|
||||
if ws[i].Section == "engine" {
|
||||
found = &ws[i]
|
||||
break
|
||||
}
|
||||
}
|
||||
if found == nil {
|
||||
t.Fatalf("a superseded engine that will not shut down produced NO warning; "+
|
||||
"Status would show a healthy router: %+v", ws)
|
||||
}
|
||||
if found.Severity != SeverityCritical {
|
||||
t.Errorf("stuck-teardown warning severity = %q, want %q", found.Severity, SeverityCritical)
|
||||
}
|
||||
if !strings.Contains(found.Name, "generation 1") {
|
||||
t.Errorf("stuck-teardown warning must name the generation, got Name=%q", found.Name)
|
||||
}
|
||||
if !strings.Contains(found.Message, "STILL RUNNING") {
|
||||
t.Errorf("stuck-teardown warning must say the instance is still running, got %q", found.Message)
|
||||
}
|
||||
if got := a.eng.Generations(); got != 2 {
|
||||
t.Fatalf("Generations() = %d, want 2 (one live + one abandoned)", got)
|
||||
}
|
||||
|
||||
// (4) another apply while the leak persists must not add a THIRD generation:
|
||||
// with a generation abandoned the swap goes close-old-then-start-new, so the
|
||||
// process still holds one live instance plus the one that will not die.
|
||||
if _, err := a.eng.Apply(mixedOn(18813)); err != nil {
|
||||
t.Fatalf("apply #3: %v", err)
|
||||
}
|
||||
if got := a.eng.Generations(); got != 2 {
|
||||
t.Fatalf("Generations() = %d after a third apply, want 2 — generations are stacking, "+
|
||||
"which is exactly the four-live-engines fault", got)
|
||||
}
|
||||
if stuck := a.eng.PendingCloses(); len(stuck) != 1 || stuck[0].Generation != 1 {
|
||||
t.Fatalf("PendingCloses() = %+v, want only the original abandoned generation 1", stuck)
|
||||
}
|
||||
|
||||
// (5) and it CLEARS: when the shutdown finally completes the warning goes away
|
||||
// on its own. A leak report that outlives the leak trains the operator to
|
||||
// ignore the panel.
|
||||
release()
|
||||
released = true
|
||||
deadline := time.Now().Add(5 * time.Second)
|
||||
for len(a.eng.PendingCloses()) > 0 && time.Now().Before(deadline) {
|
||||
time.Sleep(10 * time.Millisecond)
|
||||
}
|
||||
if got := a.eng.PendingCloses(); len(got) != 0 {
|
||||
t.Fatalf("the finished shutdown is still reported as abandoned: %+v", got)
|
||||
}
|
||||
for _, w := range a.Warnings() {
|
||||
if w.Section == "engine" {
|
||||
t.Fatalf("the engine warning outlived the leak it describes: %+v", w)
|
||||
}
|
||||
}
|
||||
if got := a.eng.Generations(); got != 1 {
|
||||
t.Fatalf("Generations() = %d after the stuck shutdown completed, want 1", got)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,47 @@
|
||||
package apply
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"testing"
|
||||
|
||||
"github.com/sagernet/sing-box/shater/engine"
|
||||
"github.com/sagernet/sing-box/shater/generate"
|
||||
"github.com/sagernet/sing-box/shater/model"
|
||||
)
|
||||
|
||||
// TestStatusReportsTraffic pins the wiring the panel's headline depends on.
|
||||
//
|
||||
// Plane says how much of the data plane is installed; it does NOT say where the
|
||||
// traffic goes, and reading it as if it did put "Protected — traffic is going
|
||||
// through the tunnel" on a router whose only rule was `default -> direct`. The
|
||||
// verdict that answers the real question travels in Status.Traffic, so it must
|
||||
// (a) start unknown, (b) surface what the last successful apply published, and
|
||||
// (c) go back to unknown the moment the engine stops carrying that config.
|
||||
func TestStatusReportsTraffic(t *testing.T) {
|
||||
a := New(engine.New(), nil)
|
||||
|
||||
// A fresh applier has applied nothing, so it knows nothing. The zero value must
|
||||
// NOT read as any verdict — least of all "tunnel".
|
||||
if got := a.Status().Traffic; got.Verdict != "" {
|
||||
t.Fatalf("fresh applier: Traffic.Verdict = %q, want \"\" (unknown)", got.Verdict)
|
||||
}
|
||||
|
||||
a.setTraffic(generate.Traffic{Verdict: generate.VerdictTunnel, Default: "auto", TunnelRules: 2})
|
||||
got := a.Status().Traffic
|
||||
if got.Verdict != generate.VerdictTunnel || got.Default != "auto" || got.TunnelRules != 2 {
|
||||
t.Fatalf("Status().Traffic = %+v, want the verdict the last apply published", got)
|
||||
}
|
||||
|
||||
// The engine is down and the config it was running is no longer in force. A
|
||||
// verdict left over from it would be the same reassuring lie, one layer down.
|
||||
// (kill_switch=open takes holdLocked's early return, which is precisely the path
|
||||
// that must still forget the verdict.)
|
||||
g := model.DefaultGlobals()
|
||||
g.KillSwitch = "open"
|
||||
a.mu.Lock()
|
||||
a.holdLocked(&model.Model{Globals: g}, errors.New("engine start failed"))
|
||||
a.mu.Unlock()
|
||||
if got := a.Status().Traffic; got.Verdict != "" {
|
||||
t.Fatalf("after hold: Traffic.Verdict = %q, want \"\" (unknown) — the engine is not carrying that config", got.Verdict)
|
||||
}
|
||||
}
|
||||
@@ -24,7 +24,9 @@ import (
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/sagernet/sing-box/shater/engine"
|
||||
"github.com/sagernet/sing-box/shater/model"
|
||||
"github.com/sagernet/sing-box/shater/netplane"
|
||||
)
|
||||
@@ -269,6 +271,50 @@ func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
|
||||
}
|
||||
}
|
||||
|
||||
// engineTeardownWarnings turns the engine's ABANDONED generations — superseded
|
||||
// sing-box instances whose shutdown overran the hard close budget and are still
|
||||
// running inside this process — into operator-facing warnings.
|
||||
//
|
||||
// Critical, without hesitation. A leaked generation is not untidiness:
|
||||
//
|
||||
// - it still holds its WireGuard devices, and two devices built from the same
|
||||
// private key evict each other at the peer (one session per public key), so
|
||||
// the leak reproduces BETWEEN generations exactly the fault
|
||||
// generate/wgdedup.go removes WITHIN a config — the tunnel flaps and neither
|
||||
// end can say why;
|
||||
// - it still holds its outbound connections and keeps probing nodes, so the
|
||||
// log fills with errors attributed to a config that is no longer applied;
|
||||
// - on a 512 MiB router each one costs real memory that is never returned.
|
||||
//
|
||||
// The generation number is carried in Name so two consecutive status reads can
|
||||
// tell "the same stuck generation" from "another one just leaked", and the
|
||||
// elapsed time is in the message because a shutdown at 8s and one at 40 minutes
|
||||
// are different problems.
|
||||
func engineTeardownWarnings(stuck []engine.StuckClose) []Warning {
|
||||
if len(stuck) == 0 {
|
||||
return nil
|
||||
}
|
||||
out := make([]Warning, 0, len(stuck))
|
||||
for _, s := range stuck {
|
||||
config := "unknown config"
|
||||
if len(s.Hash) >= 12 {
|
||||
config = "config " + s.Hash[:12]
|
||||
}
|
||||
out = append(out, Warning{
|
||||
Severity: SeverityCritical,
|
||||
Section: "engine",
|
||||
Name: fmt.Sprintf("generation %d", s.Generation),
|
||||
Message: fmt.Sprintf(
|
||||
"a superseded engine instance (%s) has been shutting down for %s and is STILL RUNNING: "+
|
||||
"it keeps its outbound connections and its WireGuard devices, so it can evict the "+
|
||||
"live tunnel at the peer and it keeps writing to the log. The current configuration "+
|
||||
"is applied and running; restart shaterd if this does not clear.",
|
||||
config, s.Elapsed.Round(time.Second)),
|
||||
})
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func severityRank(s string) int {
|
||||
switch s {
|
||||
case SeverityCritical:
|
||||
|
||||
@@ -0,0 +1,206 @@
|
||||
// Package buildtags is the contract between what shater DECLARES it supports
|
||||
// and the build tags the shipped router binary is actually compiled with.
|
||||
//
|
||||
// # Why this package exists
|
||||
//
|
||||
// The router binary is built with a deliberately trimmed tag set (D9/D23,
|
||||
// scripts/router-tags.sh) — upstream's full set registers a zoo shater/generate
|
||||
// can never emit, and a router pays for every tag in flash and in RAM. Trimming
|
||||
// is right; trimming BLIND is not. On 2026-07-25 a production router answered a
|
||||
// configured WireGuard node with
|
||||
//
|
||||
// create instance: initialize endpoint[0]: create WireGuard device:
|
||||
// gVisor is not included in this build, rebuild with -tags with_gvisor
|
||||
//
|
||||
// because `with_gvisor` had been trimmed as "unreachable code" (true for the tun
|
||||
// inbound we never emit — false for the WireGuard endpoint we ship and declare
|
||||
// [MVP]) while `with_wireguard` stayed. Nothing caught it: the test suite builds
|
||||
// with the FULL upstream tag set, so the SHIPPED tag combination was, at that
|
||||
// point, the one configuration nothing in the repo ever exercised.
|
||||
//
|
||||
// # What holds it together now
|
||||
//
|
||||
// 1. Features below names each declared feature and the build tags it needs to
|
||||
// RUN (not merely to compile). shater/buildtags's own test parses
|
||||
// scripts/router-tags.sh and fails if the shipped set does not cover them —
|
||||
// it needs no tags, no Linux and no network, so it runs in every plain
|
||||
// `go test ./...`.
|
||||
// 2. shater/generate's TestShippedTagSetConstructsDeclaredProtocols drives one
|
||||
// node of every declared protocol through box.New under whatever tags the
|
||||
// test binary was built with, skipping only what is genuinely not compiled
|
||||
// in. scripts/check-router-tags.sh runs it with the SHIPPED set, so the
|
||||
// combination we ship is proven to construct, not merely to link.
|
||||
//
|
||||
// (1) catches a trimmed dependency the moment it is trimmed; (2) catches the
|
||||
// class of failure (1) cannot model — a tag that is present but insufficient.
|
||||
//
|
||||
// Adding a protocol to shater/parse + shater/generate means adding a row here.
|
||||
package buildtags
|
||||
|
||||
import "sort"
|
||||
|
||||
// Feature is one capability the product declares, together with the build tags
|
||||
// the binary must carry for it to work at runtime.
|
||||
type Feature struct {
|
||||
// Name is the feature as a user would name it.
|
||||
Name string
|
||||
// Declared points at where we promise it (docs-shater/FEATURES.md section,
|
||||
// or the generator/registry that emits it).
|
||||
Declared string
|
||||
// Tags are ALL build tags required for the feature to work — including
|
||||
// transitive ones (with_awg alone is useless without with_wireguard, which
|
||||
// is useless without with_gvisor). Listing them transitively is deliberate:
|
||||
// the check must not depend on a dependency graph nobody maintains.
|
||||
Tags []string
|
||||
// Why explains what breaks without those tags, with the code anchor. It is
|
||||
// printed by the failing test, so a future trimmer reads the reason instead
|
||||
// of rediscovering it on a router.
|
||||
Why string
|
||||
}
|
||||
|
||||
// Features is the authoritative list. Only tag-GATED capabilities belong here:
|
||||
// tproxy, routing rules, rule-sets, the DNS filter, nft/policy routing and the
|
||||
// panel are compiled unconditionally and cannot be lost to a tag trim.
|
||||
var Features = []Feature{
|
||||
{
|
||||
Name: "WireGuard nodes (wg:// / wireguard:// links, wg-quick .conf import)",
|
||||
Declared: "FEATURES.md §Proxy engine — “VLESS, VMess, Trojan, Shadowsocks, WireGuard” [MVP]",
|
||||
Tags: []string{"with_wireguard", "with_gvisor"},
|
||||
Why: "with_wireguard registers the endpoint (shater/registry/registry_wireguard.go); " +
|
||||
"with_gvisor supplies the userspace netstack EVERY WireGuard device needs — without it " +
|
||||
"transport/wireguard/device_stack_stub.go returns tun.ErrGVisorNotIncluded from BOTH " +
|
||||
"newStackDevice and newSystemStackDevice, so box.New fails with " +
|
||||
"\"create WireGuard device: gVisor is not included in this build\" and the node is dead. " +
|
||||
"system_interface=true is not an escape hatch: it hits the same stub.",
|
||||
},
|
||||
{
|
||||
Name: "AmneziaWG obfuscation (awg:// links; jc/jmin/jmax, s1-s4, h1-h4, i1-i5)",
|
||||
Declared: "FEATURES.md §Proxy engine — “AmneziaWG 2.0 … a driving requirement” [MVP]",
|
||||
Tags: []string{"with_awg", "with_wireguard", "with_gvisor"},
|
||||
Why: "with_awg makes the AWG params reach the device (transport/wireguard/device_awg.go); " +
|
||||
"without it they parse and are silently ignored (option/wireguard.go). It rides on the " +
|
||||
"WireGuard endpoint, so it needs that feature's tags too.",
|
||||
},
|
||||
{
|
||||
Name: "Hysteria2 nodes (hysteria2:// / hy2://)",
|
||||
Declared: "FEATURES.md §Proxy engine [T1]; shater/registry registerQUICOutbounds",
|
||||
Tags: []string{"with_quic"},
|
||||
Why: "hysteria2.RegisterOutbound is compiled only under with_quic (shater/registry/registry_quic.go); without it box.New rejects the outbound as an unknown type.",
|
||||
},
|
||||
{
|
||||
Name: "TUIC nodes (tuic://)",
|
||||
Declared: "FEATURES.md §Proxy engine [T1]; shater/registry registerQUICOutbounds",
|
||||
Tags: []string{"with_quic"},
|
||||
Why: "tuic.RegisterOutbound is compiled only under with_quic (shater/registry/registry_quic.go).",
|
||||
},
|
||||
{
|
||||
Name: "VLESS/VMess over the QUIC v2ray transport (type=quic)",
|
||||
Declared: "FEATURES.md §Proxy engine — “Transports: TCP/WS/gRPC/HTTPUpgrade/H2/QUIC” [MVP]",
|
||||
Tags: []string{"with_quic"},
|
||||
Why: "transport/v2rayquic registers its constructor from an init() blank-imported only under with_quic; without it NewQUICClient returns os.ErrInvalid at dial time.",
|
||||
},
|
||||
{
|
||||
Name: "QUIC / HTTP3 DNS transports (quic://, h3://)",
|
||||
Declared: "shater/registry registerQUICTransports",
|
||||
Tags: []string{"with_quic"},
|
||||
Why: "dns/transport/quic is registered only under with_quic (shater/registry/registry_quic.go).",
|
||||
},
|
||||
{
|
||||
Name: "REALITY (vless security=reality, pbk/sid)",
|
||||
Declared: "FEATURES.md §Proxy engine — “Reality/XTLS” [MVP]; shater/parse security=reality",
|
||||
Tags: []string{"with_utls"},
|
||||
Why: "the REALITY client lives in common/tls/reality_client.go, which is itself `//go:build with_utls`; without the tag a reality config is rejected by the TLS layer.",
|
||||
},
|
||||
{
|
||||
Name: "uTLS ClientHello fingerprints (fp=chrome/firefox/safari/…)",
|
||||
Declared: "shater/generate/outbound.go TLS mapping (UTLS options)",
|
||||
Tags: []string{"with_utls"},
|
||||
Why: "common/tls/utls_client.go is `//go:build with_utls`; the stub (utls_stub.go) refuses a config that sets a fingerprint.",
|
||||
},
|
||||
{
|
||||
Name: "XHTTP / SplitHTTP transport (type=xhttp, type=splithttp)",
|
||||
Declared: "FEATURES.md §Proxy engine [T1]; shater/parse/sharelink.go case \"xhttp\"",
|
||||
Tags: []string{"with_xhttp"},
|
||||
Why: "transport/v2rayxhttp registers the \"xhttp\" transport from an init() blank-imported only under with_xhttp (shater/registry/registry_xhttp.go); without it the transport type is unknown at box.New.",
|
||||
},
|
||||
{
|
||||
Name: "badtls fast path (zero-copy TLS read-wait / ktls, used by every TLS outbound)",
|
||||
Declared: "common/badtls — linked unconditionally by the TLS client",
|
||||
Tags: []string{"badlinkname", "tfogo_checklinkname0"},
|
||||
Why: "common/badtls/*.go are `go1.25 && badlinkname`; without the tag the package degrades to read_wait_stub.go. " +
|
||||
"These two tags additionally REQUIRE -checklinkname=0 in the linker flags — the build fails at link time otherwise " +
|
||||
"(\"invalid reference to crypto/tls.(*Conn).handlePostHandshakeMessage\"), which is why " +
|
||||
"scripts/router-tags.sh carries SHATER_ROUTER_LDFLAGS next to the tag set.",
|
||||
},
|
||||
}
|
||||
|
||||
// RequiredTags is the union of every declared feature's tags, sorted.
|
||||
func RequiredTags() []string {
|
||||
seen := map[string]bool{}
|
||||
for _, f := range Features {
|
||||
for _, t := range f.Tags {
|
||||
seen[t] = true
|
||||
}
|
||||
}
|
||||
return sortedKeys(seen)
|
||||
}
|
||||
|
||||
// Compiled reports the shater-relevant build tags THIS binary was compiled with,
|
||||
// sorted. It is populated by the one-line tag_*.go twins in this package; a tag
|
||||
// with no file here is simply not tracked (and must not appear in Features).
|
||||
func Compiled() []string { return sortedKeys(compiled) }
|
||||
|
||||
// Has reports whether this binary was compiled with tag.
|
||||
func Has(tag string) bool { return compiled[tag] }
|
||||
|
||||
// MissingTags returns the tags f needs that this binary lacks, sorted. Empty
|
||||
// means the feature is fully compiled in.
|
||||
func MissingTags(f Feature) []string {
|
||||
missing := map[string]bool{}
|
||||
for _, t := range f.Tags {
|
||||
if !compiled[t] {
|
||||
missing[t] = true
|
||||
}
|
||||
}
|
||||
return sortedKeys(missing)
|
||||
}
|
||||
|
||||
// Tracked reports whether tag has a detector file (tag_*.go) in this package.
|
||||
// Features must only reference tracked tags — an untracked tag would silently
|
||||
// read as "not compiled" and turn a real check into a skip. TestFeatureTagsAreTracked
|
||||
// enforces that, and scripts/check-router-tags.sh additionally proves the
|
||||
// detectors match the tag set the compiler was actually handed.
|
||||
func Tracked(tag string) bool { return tracked[tag] }
|
||||
|
||||
// TrackedTags is every tag this package can observe, i.e. exactly the tags with
|
||||
// a tag_*.go detector. Keep the two in sync — the check script fails loudly if
|
||||
// they drift.
|
||||
func TrackedTags() []string { return sortedKeys(tracked) }
|
||||
|
||||
var tracked = map[string]bool{
|
||||
"with_gvisor": true,
|
||||
"with_quic": true,
|
||||
"with_wireguard": true,
|
||||
"with_awg": true,
|
||||
"with_utls": true,
|
||||
"with_xhttp": true,
|
||||
"with_lx_command": true,
|
||||
"badlinkname": true,
|
||||
"tfogo_checklinkname0": true,
|
||||
}
|
||||
|
||||
// compiled is filled by the tag_*.go detectors' init(). A tag with no detector
|
||||
// file compiled in is absent from the map, which reads as "not compiled".
|
||||
var compiled = map[string]bool{}
|
||||
|
||||
// mark records that tag is compiled into this binary.
|
||||
func mark(tag string) { compiled[tag] = true }
|
||||
|
||||
func sortedKeys(m map[string]bool) []string {
|
||||
out := make([]string, 0, len(m))
|
||||
for k := range m {
|
||||
out = append(out, k)
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
package buildtags
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// repoFile reads a file relative to the repo root (this package sits at
|
||||
// <repo>/shater/buildtags).
|
||||
func repoFile(t *testing.T, rel string) string {
|
||||
t.Helper()
|
||||
b, err := os.ReadFile(filepath.Join("..", "..", filepath.FromSlash(rel)))
|
||||
if err != nil {
|
||||
t.Fatalf("read %s: %v", rel, err)
|
||||
}
|
||||
return string(b)
|
||||
}
|
||||
|
||||
// shVar pulls VAR="…" out of a POSIX sh fragment.
|
||||
func shVar(t *testing.T, script, name string) string {
|
||||
t.Helper()
|
||||
re := regexp.MustCompile(`(?m)^` + regexp.QuoteMeta(name) + `="([^"]*)"`)
|
||||
m := re.FindStringSubmatch(script)
|
||||
if m == nil {
|
||||
t.Fatalf("scripts/router-tags.sh: %s=\"…\" not found (single line, double quotes)", name)
|
||||
}
|
||||
return m[1]
|
||||
}
|
||||
|
||||
// routerTagSet returns the shipped tag set as a set, read from the ONE file that
|
||||
// defines it.
|
||||
func routerTagSet(t *testing.T) map[string]bool {
|
||||
t.Helper()
|
||||
set := map[string]bool{}
|
||||
for _, tag := range strings.Split(shVar(t, repoFile(t, "scripts/router-tags.sh"), "SHATER_ROUTER_TAGS"), ",") {
|
||||
if tag = strings.TrimSpace(tag); tag != "" {
|
||||
set[tag] = true
|
||||
}
|
||||
}
|
||||
if len(set) == 0 {
|
||||
t.Fatal("SHATER_ROUTER_TAGS is empty")
|
||||
}
|
||||
return set
|
||||
}
|
||||
|
||||
// TestRouterTagSetCoversDeclaredFeatures is THE guard the 2026-07-25 WireGuard
|
||||
// outage was missing (D23): it reads the tag set the router binary is actually
|
||||
// built with and fails if a feature we DECLARE supported has lost the build tag
|
||||
// it needs to run.
|
||||
//
|
||||
// It deliberately needs no build tags, no Linux, no privileges and no network,
|
||||
// so it runs in every plain `go test ./...` — including on the Windows dev host,
|
||||
// where nothing else can exercise the shipped configuration. The behavioural
|
||||
// half (does the shipped combination actually CONSTRUCT?) is
|
||||
// shater/generate.TestShippedTagSetConstructsDeclaredProtocols, run with this
|
||||
// same set by scripts/check-router-tags.sh.
|
||||
func TestRouterTagSetCoversDeclaredFeatures(t *testing.T) {
|
||||
shipped := routerTagSet(t)
|
||||
|
||||
for _, f := range Features {
|
||||
var missing []string
|
||||
for _, tag := range f.Tags {
|
||||
if !shipped[tag] {
|
||||
missing = append(missing, tag)
|
||||
}
|
||||
}
|
||||
if len(missing) > 0 {
|
||||
t.Errorf("the shipped router binary would NOT support a feature we declare.\n"+
|
||||
" feature : %s\n"+
|
||||
" declared: %s\n"+
|
||||
" missing : %s (not in SHATER_ROUTER_TAGS, scripts/router-tags.sh)\n"+
|
||||
" why : %s\n"+
|
||||
"Either add the tag back, or stop declaring the feature — those are the only two honest options.",
|
||||
f.Name, f.Declared, strings.Join(missing, ", "), f.Why)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestFeatureTagsAreTracked keeps Features honest: every tag it names must have
|
||||
// a tag_*.go detector, or Compiled()/MissingTags() would report it absent even
|
||||
// when it is compiled in — and the behavioural test would silently SKIP the
|
||||
// feature instead of checking it. A false green is worse than a red.
|
||||
func TestFeatureTagsAreTracked(t *testing.T) {
|
||||
for _, f := range Features {
|
||||
for _, tag := range f.Tags {
|
||||
if !Tracked(tag) {
|
||||
t.Errorf("feature %q requires tag %q, which has no detector: add shater/buildtags/tag_%s.go and the entry in the tracked map", f.Name, tag, tag)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestTrackedTagsHaveDetectorFiles pairs the tracked map with the files on disk,
|
||||
// so a renamed/deleted detector cannot quietly make a tag read as absent.
|
||||
func TestTrackedTagsHaveDetectorFiles(t *testing.T) {
|
||||
for _, tag := range TrackedTags() {
|
||||
name := "tag_" + tag + ".go"
|
||||
body, err := os.ReadFile(name)
|
||||
if err != nil {
|
||||
t.Errorf("tracked tag %q has no detector file %s: %v", tag, name, err)
|
||||
continue
|
||||
}
|
||||
if !strings.Contains(string(body), "//go:build "+tag) || !strings.Contains(string(body), `mark("`+tag+`")`) {
|
||||
t.Errorf("%s must be `//go:build %s` and call mark(%q)", name, tag, tag)
|
||||
}
|
||||
}
|
||||
files, err := filepath.Glob("tag_*.go")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, f := range files {
|
||||
tag := strings.TrimSuffix(strings.TrimPrefix(f, "tag_"), ".go")
|
||||
if !Tracked(tag) {
|
||||
t.Errorf("detector %s exists but %q is not in the tracked map", f, tag)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestBuildScriptUsesTheSharedTagSet stops the split that caused the outage from
|
||||
// coming back: the ship build must SOURCE scripts/router-tags.sh, not carry its
|
||||
// own copy of the tag list. A second copy is a second truth, and the second one
|
||||
// is the one nobody checks.
|
||||
func TestBuildScriptUsesTheSharedTagSet(t *testing.T) {
|
||||
build := repoFile(t, "scripts/build-shaterd.sh")
|
||||
if !strings.Contains(build, "router-tags.sh") {
|
||||
t.Fatal("scripts/build-shaterd.sh must source scripts/router-tags.sh")
|
||||
}
|
||||
if regexp.MustCompile(`(?m)^\s*ROUTER_TAGS="with_`).MatchString(build) {
|
||||
t.Fatal("scripts/build-shaterd.sh re-inlines a literal tag list; the set must come from scripts/router-tags.sh only")
|
||||
}
|
||||
}
|
||||
|
||||
// TestRouterLdflagsSatisfyTagRequirements: `badlinkname` is not self-contained —
|
||||
// the LINK step fails without -checklinkname=0. The flag therefore belongs to
|
||||
// the tag set, and lives beside it; assert the pair never separates.
|
||||
func TestRouterLdflagsSatisfyTagRequirements(t *testing.T) {
|
||||
script := repoFile(t, "scripts/router-tags.sh")
|
||||
ldflags := shVar(t, script, "SHATER_ROUTER_LDFLAGS")
|
||||
if routerTagSet(t)["badlinkname"] && !strings.Contains(ldflags, "-checklinkname=0") {
|
||||
t.Fatalf("SHATER_ROUTER_TAGS carries badlinkname but SHATER_ROUTER_LDFLAGS (%q) lacks -checklinkname=0: the build will fail at link time", ldflags)
|
||||
}
|
||||
if !strings.Contains(repoFile(t, "scripts/build-shaterd.sh"), "SHATER_ROUTER_LDFLAGS") {
|
||||
t.Fatal("scripts/build-shaterd.sh must use $SHATER_ROUTER_LDFLAGS, not a hand-copied -checklinkname=0")
|
||||
}
|
||||
}
|
||||
|
||||
// TestCompiledTagsMatchTheShippedSet proves the DETECTORS are telling the truth:
|
||||
// when the test binary is compiled with exactly the shipped tag set, Compiled()
|
||||
// must equal that set (restricted to tracked tags). Without this, a typo'd or
|
||||
// deleted detector would make the behavioural test skip a protocol and pass.
|
||||
//
|
||||
// It only runs under scripts/check-router-tags.sh (which compiles with that very
|
||||
// set and exports SHATER_ROUTER_TAG_CHECK=1); a plain `go test ./...` compiles
|
||||
// with no tags at all, where the comparison is meaningless.
|
||||
func TestCompiledTagsMatchTheShippedSet(t *testing.T) {
|
||||
if os.Getenv("SHATER_ROUTER_TAG_CHECK") != "1" {
|
||||
t.Skip("not a router-tag-set run; use scripts/check-router-tags.sh")
|
||||
}
|
||||
var want []string
|
||||
for tag := range routerTagSet(t) {
|
||||
if Tracked(tag) {
|
||||
want = append(want, tag)
|
||||
}
|
||||
}
|
||||
sort.Strings(want)
|
||||
got := Compiled()
|
||||
if strings.Join(got, ",") != strings.Join(want, ",") {
|
||||
t.Fatalf("compiled tags do not match the shipped set\n compiled: %v\n shipped : %v\n"+
|
||||
"Either the build ran with the wrong -tags, or a tag_*.go detector is broken.", got, want)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build badlinkname
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the badlinkname build tag — see buildtags.go. There is no !badlinkname twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("badlinkname") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build tfogo_checklinkname0
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the tfogo_checklinkname0 build tag — see buildtags.go. There is no !tfogo_checklinkname0 twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("tfogo_checklinkname0") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_awg
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_awg build tag — see buildtags.go. There is no !with_awg twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_awg") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_gvisor
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_gvisor build tag — see buildtags.go. There is no !with_gvisor twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_gvisor") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_lx_command
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_lx_command build tag — see buildtags.go. There is no !with_lx_command twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_lx_command") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_quic
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_quic build tag — see buildtags.go. There is no !with_quic twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_quic") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_utls
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_utls build tag — see buildtags.go. There is no !with_utls twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_utls") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_wireguard
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_wireguard build tag — see buildtags.go. There is no !with_wireguard twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_wireguard") }
|
||||
@@ -0,0 +1,7 @@
|
||||
//go:build with_xhttp
|
||||
|
||||
package buildtags
|
||||
|
||||
// Detector for the with_xhttp build tag — see buildtags.go. There is no !with_xhttp twin:
|
||||
// an absent detector means "not compiled in", which is exactly the truth.
|
||||
func init() { mark("with_xhttp") }
|
||||
@@ -68,6 +68,12 @@ type nodeView struct {
|
||||
// ParseError is the reason the engine will SKIP this node, verbatim from
|
||||
// parse.ParseShareLink. Empty on every usable node.
|
||||
ParseError string `json:"parse_error,omitempty"`
|
||||
// ParseWarnings lists what the share link asked for that the engine cannot
|
||||
// do (hysteria2 port hopping, tuic congestion control, …). The node IS
|
||||
// usable — that is the difference from parse_error — but it does not behave
|
||||
// exactly as its link describes, and that gap belongs on screen rather than
|
||||
// in a code comment.
|
||||
ParseWarnings []string `json:"parse_warnings,omitempty"`
|
||||
}
|
||||
|
||||
// readModel is the model source. A package var so the tests can exercise the
|
||||
@@ -111,6 +117,7 @@ func nodeViews(m *model.Model) []nodeView {
|
||||
}
|
||||
v.Server = p.Server
|
||||
v.Port = p.Port
|
||||
v.ParseWarnings = p.Warnings
|
||||
} else if n.URI != "" {
|
||||
v.ParseError = err.Error()
|
||||
}
|
||||
|
||||
+126
-31
@@ -61,6 +61,25 @@ type Engine struct {
|
||||
current option.Options // options the running Box was built from
|
||||
hash string // stable hash of current (hex sha256 of canonical JSON)
|
||||
|
||||
// instanceCancel cancels the context THIS box was built on — its own
|
||||
// cancellable child of e.ctx, one per box (see newBox). box.New does not
|
||||
// derive a cancellable context of its own, so without this every goroutine
|
||||
// inside a retired box that waits on ctx.Done() waits forever. Upstream's
|
||||
// runner does exactly this and calls cancel before Close
|
||||
// (cmd/sing-box/cmd_run.go). nil when nothing is running.
|
||||
instanceCancel context.CancelFunc
|
||||
// instanceGen is the 1-based sequence number of the running box, bumped on
|
||||
// every adopted swap. It is what names a generation in the log and in
|
||||
// PendingCloses when a shutdown overruns its budget.
|
||||
instanceGen uint64
|
||||
|
||||
// pending holds the retirements in flight (see teardown.go). Its OWN leaf
|
||||
// lock, deliberately not mu: PendingCloses is read by the status path, and
|
||||
// the moment that read matters most is while an apply is holding mu waiting
|
||||
// out a shutdown that will not finish.
|
||||
pendingMu sync.Mutex
|
||||
pending []*pendingClose
|
||||
|
||||
// defaultLogWriter, when non-nil, is handed to EVERY box.New this engine
|
||||
// performs (box.Options.DefaultLogWriter): the daemon points it at its
|
||||
// long-lived logsink once, and each Apply-swapped box then logs into that
|
||||
@@ -174,6 +193,17 @@ func (e *Engine) Apply(opts option.Options) (changed bool, err error) {
|
||||
}
|
||||
|
||||
func (e *Engine) applyLocked(opts option.Options) (bool, error) {
|
||||
// Stand down the self-check of urltest groups no rule reaches (see
|
||||
// selfcheck.go for the whole argument). This MUTATES opts in place — the
|
||||
// option structs are pointers behind an `any` — and it must run BEFORE the
|
||||
// hash below, so the hash describes the config that is really built: a rule
|
||||
// change that flips a group used<->unused is then a real change that
|
||||
// triggers a swap, and an unchanged config hashes identically on every
|
||||
// reconcile because the stand-down is deterministic.
|
||||
if stood := standDownUnusedSelfCheck(opts); stood > 0 && e.log != nil {
|
||||
e.log.Info("apply: stood down self-check on ", stood, " unused urltest group(s); the observatory is their only prober")
|
||||
}
|
||||
|
||||
newHash, err := e.hashOptions(opts)
|
||||
if err != nil {
|
||||
return false, E.Cause(err, "hash options")
|
||||
@@ -184,12 +214,18 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
|
||||
return false, nil
|
||||
}
|
||||
|
||||
// (2) build + validate. box.New constructs and validates every adapter.
|
||||
nb, err := box.New(box.Options{
|
||||
Context: e.ctx,
|
||||
Options: opts,
|
||||
DefaultLogWriter: e.defaultLogWriter,
|
||||
})
|
||||
// (1b) Do not stack generations. A superseded box whose shutdown overran its
|
||||
// budget is still holding its outbound connections and its WireGuard devices;
|
||||
// building another one on top of it is how one process ends up carrying four
|
||||
// live generations. Give any abandoned teardown one more budget to finish
|
||||
// BEFORE we create anything (see teardown.go awaitAbandonedLocked). Placed
|
||||
// after the hash gate on purpose: a no-op reconcile — cron, every minute —
|
||||
// must stay free.
|
||||
stuck := e.awaitAbandonedLocked()
|
||||
|
||||
// (2) build + validate on its OWN cancellable context. box.New constructs and
|
||||
// validates every adapter.
|
||||
nb, nbCancel, err := e.newBox(opts)
|
||||
if err != nil {
|
||||
// Validation failed: keep the running instance, do not swap.
|
||||
return false, E.Cause(err, "create instance")
|
||||
@@ -211,8 +247,15 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
|
||||
// close-old-then-start-new DIRECTLY — skipping the stall. (box.New above
|
||||
// already validated opts, so we never tear the old box down for a config
|
||||
// that would fail to build.)
|
||||
if e.instance != nil && sharesCacheFileLock(e.current, opts) {
|
||||
return e.closeOldThenStart(nb, opts, newHash)
|
||||
//
|
||||
// A generation that is STILL abandoned after the wait above forces the same
|
||||
// path for a different reason: start-new-first would put a second LIVE box
|
||||
// alongside a third that refuses to die, all three contending for the same
|
||||
// tproxy port, the same cache file and — the expensive one — the same
|
||||
// WireGuard private keys. Closing the current box first keeps the process to
|
||||
// at most one live instance plus the abandoned one.
|
||||
if e.instance != nil && (stuck > 0 || sharesCacheFileLock(e.current, opts)) {
|
||||
return e.closeOldThenStart(nb, nbCancel, opts, newHash)
|
||||
}
|
||||
|
||||
// (3b) start-new-first (zero-downtime when there is no resource conflict).
|
||||
@@ -223,25 +266,59 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
|
||||
// close-old-then-start-new then. Any OTHER Start failure keeps today's
|
||||
// behavior: discard the new box, keep the old one running.
|
||||
if e.instance == nil || !isSwapConflict(err) {
|
||||
nbCancel()
|
||||
_ = nb.Close()
|
||||
return false, E.Cause(err, "start instance")
|
||||
}
|
||||
return e.closeOldThenStart(nb, opts, newHash)
|
||||
return e.closeOldThenStart(nb, nbCancel, opts, newHash)
|
||||
}
|
||||
|
||||
// Swap succeeded. The previously running config becomes last-good.
|
||||
old := e.instance
|
||||
if old != nil {
|
||||
if e.instance != nil {
|
||||
e.lastGood = e.current
|
||||
e.hasLastGood = true
|
||||
_ = old.Close()
|
||||
// Retire the old generation under the hard budget. Its error — including
|
||||
// ErrCloseTimeout — is deliberately NOT returned: the new box is started
|
||||
// and carrying traffic, so this apply SUCCEEDED, and failing it here would
|
||||
// abort the caller's netplane stage and leave a stale ruleset loaded over a
|
||||
// perfectly healthy engine. An abandoned generation is surfaced through
|
||||
// PendingCloses() (critical warning in `shaterd status` and the panel) and
|
||||
// an ERROR line naming it — visible, but not mistaken for a failed apply.
|
||||
_ = e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, e.hash)
|
||||
}
|
||||
e.instance = nb
|
||||
e.current = opts
|
||||
e.hash = newHash
|
||||
e.adoptLocked(nb, nbCancel, opts, newHash)
|
||||
return true, nil
|
||||
}
|
||||
|
||||
// newBox builds a box on its OWN cancellable child of the engine context and
|
||||
// returns the cancel alongside it. Every goroutine the box starts inherits that
|
||||
// context, so cancelling it is what unwinds the ones Close does not reach; see
|
||||
// teardown.go for why the shared, never-cancelled context was the defect.
|
||||
func (e *Engine) newBox(opts option.Options) (*box.Box, context.CancelFunc, error) {
|
||||
ctx, cancel := context.WithCancel(e.ctx)
|
||||
b, err := box.New(box.Options{
|
||||
Context: ctx,
|
||||
Options: opts,
|
||||
DefaultLogWriter: e.defaultLogWriter,
|
||||
})
|
||||
if err != nil {
|
||||
cancel()
|
||||
return nil, nil, err
|
||||
}
|
||||
return b, cancel, nil
|
||||
}
|
||||
|
||||
// adoptLocked installs a started box as THE running instance and gives it the
|
||||
// next generation number. Caller holds e.mu and has already retired whatever was
|
||||
// running before.
|
||||
func (e *Engine) adoptLocked(b *box.Box, cancel context.CancelFunc, opts option.Options, hash string) {
|
||||
e.instanceGen++
|
||||
e.instance = b
|
||||
e.instanceCancel = cancel
|
||||
e.current = opts
|
||||
e.hash = hash
|
||||
}
|
||||
|
||||
// closeOldThenStart is the close-old-then-start-new swap. It is taken both
|
||||
// proactively (the incoming config shares the running box's cache_file lock, so
|
||||
// start-new-first cannot work) and reactively (start-new-first hit a swap
|
||||
@@ -256,27 +333,36 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
|
||||
// already validated by the caller's box.New, so the rebuild below is expected to
|
||||
// succeed; the restore path guards the rare case it does not. The caller holds
|
||||
// e.mu.
|
||||
func (e *Engine) closeOldThenStart(discard *box.Box, opts option.Options, newHash string) (bool, error) {
|
||||
func (e *Engine) closeOldThenStart(discard *box.Box, discardCancel context.CancelFunc, opts option.Options, newHash string) (bool, error) {
|
||||
// (a) Drop the pre-built box (a Start-failed box cannot be restarted, and the
|
||||
// proactively-built one must not hold the cache_file lock while we rebuild).
|
||||
// It was never adopted, so it is not a generation — close it inline, but
|
||||
// cancel its context first exactly like a retired one.
|
||||
if discardCancel != nil {
|
||||
discardCancel()
|
||||
}
|
||||
if discard != nil {
|
||||
_ = discard.Close()
|
||||
}
|
||||
|
||||
// (b) Free the port + cache_file lock by closing the old instance. Remember
|
||||
// its config so we can restore it if the fresh box cannot come up.
|
||||
// (b) Free the port + cache_file lock by closing the old instance, under the
|
||||
// hard budget. Remember its config so we can restore it if the fresh box
|
||||
// cannot come up. An overrun here is reported by retireLocked and tracked in
|
||||
// PendingCloses; we still proceed, because the resources it was supposed to
|
||||
// free are exactly what the operator is waiting on.
|
||||
prevOpts := e.current
|
||||
prevHash := e.hash
|
||||
prevLastGood := e.lastGood
|
||||
prevHasLastGood := e.hasLastGood
|
||||
_ = e.instance.Close()
|
||||
e.instance = nil
|
||||
_ = e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, prevHash)
|
||||
e.instance, e.instanceCancel = nil, nil
|
||||
|
||||
// (c) Build a FRESH box for opts (the discarded one cannot be reused).
|
||||
nb2, err := box.New(box.Options{Context: e.ctx, Options: opts, DefaultLogWriter: e.defaultLogWriter})
|
||||
nb2, cancel2, err := e.newBox(opts)
|
||||
if err == nil {
|
||||
err = nb2.Start()
|
||||
if err != nil {
|
||||
cancel2()
|
||||
_ = nb2.Close()
|
||||
}
|
||||
}
|
||||
@@ -284,28 +370,25 @@ func (e *Engine) closeOldThenStart(discard *box.Box, opts option.Options, newHas
|
||||
// (d) Success: the old config we just closed becomes last-good.
|
||||
e.lastGood = prevOpts
|
||||
e.hasLastGood = true
|
||||
e.instance = nb2
|
||||
e.current = opts
|
||||
e.hash = newHash
|
||||
e.adoptLocked(nb2, cancel2, opts, newHash)
|
||||
return true, nil
|
||||
}
|
||||
|
||||
// (e) The fresh box could not come up and the old one is already closed —
|
||||
// interception is currently down. Try to RESTORE the previous config so we
|
||||
// do not leave the tunnel dead.
|
||||
rb, rerr := box.New(box.Options{Context: e.ctx, Options: prevOpts, DefaultLogWriter: e.defaultLogWriter})
|
||||
rb, rcancel, rerr := e.newBox(prevOpts)
|
||||
if rerr == nil {
|
||||
rerr = rb.Start()
|
||||
if rerr != nil {
|
||||
rcancel()
|
||||
_ = rb.Close()
|
||||
}
|
||||
}
|
||||
if rerr == nil {
|
||||
// Old config restored: keep current/hash/last-good exactly as they were
|
||||
// (do NOT advance them). Report that opts was not applied.
|
||||
e.instance = rb
|
||||
e.current = prevOpts
|
||||
e.hash = prevHash
|
||||
e.adoptLocked(rb, rcancel, prevOpts, prevHash)
|
||||
e.lastGood = prevLastGood
|
||||
e.hasLastGood = prevHasLastGood
|
||||
return false, E.Cause(err, "start instance (config not applied; previous config restored)")
|
||||
@@ -314,7 +397,7 @@ func (e *Engine) closeOldThenStart(discard *box.Box, opts option.Options, newHas
|
||||
// Restore ALSO failed: the engine is now STOPPED. e.instance stays nil (never
|
||||
// pointing at a closed box). The fail-closed nft kill-switch keeps the LAN
|
||||
// safe (no unproxied leak) even though interception is down.
|
||||
e.instance = nil
|
||||
e.instance, e.instanceCancel = nil, nil
|
||||
e.hash = ""
|
||||
return false, E.Cause(E.Errors(err, rerr), "start instance failed and could not restore previous config; engine stopped")
|
||||
}
|
||||
@@ -410,14 +493,26 @@ func (e *Engine) HasLastGood() bool {
|
||||
}
|
||||
|
||||
// Close stops the running instance, if any. It is idempotent.
|
||||
//
|
||||
// Unlike the swap path, this one DOES return ErrCloseTimeout: here the shutdown
|
||||
// is the whole operation, so "it did not stop" is the result, not a footnote. The
|
||||
// caller (apply.Teardown) records it as the teardown's error while still
|
||||
// completing the netplane teardown — the data plane must come down even when a
|
||||
// box will not.
|
||||
func (e *Engine) Close() error {
|
||||
e.mu.Lock()
|
||||
defer e.mu.Unlock()
|
||||
if e.instance == nil {
|
||||
// Nothing running, but a previously abandoned generation may still be:
|
||||
// give it a last budget so a teardown followed by a restart does not carry
|
||||
// the leak across.
|
||||
if e.awaitAbandonedLocked() > 0 {
|
||||
return ErrCloseTimeout
|
||||
}
|
||||
return nil
|
||||
}
|
||||
err := e.instance.Close()
|
||||
e.instance = nil
|
||||
err := e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, e.hash)
|
||||
e.instance, e.instanceCancel = nil, nil
|
||||
e.hash = ""
|
||||
return err
|
||||
}
|
||||
|
||||
+203
-14
@@ -1,6 +1,7 @@
|
||||
package engine
|
||||
|
||||
import (
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"github.com/sagernet/sing-box/adapter"
|
||||
@@ -320,11 +321,15 @@ func parseGroupCopyTag(group, tag string) (member string, ok bool) {
|
||||
return member, true
|
||||
}
|
||||
|
||||
// ChainHealth is one configured chain's reachability, mirroring GroupHealth.Used
|
||||
// for the chain card (plan §5.E): a chain no enabled rule routes through is never
|
||||
// probed — the observatory walks only reachable paths — and the panel renders it
|
||||
// "unused" rather than as a health problem. A chain has no membership counters: it
|
||||
// is a fixed path, and its end-to-end health is the exit test's job, not a roll-up.
|
||||
// ChainHealth is one configured chain's reachability plus its PER-HOP health,
|
||||
// mirroring GroupHealth.Used for the chain card (plan §5.E): a chain no enabled
|
||||
// rule routes through is never probed — the observatory walks only reachable
|
||||
// paths — and the panel renders it "unused" rather than as a health problem.
|
||||
// A chain has no membership counters of its own: it is a fixed path, and its
|
||||
// end-to-end health is the exit probe's job. What it DOES have is hops, and
|
||||
// since the observatory now probes every hop wrapper (probeplan.go walkDetour),
|
||||
// each hop's health is on the board and is projected here so an operator can
|
||||
// see WHICH hop died instead of only that the chain did.
|
||||
type ChainHealth struct {
|
||||
// Name is the chain's model name (config chain "chain:<name>"), what the Targets
|
||||
// page lists and what a rule targets.
|
||||
@@ -336,34 +341,218 @@ type ChainHealth struct {
|
||||
// when the observatory is disabled or not yet configured: no badge is better
|
||||
// than a wrong one (the same rule as GroupHealth.Used).
|
||||
Used bool `json:"used"`
|
||||
// Hops is the per-hop health readout, L1..Ln in wire order (see ChainHopHealth).
|
||||
// It is EMPTY for a chain the running box never materialised: a chain no rule
|
||||
// references is resolved lazily and never built, and a 1-hop chain without an
|
||||
// egress entry resolves straight to its target with no wrapper — in both cases
|
||||
// there are no "chain-<name>-h…" outbounds to project. An absent "hops" key
|
||||
// therefore means "nothing materialised to report on", NEVER "this chain has
|
||||
// no hops" — the model, not this projection, knows how many hops were
|
||||
// configured.
|
||||
Hops []ChainHopHealth `json:"hops,omitempty"`
|
||||
}
|
||||
|
||||
// ChainHealth reports the reachability (used/unused) of every named chain, one row
|
||||
// per name, in the order given. It is the chain analogue of GroupHealth.Used: a
|
||||
// chain the observatory's used-set does not cover is reported Used=false so the
|
||||
// panel can mark it "unused" instead of running an exit test against a path nothing
|
||||
// routes through.
|
||||
// ChainHopHealth is one hop of one materialised chain, as the health board saw
|
||||
// it — a projection, like everything in this file: nothing here dials.
|
||||
//
|
||||
// A NODE hop is a single measurement: the observatory dials the hop wrapper
|
||||
// "chain-<name>-h<i>", which pulls exactly the path prefix up to and including
|
||||
// this hop, so Total=1, the counters follow the hop's own state, DelayMs and
|
||||
// AgeSeconds are its own observation, and Selected is "" (a fixed hop selects
|
||||
// nothing).
|
||||
//
|
||||
// A GROUP hop rolls up its member copies "chain-<name>-h<i>-<member>", each of
|
||||
// which the observatory probes through its own prefix of the chain. The
|
||||
// counters obey the same invariants as GroupHealth — Tested == Alive+Dead and
|
||||
// Alive+Dead+Untested == Total — so the panel needs no arithmetic of its own.
|
||||
// State summarises them: "alive" when at least one member is alive (the hop can
|
||||
// carry traffic), "dead" when at least one was tested and none is alive (a
|
||||
// positive finding of a dead hop), "untested" when nothing was tested. Selected
|
||||
// is the node NAME the wrapper currently picks; DelayMs/AgeSeconds are the
|
||||
// SELECTED member's observation, or the freshest ALIVE member's when the
|
||||
// selection has no measurement of its own — the number shown must always be a
|
||||
// measurement somebody took, never an average nobody did.
|
||||
type ChainHopHealth struct {
|
||||
Index int `json:"index"` // 1-based position on the wire, L1..Ln
|
||||
Tag string `json:"tag"` // "chain-<name>-h<i>" — the wrapper actually dialled
|
||||
Kind string `json:"kind"` // "node" | "group"
|
||||
Exit bool `json:"exit"` // the LAST hop: where traffic leaves to the internet
|
||||
State string `json:"state"` // "alive" | "dead" | "untested"
|
||||
DelayMs int `json:"delay_ms"`
|
||||
AgeSeconds int64 `json:"age_seconds"` // -1 when unknown
|
||||
Selected string `json:"selected"` // group hop: the node NAME it currently selects; "" otherwise
|
||||
Total int `json:"total"`
|
||||
Tested int `json:"tested"`
|
||||
Alive int `json:"alive"`
|
||||
Dead int `json:"dead"`
|
||||
Untested int `json:"untested"`
|
||||
}
|
||||
|
||||
// ChainHealth reports the reachability (used/unused) of every named chain plus
|
||||
// the per-hop health of each one the running box materialised, one row per
|
||||
// name, in the order given. The Used half is the chain analogue of
|
||||
// GroupHealth.Used: a chain the observatory's used-set does not cover is
|
||||
// reported Used=false so the panel can mark it "unused" instead of a health
|
||||
// readout against a path nothing routes through.
|
||||
//
|
||||
// names come from the desired-state model, NOT the running box: a chain no rule
|
||||
// references is never materialised (generate/chain.go resolveChain is lazy), so it
|
||||
// is invisible to a box-only enumeration — yet the panel lists it from the config
|
||||
// and must be able to badge it. The engine supplies the only fact a box read can
|
||||
// add here, the observatory's published used-set. Pure apart from that read; nil
|
||||
// names or a stopped engine (nil used-set) yield an empty/used-everything result.
|
||||
// and must be able to badge it. The engine supplies what only it can: the
|
||||
// observatory's published used-set, and the running box's outbound/endpoint
|
||||
// pool the hop projection reads. Still no dialling anywhere; nil names or a
|
||||
// stopped engine (nil used-set, empty pool) yield an empty/used-everything
|
||||
// result with no hops.
|
||||
func (e *Engine) ChainHealth(names []string) []ChainHealth {
|
||||
out := make([]ChainHealth, 0, len(names))
|
||||
used := e.observatoryUsed()
|
||||
usedChains := usedChainNames(used)
|
||||
pool := e.runningPool()
|
||||
view := e.HealthView()
|
||||
for _, name := range names {
|
||||
name = strings.TrimSpace(name)
|
||||
if name == "" {
|
||||
continue
|
||||
}
|
||||
out = append(out, ChainHealth{Name: name, Used: used == nil || usedChains[name]})
|
||||
out = append(out, ChainHealth{
|
||||
Name: name,
|
||||
Used: used == nil || usedChains[name],
|
||||
Hops: chainHopHealthOf(pool, name, view),
|
||||
})
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// runningPool snapshots the running box's outbounds AND endpoints into one
|
||||
// list. The endpoints matter: a chain hop rebuilt from a wireguard/AmneziaWG
|
||||
// node is an ENDPOINT copy, invisible in Outbounds() — the same trap
|
||||
// groupTargets documents — and a hop projection that missed it would silently
|
||||
// drop the very hop this feature exists to localise (the production chain's
|
||||
// first hop IS an AWG endpoint). Empty (never nil-unsafe) on a stopped engine.
|
||||
func (e *Engine) runningPool() []adapter.Outbound {
|
||||
var pool []adapter.Outbound
|
||||
inst := e.Instance()
|
||||
if inst == nil {
|
||||
return pool
|
||||
}
|
||||
if om := inst.Outbound(); om != nil {
|
||||
pool = append(pool, om.Outbounds()...)
|
||||
}
|
||||
if em := inst.Endpoint(); em != nil {
|
||||
for _, ep := range em.Endpoints() {
|
||||
pool = append(pool, ep)
|
||||
}
|
||||
}
|
||||
return pool
|
||||
}
|
||||
|
||||
// chainHopHealthOf projects one chain's hop wrappers out of an outbound pool
|
||||
// against one health view. Pure apart from the view reads — the same
|
||||
// unit-testing contract as groupHealthOf: hand it a fake pool and a hand-built
|
||||
// store and every branch is reachable without a box.
|
||||
//
|
||||
// A pool entry belongs to chain <name> when its tag is exactly
|
||||
// "chain-<name>-h<digits>" — the hop WRAPPER the observatory dials. Member
|
||||
// copies ("chain-<name>-h<i>-<member>") are not hops themselves; they are
|
||||
// reached through the wrapper's own member list (adapter.OutboundGroup.All), so
|
||||
// the roll-up sees exactly what the wrapper can select, in its order.
|
||||
func chainHopHealthOf(pool []adapter.Outbound, name string, view HealthView) []ChainHopHealth {
|
||||
prefix := "chain-" + name + "-h"
|
||||
var hops []ChainHopHealth
|
||||
for _, ob := range pool {
|
||||
tag := ob.Tag()
|
||||
rest, ok := strings.CutPrefix(tag, prefix)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
idx, ok := parseAllDigits(rest)
|
||||
if !ok {
|
||||
continue // a member copy, or another chain sharing the prefix
|
||||
}
|
||||
if g, isGroup := ob.(adapter.OutboundGroup); isGroup {
|
||||
hops = append(hops, chainGroupHop(g, name, idx, view))
|
||||
} else {
|
||||
hops = append(hops, chainNodeHop(tag, idx, view))
|
||||
}
|
||||
}
|
||||
sort.Slice(hops, func(i, j int) bool { return hops[i].Index < hops[j].Index })
|
||||
if len(hops) > 0 {
|
||||
// The largest index is the exit — the wrapper whose probe leaves to the
|
||||
// internet. Marked after sorting so the flag cannot depend on pool order.
|
||||
hops[len(hops)-1].Exit = true
|
||||
}
|
||||
return hops
|
||||
}
|
||||
|
||||
// chainNodeHop is the one-measurement hop: the wrapper itself was dialled by
|
||||
// the observatory, so its own board state IS the hop's health and the counters
|
||||
// degenerate to whichever bucket that state fills.
|
||||
func chainNodeHop(tag string, idx int, view HealthView) ChainHopHealth {
|
||||
hop := ChainHopHealth{Index: idx, Tag: tag, Kind: "node", Total: 1}
|
||||
hop.State, hop.DelayMs, hop.AgeSeconds = view.State(tag)
|
||||
switch hop.State {
|
||||
case HealthAlive:
|
||||
hop.Alive = 1
|
||||
case HealthDead:
|
||||
hop.Dead = 1
|
||||
default:
|
||||
hop.Untested = 1
|
||||
}
|
||||
hop.Tested = hop.Alive + hop.Dead
|
||||
return hop
|
||||
}
|
||||
|
||||
// chainGroupHop rolls a group hop up over its member copies. The counters carry
|
||||
// the GroupHealth invariants; the summary State answers the only question a hop
|
||||
// row asks — "can this hop carry the chain": alive while anything answers,
|
||||
// dead only on a positive all-tested-dead finding, untested when nothing is
|
||||
// known (never dead-by-absence, the same honesty rule as everywhere else).
|
||||
func chainGroupHop(g adapter.OutboundGroup, chain string, idx int, view HealthView) ChainHopHealth {
|
||||
hop := ChainHopHealth{Index: idx, Tag: g.Tag(), Kind: "group", AgeSeconds: -1}
|
||||
selected := g.Now()
|
||||
if selected != "" {
|
||||
hop.Selected = chainMemberName(selected, chain)
|
||||
}
|
||||
|
||||
// The number a hop row shows must be a real observation: the selected
|
||||
// member's when it has one, else the freshest alive member's.
|
||||
freshDelay, freshAge := 0, int64(-1)
|
||||
selDelay, selAge := 0, int64(-1)
|
||||
for _, tag := range g.All() {
|
||||
state, delayMs, age := view.State(tag)
|
||||
switch state {
|
||||
case HealthAlive:
|
||||
hop.Alive++
|
||||
if age >= 0 && (freshAge < 0 || age < freshAge) {
|
||||
freshDelay, freshAge = delayMs, age
|
||||
}
|
||||
case HealthDead:
|
||||
hop.Dead++
|
||||
default:
|
||||
hop.Untested++
|
||||
}
|
||||
hop.Total++
|
||||
if tag == selected && age >= 0 {
|
||||
selDelay, selAge = delayMs, age
|
||||
}
|
||||
}
|
||||
hop.Tested = hop.Alive + hop.Dead
|
||||
switch {
|
||||
case hop.Alive > 0:
|
||||
hop.State = HealthAlive
|
||||
case hop.Tested > 0:
|
||||
hop.State = HealthDead
|
||||
default:
|
||||
hop.State = HealthUntested
|
||||
}
|
||||
if selAge >= 0 {
|
||||
hop.DelayMs, hop.AgeSeconds = selDelay, selAge
|
||||
} else {
|
||||
hop.DelayMs, hop.AgeSeconds = freshDelay, freshAge
|
||||
}
|
||||
return hop
|
||||
}
|
||||
|
||||
// usedChainNames recovers the set of chain NAMES the observatory's used-set covers.
|
||||
//
|
||||
// The used-set is keyed by outbound TAG (the generator's schema, materialised in
|
||||
|
||||
@@ -309,8 +309,145 @@ func TestChainHealthNilUsedSet(t *testing.T) {
|
||||
t.Fatalf("ChainHealth = %+v, want %+v (nil used-set ⇒ all used, blanks dropped)", got, want)
|
||||
}
|
||||
for i, w := range want {
|
||||
if got[i] != w {
|
||||
if got[i].Name != w.Name || got[i].Used != w.Used {
|
||||
t.Errorf("ChainHealth[%d] = %+v, want %+v", i, got[i], w)
|
||||
}
|
||||
// A stopped engine materialised nothing: hops must be empty, and the
|
||||
// contract says empty means "nothing materialised", not "no hops".
|
||||
if len(got[i].Hops) != 0 {
|
||||
t.Errorf("ChainHealth[%d].Hops = %+v, want empty on a stopped engine", i, got[i].Hops)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestChainHopHealthProjection drives the per-hop readout over a 3-hop chain —
|
||||
// node hop L1, group hop L2, group hop L3 (the exit) — with exactly ONE hop's
|
||||
// members dead. The dead state must land on THAT hop's index and nowhere else,
|
||||
// the counters must obey the GroupHealth invariants on every hop, and the exit
|
||||
// flag must sit on the largest index. This is the "which hop died" question the
|
||||
// whole hop surface exists to answer.
|
||||
func TestChainHopHealthProjection(t *testing.T) {
|
||||
hist := urltest.NewHistoryStorage()
|
||||
now := time.Now()
|
||||
|
||||
// L1 (node hop wrapper): alive — the prefix up to hop 1 works.
|
||||
hist.StoreURLTestHistory("chain-c-h1", &adapter.URLTestHistory{LastOK: now.Add(-5 * time.Second), Delay: 40})
|
||||
// L2 (group hop): BOTH member copies dead — this is the hop that died.
|
||||
hist.StoreURLTestHistory("chain-c-h2-x", &adapter.URLTestHistory{LastFail: now.Add(-3 * time.Second)})
|
||||
hist.StoreURLTestHistory("chain-c-h2-y", &adapter.URLTestHistory{LastFail: now.Add(-2 * time.Second)})
|
||||
// L3 (group hop, exit): one member alive, one never measured.
|
||||
hist.StoreURLTestHistory("chain-c-h3-a", &adapter.URLTestHistory{LastOK: now.Add(-7 * time.Second), Delay: 200})
|
||||
// chain-c-h3-b: nothing at all.
|
||||
|
||||
pool := []adapter.Outbound{
|
||||
&depOutbound{failingOutbound{tag: "chain-c-h1"}, nil},
|
||||
&depGroup{fakeGroup{
|
||||
tag: "chain-c-h2", kind: C.TypeSelector,
|
||||
all: []string{"chain-c-h2-x", "chain-c-h2-y"},
|
||||
now: "chain-c-h2-x",
|
||||
}, []string{"chain-c-h2-x", "chain-c-h2-y"}},
|
||||
&depGroup{fakeGroup{
|
||||
tag: "chain-c-h3", kind: C.TypeSelector,
|
||||
all: []string{"chain-c-h3-a", "chain-c-h3-b"},
|
||||
now: "chain-c-h3-a",
|
||||
}, []string{"chain-c-h3-a", "chain-c-h3-b"}},
|
||||
// Noise the projection must ignore: a member copy is not a hop, another
|
||||
// chain's wrapper is not this chain's.
|
||||
&depOutbound{failingOutbound{tag: "chain-c-h2-x"}, nil},
|
||||
&depOutbound{failingOutbound{tag: "chain-other-h1"}, nil},
|
||||
}
|
||||
|
||||
hops := chainHopHealthOf(pool, "c", newHealthView(hist, healthTTLFloor, now))
|
||||
if len(hops) != 3 {
|
||||
t.Fatalf("got %d hops, want 3: %+v", len(hops), hops)
|
||||
}
|
||||
|
||||
// Ordered by index, exit on the largest.
|
||||
for i, wantIdx := range []int{1, 2, 3} {
|
||||
if hops[i].Index != wantIdx {
|
||||
t.Fatalf("hops out of order: %+v", hops)
|
||||
}
|
||||
if got, want := hops[i].Exit, wantIdx == 3; got != want {
|
||||
t.Errorf("hop %d Exit = %v, want %v", wantIdx, got, want)
|
||||
}
|
||||
}
|
||||
|
||||
h1, h2, h3 := hops[0], hops[1], hops[2]
|
||||
|
||||
// L1: one measurement, its own numbers, no selection.
|
||||
if h1.Kind != "node" || h1.State != HealthAlive || h1.Total != 1 || h1.Alive != 1 ||
|
||||
h1.DelayMs != 40 || h1.AgeSeconds != 5 || h1.Selected != "" {
|
||||
t.Errorf("h1 = %+v, want an alive node hop with its own 40ms/5s and no selection", h1)
|
||||
}
|
||||
|
||||
// L2: the dead hop. Both members tested, none alive => a POSITIVE dead
|
||||
// finding on exactly this index. The selected member (x) is dead, and a
|
||||
// dead observation carries delay 0 with the failure's age.
|
||||
if h2.Kind != "group" || h2.State != HealthDead {
|
||||
t.Fatalf("h2 = %+v, want the DEAD group hop — this is the answer to 'which hop died'", h2)
|
||||
}
|
||||
if h2.Total != 2 || h2.Tested != 2 || h2.Alive != 0 || h2.Dead != 2 || h2.Untested != 0 {
|
||||
t.Errorf("h2 counters = %+v, want 2 tested / 2 dead", h2)
|
||||
}
|
||||
if h2.Selected != "x" {
|
||||
t.Errorf("h2 Selected = %q, want the node NAME x", h2.Selected)
|
||||
}
|
||||
if h2.DelayMs != 0 || h2.AgeSeconds != 3 {
|
||||
t.Errorf("h2 delay/age = %d/%d, want 0/3 (the selected member's failure observation)", h2.DelayMs, h2.AgeSeconds)
|
||||
}
|
||||
|
||||
// L3: alive (one member answers), one member honestly untested — never
|
||||
// folded into dead. The selected member is the alive one, so its numbers show.
|
||||
if h3.Kind != "group" || h3.State != HealthAlive {
|
||||
t.Fatalf("h3 = %+v, want an alive exit hop", h3)
|
||||
}
|
||||
if h3.Total != 2 || h3.Tested != 1 || h3.Alive != 1 || h3.Dead != 0 || h3.Untested != 1 {
|
||||
t.Errorf("h3 counters = %+v, want 1 alive / 1 untested", h3)
|
||||
}
|
||||
if h3.Selected != "a" || h3.DelayMs != 200 || h3.AgeSeconds != 7 {
|
||||
t.Errorf("h3 = %+v, want selection a with its 200ms/7s", h3)
|
||||
}
|
||||
|
||||
// The invariants, on every hop, exactly as GroupHealth promises.
|
||||
for _, h := range hops {
|
||||
if h.Tested != h.Alive+h.Dead {
|
||||
t.Errorf("hop %d: Tested = %d, want alive+dead = %d", h.Index, h.Tested, h.Alive+h.Dead)
|
||||
}
|
||||
if h.Alive+h.Dead+h.Untested != h.Total {
|
||||
t.Errorf("hop %d: alive+dead+untested = %d, want total = %d",
|
||||
h.Index, h.Alive+h.Dead+h.Untested, h.Total)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A group hop whose SELECTION has no measurement shows the freshest ALIVE
|
||||
// member's numbers instead — the shown number must always be an observation
|
||||
// somebody took.
|
||||
func TestChainHopSelectionWithoutMeasurementFallsBack(t *testing.T) {
|
||||
hist := urltest.NewHistoryStorage()
|
||||
now := time.Now()
|
||||
hist.StoreURLTestHistory("chain-c-h1-a", &adapter.URLTestHistory{LastOK: now.Add(-30 * time.Second), Delay: 90})
|
||||
hist.StoreURLTestHistory("chain-c-h1-b", &adapter.URLTestHistory{LastOK: now.Add(-4 * time.Second), Delay: 150})
|
||||
// The selection points at c, which has no observation at all.
|
||||
pool := []adapter.Outbound{
|
||||
&depGroup{fakeGroup{
|
||||
tag: "chain-c-h1", kind: C.TypeSelector,
|
||||
all: []string{"chain-c-h1-a", "chain-c-h1-b", "chain-c-h1-c"},
|
||||
now: "chain-c-h1-c",
|
||||
}, nil},
|
||||
}
|
||||
hops := chainHopHealthOf(pool, "c", newHealthView(hist, healthTTLFloor, now))
|
||||
if len(hops) != 1 {
|
||||
t.Fatalf("got %d hops, want 1", len(hops))
|
||||
}
|
||||
h := hops[0]
|
||||
if h.Selected != "c" {
|
||||
t.Errorf("Selected = %q, want c (the selection is reported even unmeasured)", h.Selected)
|
||||
}
|
||||
if h.DelayMs != 150 || h.AgeSeconds != 4 {
|
||||
t.Errorf("delay/age = %d/%d, want the freshest ALIVE member's 150/4", h.DelayMs, h.AgeSeconds)
|
||||
}
|
||||
if h.State != HealthAlive || h.Alive != 2 || h.Untested != 1 {
|
||||
t.Errorf("hop = %+v, want alive with 2 alive / 1 untested", h)
|
||||
}
|
||||
}
|
||||
+228
-59
@@ -12,33 +12,46 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/sagernet/sing-box/adapter"
|
||||
"github.com/sagernet/sing-box/common/urltest"
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
)
|
||||
|
||||
// Exit test — "what am I actually going out through, and how fast" (F2, plan §5.E).
|
||||
// Group/chain test — "ask the observatory to refresh, then report what it
|
||||
// measured, plus the exit address" (F2, plan §5.E; reworked under the one-probe
|
||||
// rule).
|
||||
//
|
||||
// # Why this is not just another node probe
|
||||
// # This file no longer measures latency. On purpose.
|
||||
//
|
||||
// The observatory (observatory.go) answers "which nodes are alive". It cannot
|
||||
// answer the question an operator actually asks after switching a group: *through
|
||||
// which address am I leaving the country right now*. A group is an indirection —
|
||||
// a selector or urltest over many members — so the delay of the group and the
|
||||
// public address it exits from are properties of the CURRENT selection, not of
|
||||
// any node the panel can point at. A CHAIN is the same question over a longer
|
||||
// path: its exit wrapper "chain-<name>-hN" tunnels through every hop, so dialling
|
||||
// it measures the whole L1..Ln path end to end.
|
||||
// It used to: one urltest.URLTest per target, dialled right here. That was the
|
||||
// defect, not a feature. The observatory (observatory.go) probes every path the
|
||||
// routing rules actually use — per-chain member copies, egress-bound group
|
||||
// copies, chain exits — and it is the ONLY thing allowed to dial for health.
|
||||
// A second dial path from this file measured the WRONG thing: pressing "Test"
|
||||
// dialled the base groups directly from the router over the default WAN, a path
|
||||
// no rule routes through, and (worse) URLTest's DialContext touched the group,
|
||||
// arming its own 30-minute probing ticker, which kept writing those false
|
||||
// direct measurements under the members' base tags. For a node that is blocked
|
||||
// on the direct WAN and alive only behind a tunnel, that reading is not stale —
|
||||
// it is FALSE, and it poisoned selection and the panel alike.
|
||||
//
|
||||
// So one test per target (group or chain) measures two things:
|
||||
// So a manual test is now a READ with a refresh request in front of it:
|
||||
//
|
||||
// delay_ms — reusing urltest.URLTest, the SAME primitive the observatory uses
|
||||
// for a node. There is deliberately no second latency mechanism
|
||||
// here: the target is just an outbound, and probing it exercises
|
||||
// exactly the path traffic will take.
|
||||
// exit_ip — a real HTTP request THROUGH the target to a service that echoes
|
||||
// the client address back. Nothing else can produce this number: the
|
||||
// router cannot know its own public address, and the proxy protocol
|
||||
// does not report it.
|
||||
// delay_ms / ok — the observatory's OWN measurement of the target's dial
|
||||
// path. TestGroups rewinds the observatory (RefreshObservatory)
|
||||
// so the numbers are fresh — taken AFTER the button press —
|
||||
// then each target waits for an observation newer than the
|
||||
// run's start instant and reports it verbatim. tested_unix is
|
||||
// the instant that observation was taken, not "now".
|
||||
// exit_ip — a real HTTP request THROUGH the target to a service that
|
||||
// echoes the client address back. This is the ONE connection
|
||||
// this file still opens, and it stays because no probe can
|
||||
// answer it: the router cannot know its own public address,
|
||||
// and the proxy protocol does not report it. It is NOT a
|
||||
// second health probe — it travels the target's own routed
|
||||
// path (the same outbound object the rules dial), and it is
|
||||
// only ever issued for a target that (a) the observatory's
|
||||
// used-set covers and (b) just resolved ALIVE. A target
|
||||
// outside the rules, or one whose path is down, gets an empty
|
||||
// exit_ip — never a connection.
|
||||
//
|
||||
// # The direct-egress trap
|
||||
//
|
||||
@@ -51,16 +64,15 @@ import (
|
||||
// exit_ip instead of somebody else's number.
|
||||
|
||||
const (
|
||||
// groupTestDelayTimeout bounds the latency probe of one group.
|
||||
groupTestDelayTimeout = 5 * time.Second
|
||||
// groupTestExitTimeout bounds the whole exit-address lookup for one group
|
||||
// (connect through the tunnel + TLS + response). Short on purpose: this runs
|
||||
// while a human waits on a panel button, and a slow answer is worth less than a
|
||||
// prompt "could not determine".
|
||||
groupTestExitTimeout = 6 * time.Second
|
||||
// groupTestConcurrency bounds how many groups are tested in parallel. Small:
|
||||
// every one of these opens a real tunnelled connection, and a router with a
|
||||
// handful of groups is the normal case.
|
||||
// groupTestConcurrency bounds how many EXIT-ADDRESS lookups run in parallel.
|
||||
// Small: each one opens a real tunnelled connection, and a router with a
|
||||
// handful of groups is the normal case. The board polling itself is not
|
||||
// bounded by this — it is sleep-and-read, no I/O.
|
||||
groupTestConcurrency = 4
|
||||
// exitBodyLimit caps what is read from the exit-address service. These responses
|
||||
// are a few hundred bytes; anything larger is a hijacked/captive-portal answer
|
||||
@@ -68,19 +80,58 @@ const (
|
||||
exitBodyLimit = 4 << 10
|
||||
)
|
||||
|
||||
// groupTestWaitDeadline / groupTestPollEvery pace the wait for the observatory:
|
||||
// each covered target polls the health board once a second until its measured
|
||||
// tag carries an observation newer than the run's start, giving up after the
|
||||
// deadline. 120s covers a forced pass of a large plan (batches chain
|
||||
// back-to-back during a forced pass, see observatoryTickOnce) with room for the
|
||||
// probe timeouts of a mostly-dead population. Package variables, not constants,
|
||||
// for exactly one reason: the timeout tests must not take two minutes.
|
||||
var (
|
||||
groupTestWaitDeadline = 120 * time.Second
|
||||
groupTestPollEvery = time.Second
|
||||
)
|
||||
|
||||
// The fixed result texts for the targets that are never (or not yet) measured.
|
||||
// They are contract, not decoration — the panel shows them verbatim.
|
||||
const (
|
||||
// groupTestErrNotRouted: the target is outside the observatory's used-set,
|
||||
// so no rule routes through it and nothing measures it. Dialling it anyway
|
||||
// would recreate the false direct measurement this rework removed.
|
||||
groupTestErrNotRouted = "not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use"
|
||||
// groupTestErrNotReached: the deadline passed without a fresh observation.
|
||||
groupTestErrNotReached = "the observatory has not reached this target yet — it refreshes on the global probe interval"
|
||||
// groupTestErrProbingOff: the observatory is disabled (the GroupHealth
|
||||
// master switch), so a refresh request has nothing to wake and waiting for
|
||||
// the deadline would just delay the same answer by two minutes.
|
||||
groupTestErrProbingOff = "background probing is disabled, so there is nothing to measure this target with"
|
||||
// groupTestErrPathDead: the fresh observation exists and it is a FAILURE —
|
||||
// the observatory probed the target's path after the button press and the
|
||||
// path did not answer. An honest negative, not a missing measurement.
|
||||
groupTestErrPathDead = "the observatory's probe through this path failed"
|
||||
)
|
||||
|
||||
// GroupTestResult is one target's test outcome — a group's or a chain's (Group
|
||||
// then carries the chain's model name). The JSON tags are the panel contract —
|
||||
// see the shater API docs for /api/groups/test.
|
||||
// see the shater API docs for /api/groups/test. The field names and types are
|
||||
// FROZEN; what changed in the rework is where the numbers come from.
|
||||
//
|
||||
// DelayMs and OK are a READ of the observatory's measurement of the target's
|
||||
// dial path — not a fresh dial performed by this file. OK is true exactly when
|
||||
// the health board's state for the measured tag is alive; DelayMs is that
|
||||
// observation's RTT; TestedUnix is the instant the OBSERVATION was taken (now
|
||||
// minus its age), so a result honestly says how old its number is instead of
|
||||
// stamping the poll time over it.
|
||||
//
|
||||
// Selected is the group's current pick (OutboundGroup.Now()); for a chain it is
|
||||
// the node NAME the chain's last group hop currently selects, "" when the chain
|
||||
// has no group hop (a fixed path selects nothing).
|
||||
//
|
||||
// OK reports whether the LATENCY measurement succeeded, which is the test's primary
|
||||
// question. A failed exit-address lookup deliberately does NOT clear it: knowing the
|
||||
// target is up and fast is useful on its own, and a probe service being unreachable
|
||||
// says nothing about the tunnel. In that case OK stays true and ExitIP is empty —
|
||||
// "not determined", never a guess and never somebody else's address.
|
||||
// A failed exit-address lookup deliberately does NOT clear OK: knowing the
|
||||
// target is up and fast is useful on its own, and a probe service being
|
||||
// unreachable says nothing about the tunnel. In that case OK stays true and
|
||||
// ExitIP is empty — "not determined", never a guess and never somebody else's
|
||||
// address.
|
||||
type GroupTestResult struct {
|
||||
Group string `json:"group"`
|
||||
Selected string `json:"selected"`
|
||||
@@ -259,19 +310,30 @@ func parseAllDigits(s string) (int, bool) {
|
||||
// and chains (empty/nil = every group and every chain in the running box),
|
||||
// returning started=false when a run is already in flight.
|
||||
//
|
||||
// It is a SINGLETON: a second request while a run is in flight is refused rather
|
||||
// than queued or run in parallel, because these runs open real tunnelled
|
||||
// connections and a panel that double-fires a button must not multiply the load
|
||||
// on the uplink. The observatory checks the same guard and skips its tick while
|
||||
// a run is in flight (observatory.go), so a manual test never competes with
|
||||
// background probing for the uplink.
|
||||
// It does NOT probe. It records the run's start instant, asks the observatory
|
||||
// for an out-of-turn full pass (RefreshObservatory), and then each target waits
|
||||
// for the health board to carry an observation NEWER than that instant on the
|
||||
// tag the observatory actually measures for it — see testOneTarget. The only
|
||||
// connection a run may still open is the exit-address lookup, and only for a
|
||||
// target that resolved alive on a rule-routed path.
|
||||
//
|
||||
// probeURL is the latency-probe URL; "" falls back to urltest's gstatic default.
|
||||
// It is a SINGLETON: a second request while a run is in flight is refused
|
||||
// rather than queued, because two overlapping runs would each rewind the
|
||||
// observatory's cursor and neither pass would ever complete — and the panel's
|
||||
// progress contract assumes one run's counters at a time anyway.
|
||||
//
|
||||
// Apply-swap safety: the target outbounds are snapshotted up front, so a config swap
|
||||
// mid-run cannot change what is being tested. A target torn down mid-run simply fails
|
||||
// its probe and is reported not-ok.
|
||||
// probeURL is accepted and IGNORED. The probe URL is a global observatory
|
||||
// setting now (ObservatoryConfig.ProbeURL, set at apply time); a per-run URL
|
||||
// would mean this run measures something different from what the board holds,
|
||||
// which is exactly the two-instruments split the rework removed. The parameter
|
||||
// stays so the callers (shater/apply, shater/panel) keep compiling and the
|
||||
// control-plane API shape does not churn.
|
||||
//
|
||||
// Apply-swap safety: the target outbounds are snapshotted up front, so a config
|
||||
// swap mid-run cannot change what is being tested. A target torn down mid-run
|
||||
// simply never receives a fresh observation and resolves on the deadline.
|
||||
func (e *Engine) TestGroups(names []string, probeURL string) (started bool) {
|
||||
_ = probeURL // ignored — see the doc comment above
|
||||
if !e.groupTestRunning.CompareAndSwap(false, true) {
|
||||
return false
|
||||
}
|
||||
@@ -285,6 +347,13 @@ func (e *Engine) TestGroups(names []string, probeURL string) (started bool) {
|
||||
// GroupTestStatus for why the scope is a set of names.
|
||||
e.setGroupTestScope(scopeOf(targets, missing))
|
||||
|
||||
// The freshness watermark: only an observation taken AFTER this instant may
|
||||
// answer this run. Recorded BEFORE the refresh request so a probe that lands
|
||||
// between the two can never be missed, only double-counted as fresh — the
|
||||
// harmless direction.
|
||||
t0 := time.Now()
|
||||
e.RefreshObservatory()
|
||||
|
||||
go func() {
|
||||
defer e.groupTestRunning.Store(false)
|
||||
|
||||
@@ -303,15 +372,22 @@ func (e *Engine) TestGroups(names []string, probeURL string) (started bool) {
|
||||
e.setGroupTestResults(results)
|
||||
e.groupTestDone.Add(int64(len(missing)))
|
||||
|
||||
// The used-set and the enabled bit are snapshotted once for the whole
|
||||
// run: they only change on an apply, and a run that straddles an apply
|
||||
// is already best-effort (see the swap-safety note above).
|
||||
used := e.observatoryUsed()
|
||||
obsEnabled, _, _ := e.ObservatoryStatus()
|
||||
|
||||
// One goroutine per target: they spend their life sleeping on the board
|
||||
// poll, so there is nothing to bound — the semaphore below bounds the
|
||||
// exit-address lookups, the only real connections left in a run.
|
||||
sem := make(chan struct{}, groupTestConcurrency)
|
||||
var wg sync.WaitGroup
|
||||
for i, tgt := range targets {
|
||||
wg.Add(1)
|
||||
sem <- struct{}{}
|
||||
go func(i int, tgt groupTestTarget) {
|
||||
defer wg.Done()
|
||||
defer func() { <-sem }()
|
||||
res := e.testOneTarget(tgt, probeURL)
|
||||
res := e.testOneTarget(tgt, used, obsEnabled, t0, sem)
|
||||
e.storeGroupTestResult(i, res)
|
||||
e.groupTestDone.Add(1)
|
||||
}(i, tgt)
|
||||
@@ -393,29 +469,122 @@ func (e *Engine) groupTargets(names []string) (targets []groupTestTarget, missin
|
||||
return targets, missing
|
||||
}
|
||||
|
||||
// testOneTarget measures one target: what it currently selects, the latency
|
||||
// through it, and the public address it exits from. For a chain the dialled
|
||||
// outbound is the exit wrapper, so the delay and the exit address are end-to-end
|
||||
// properties of the whole L1..Ln path.
|
||||
func (e *Engine) testOneTarget(t groupTestTarget, probeURL string) GroupTestResult {
|
||||
// testOneTarget resolves one target WITHOUT probing it: it decides whether the
|
||||
// observatory measures this target at all, and if so waits for a fresh
|
||||
// observation and reports it. The three ways out, in order:
|
||||
//
|
||||
// 1. the used-set does not cover the target — no enabled rule routes through
|
||||
// it, so nothing measures it and nothing SHOULD: resolved immediately with
|
||||
// groupTestErrNotRouted, never dialled, empty exit address. A nil used-set
|
||||
// means "unknown" (observatory not yet configured) and is treated as
|
||||
// covered — no refusal is better than a wrong one;
|
||||
// 2. the observatory is disabled — the refresh request went nowhere, so the
|
||||
// wait below could only ever end on its deadline: resolved immediately
|
||||
// with groupTestErrProbingOff instead of stalling the panel for two
|
||||
// minutes to say the same thing;
|
||||
// 3. covered and enabled — poll the health board once a second until the
|
||||
// MEASURED TAG carries an observation newer than since, then report that
|
||||
// observation verbatim (readFreshObservation). On the deadline:
|
||||
// groupTestErrNotReached.
|
||||
//
|
||||
// The exit-address lookup runs ONLY on the alive path of (3) — a rule-routed
|
||||
// target whose path just answered a probe — and through sem, so a run never
|
||||
// opens more than groupTestConcurrency tunnelled connections at once.
|
||||
func (e *Engine) testOneTarget(t groupTestTarget, used map[string]bool, obsEnabled bool, since time.Time, sem chan struct{}) GroupTestResult {
|
||||
res := GroupTestResult{Group: t.name, TestedUnix: time.Now().Unix()}
|
||||
if t.sel != nil {
|
||||
res.Selected = t.sel()
|
||||
}
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), groupTestDelayTimeout)
|
||||
delay, err := urltest.URLTest(ctx, probeURL, t.ob)
|
||||
cancel()
|
||||
if err != nil {
|
||||
res.Error = err.Error()
|
||||
if used != nil && !used[t.ob.Tag()] {
|
||||
res.Error = groupTestErrNotRouted
|
||||
return res
|
||||
}
|
||||
if !obsEnabled {
|
||||
res.Error = groupTestErrProbingOff
|
||||
return res
|
||||
}
|
||||
res.OK = true
|
||||
res.DelayMs = int(delay)
|
||||
|
||||
// The exit address is best-effort by design: see GroupTestResult.OK.
|
||||
res.ExitIP, res.ExitCountry = e.exitAddress(t.ob)
|
||||
return res
|
||||
deadline := time.Now().Add(groupTestWaitDeadline)
|
||||
for {
|
||||
if e.readFreshObservation(&res, t, since) {
|
||||
if res.OK {
|
||||
// The exit address is best-effort by design: see GroupTestResult.
|
||||
sem <- struct{}{}
|
||||
res.ExitIP, res.ExitCountry = e.exitAddress(t.ob)
|
||||
<-sem
|
||||
}
|
||||
return res
|
||||
}
|
||||
if !time.Now().Before(deadline) {
|
||||
res.Error = groupTestErrNotReached
|
||||
return res
|
||||
}
|
||||
time.Sleep(groupTestPollEvery)
|
||||
}
|
||||
}
|
||||
|
||||
// measuredTagOf is the tag the observatory actually probes for this target's
|
||||
// dial path — the tag whose board entry answers "how is this target doing":
|
||||
//
|
||||
// - a GROUP target dials whatever it currently selects, so the group's health
|
||||
// IS its selection's health: OutboundGroup.Now(). For an egress-bound group
|
||||
// that is a per-group copy tag, which is exactly what the plan probes;
|
||||
// - a CHAIN whose exit wrapper is a PLAIN outbound is probed end-to-end under
|
||||
// that wrapper tag (probeplan.go: a chain exit is its own measurement);
|
||||
// - a CHAIN whose exit wrapper is a GROUP (the last hop is a group) has its
|
||||
// member copies probed instead of the wrapper, so the wrapper's Now() — the
|
||||
// member copy the chain currently dials through — is the measured tag. The
|
||||
// wrapper here IS the last group hop, so this is chainLastGroupHop's Now()
|
||||
// without a second lookup;
|
||||
// - fallback: the target's own tag, for a group that has not selected yet
|
||||
// (cold selector mid-swap). Its board entry is almost certainly empty, and
|
||||
// the caller then honestly reports "not reached" rather than inventing one.
|
||||
func measuredTagOf(t groupTestTarget) string {
|
||||
if g, ok := t.ob.(adapter.OutboundGroup); ok {
|
||||
if now := g.Now(); now != "" {
|
||||
return now
|
||||
}
|
||||
}
|
||||
return t.ob.Tag()
|
||||
}
|
||||
|
||||
// readFreshObservation reads the board once: if the target's measured tag holds
|
||||
// an observation taken at or after since, it is written into res (state, delay,
|
||||
// the observation's own timestamp, and the current selection so Selected and
|
||||
// the measurement describe the same pick) and true is returned. Otherwise res
|
||||
// is left for the next poll.
|
||||
//
|
||||
// Precision note: the board reports ages in whole seconds, so "at or after
|
||||
// since" is accurate to one second — an observation taken up to a second
|
||||
// BEFORE the refresh can slip through as fresh. That is the acceptable
|
||||
// direction: it is still a real measurement of the same path, at most a second
|
||||
// older than requested; the strict direction (discarding genuinely fresh
|
||||
// observations) would make every run one probe interval slower for nothing.
|
||||
func (e *Engine) readFreshObservation(res *GroupTestResult, t groupTestTarget, since time.Time) bool {
|
||||
// Re-read the selection at every poll: a forced observatory pass is exactly
|
||||
// the kind of event that makes a urltest group switch members, and the
|
||||
// measurement below is taken against the CURRENT pick.
|
||||
if t.sel != nil {
|
||||
res.Selected = t.sel()
|
||||
}
|
||||
tag := measuredTagOf(t)
|
||||
now := time.Now()
|
||||
state, delayMs, age := e.HealthView().State(tag)
|
||||
if age < 0 {
|
||||
return false // no observation at all (untested)
|
||||
}
|
||||
observedAt := now.Add(-time.Duration(age) * time.Second)
|
||||
if observedAt.Before(since) {
|
||||
return false // an old reading; the refresh has not reached this tag yet
|
||||
}
|
||||
res.OK = state == HealthAlive
|
||||
res.DelayMs = delayMs
|
||||
res.TestedUnix = observedAt.Unix()
|
||||
if !res.OK {
|
||||
res.Error = groupTestErrPathDead
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// exitProbe is one exit-address service: a URL and the parser for its body.
|
||||
|
||||
+142
-15
@@ -3,6 +3,7 @@ package engine
|
||||
import (
|
||||
"context"
|
||||
"crypto/tls"
|
||||
"fmt"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
@@ -143,10 +144,10 @@ func TestChainMemberName(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// dialableOutbound routes every dial to a fixed local address — a stand-in for a
|
||||
// chain exit whose whole path is up. urltest.URLTest dials the probe URL's host
|
||||
// through the outbound, so pointing every dial at a local HTTP server makes the
|
||||
// latency probe succeed without any network.
|
||||
// dialableOutbound routes every dial to a fixed local address — the sink the
|
||||
// exit-address lookup lands in during tests. Since the rework nothing else in
|
||||
// this file dials at all; the local server refuses to be an exit-address
|
||||
// service (wrong status, no TLS), so the lookup honestly comes back empty.
|
||||
type dialableOutbound struct {
|
||||
failingOutbound
|
||||
addr string
|
||||
@@ -158,20 +159,42 @@ func (d *dialableOutbound) DialContext(ctx context.Context, network string, _ M.
|
||||
return (&net.Dialer{}).DialContext(ctx, network, d.addr)
|
||||
}
|
||||
|
||||
// TestChainExitTestMeasuresEndToEnd is the §6-S4 acceptance path for chains: the
|
||||
// exit test dials the chain's EXIT TAG, returns a measured delay, carries the
|
||||
// chain's model name (not the wrapper tag) as the result's Group, and reports the
|
||||
// last group hop's pick as Selected. The exit address is measured through the
|
||||
// same outbound (unreachable from a test => empty, never a guess).
|
||||
func TestChainExitTestMeasuresEndToEnd(t *testing.T) {
|
||||
probe := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
// countingOutbound counts every DialContext/ListenPacket. It is the tripwire of
|
||||
// the rework: a target the observatory does not cover must NEVER be dialled,
|
||||
// and the counter is the proof.
|
||||
type countingOutbound struct {
|
||||
failingOutbound
|
||||
dials int
|
||||
}
|
||||
|
||||
func (c *countingOutbound) DialContext(ctx context.Context, network string, dst M.Socksaddr) (net.Conn, error) {
|
||||
c.dials++
|
||||
return nil, fmt.Errorf("dial refused (test)")
|
||||
}
|
||||
|
||||
func (c *countingOutbound) ListenPacket(ctx context.Context, dst M.Socksaddr) (net.PacketConn, error) {
|
||||
c.dials++
|
||||
return nil, fmt.Errorf("dial refused (test)")
|
||||
}
|
||||
|
||||
// groupTestSem is a fresh exit-lookup semaphore for direct testOneTarget calls.
|
||||
func groupTestSem() chan struct{} { return make(chan struct{}, groupTestConcurrency) }
|
||||
|
||||
// TestChainTestReportsObservatoryMeasurement is the §6-S4 acceptance path for
|
||||
// chains under the one-probe rule: the manual test does NOT dial the exit — it
|
||||
// reads the observatory's board entry for the exit tag, reports its delay and
|
||||
// ITS timestamp, carries the chain's model name as Group and the last group
|
||||
// hop's pick as Selected. The exit-address lookup still travels the exit
|
||||
// outbound (unreachable from a test => empty, never a guess).
|
||||
func TestChainTestReportsObservatoryMeasurement(t *testing.T) {
|
||||
sink := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}))
|
||||
defer probe.Close()
|
||||
defer sink.Close()
|
||||
|
||||
exit := &dialableOutbound{
|
||||
failingOutbound: failingOutbound{tag: "chain-x-h2"},
|
||||
addr: probe.Listener.Addr().String(),
|
||||
addr: sink.Listener.Addr().String(),
|
||||
deps: []string{"chain-x-h1"},
|
||||
}
|
||||
pool := []adapter.Outbound{
|
||||
@@ -187,14 +210,31 @@ func TestChainExitTestMeasuresEndToEnd(t *testing.T) {
|
||||
if len(targets) != 1 || targets[0].ob.Tag() != "chain-x-h2" {
|
||||
t.Fatalf("targets = %+v, want the chain dialled via its exit tag", targets)
|
||||
}
|
||||
if got := measuredTagOf(targets[0]); got != "chain-x-h2" {
|
||||
t.Fatalf("measuredTagOf = %q, want the plain exit wrapper itself", got)
|
||||
}
|
||||
|
||||
e := New()
|
||||
res := e.testOneTarget(targets[0], probe.URL)
|
||||
// The observatory measured the exit 10 seconds ago, AFTER the run's start
|
||||
// instant below: that observation — delay and timestamp both — is the answer.
|
||||
observedAt := time.Now().Add(-10 * time.Second)
|
||||
e.URLTestHistory().StoreURLTestHistory("chain-x-h2", &adapter.URLTestHistory{LastOK: observedAt, Delay: 77})
|
||||
since := time.Now().Add(-time.Minute)
|
||||
|
||||
res := e.testOneTarget(targets[0], map[string]bool{"chain-x-h2": true}, true, since, groupTestSem())
|
||||
if res.Group != "x" {
|
||||
t.Errorf("Group = %q, want the chain's model name x", res.Group)
|
||||
}
|
||||
if !res.OK || res.Error != "" {
|
||||
t.Fatalf("result = %+v, want a successful measurement through the exit tag (a local roundtrip may legitimately read 0ms)", res)
|
||||
t.Fatalf("result = %+v, want the board's alive observation reported", res)
|
||||
}
|
||||
if res.DelayMs != 77 {
|
||||
t.Errorf("DelayMs = %d, want the observatory's 77 — this file measures nothing itself", res.DelayMs)
|
||||
}
|
||||
// tested_unix is the OBSERVATION's instant (whole-second precision), never
|
||||
// the poll's.
|
||||
if got, want := res.TestedUnix, observedAt.Unix(); got < want-1 || got > want+1 {
|
||||
t.Errorf("TestedUnix = %d, want the observation's instant ~%d", got, want)
|
||||
}
|
||||
if res.Selected != "relay" {
|
||||
t.Errorf("Selected = %q, want the last group hop's pick relay", res.Selected)
|
||||
@@ -204,6 +244,93 @@ func TestChainExitTestMeasuresEndToEnd(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestTestOneTargetUnroutedNeverDialled is the rework's core promise: a target
|
||||
// outside the observatory's used-set is resolved immediately with the
|
||||
// documented explanation and ZERO dials — no latency probe, no exit-address
|
||||
// lookup, nothing. Dialling it would manufacture exactly the false direct
|
||||
// measurement the rework removed.
|
||||
func TestTestOneTargetUnroutedNeverDialled(t *testing.T) {
|
||||
e := New()
|
||||
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "idle"}}
|
||||
tgt := groupTestTarget{name: "idle", ob: ob}
|
||||
|
||||
res := e.testOneTarget(tgt, map[string]bool{"something-else": true}, true, time.Now(), groupTestSem())
|
||||
if ob.dials != 0 {
|
||||
t.Fatalf("an unrouted target was dialled %d time(s); it must never be", ob.dials)
|
||||
}
|
||||
if res.OK {
|
||||
t.Fatal("an unrouted target reported ok=true")
|
||||
}
|
||||
if res.Error != groupTestErrNotRouted {
|
||||
t.Fatalf("Error = %q, want the not-routed explanation", res.Error)
|
||||
}
|
||||
if res.ExitIP != "" || res.ExitCountry != "" {
|
||||
t.Fatalf("an unrouted target must carry no exit address: %+v", res)
|
||||
}
|
||||
}
|
||||
|
||||
// A disabled observatory resolves covered targets immediately too: the refresh
|
||||
// request went nowhere, so waiting out the deadline would only delay the same
|
||||
// honest answer — and still, nothing is dialled.
|
||||
func TestTestOneTargetProbingOffNeverDialled(t *testing.T) {
|
||||
e := New()
|
||||
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "auto"}}
|
||||
tgt := groupTestTarget{name: "auto", ob: ob}
|
||||
|
||||
res := e.testOneTarget(tgt, nil, false, time.Now(), groupTestSem())
|
||||
if ob.dials != 0 {
|
||||
t.Fatalf("target dialled %d time(s) with probing off; want 0", ob.dials)
|
||||
}
|
||||
if res.OK || res.Error != groupTestErrProbingOff {
|
||||
t.Fatalf("result = %+v, want ok=false with the probing-off explanation", res)
|
||||
}
|
||||
}
|
||||
|
||||
// The deadline path: a covered target whose measured tag never receives a fresh
|
||||
// observation resolves with the not-reached explanation — and, again, without a
|
||||
// single dial of its own.
|
||||
func TestTestOneTargetDeadlineWithoutObservation(t *testing.T) {
|
||||
// Shrink the wait so the test answers in milliseconds; restore afterwards.
|
||||
oldDeadline, oldPoll := groupTestWaitDeadline, groupTestPollEvery
|
||||
groupTestWaitDeadline, groupTestPollEvery = 30*time.Millisecond, 5*time.Millisecond
|
||||
defer func() { groupTestWaitDeadline, groupTestPollEvery = oldDeadline, oldPoll }()
|
||||
|
||||
e := New()
|
||||
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "auto"}}
|
||||
tgt := groupTestTarget{name: "auto", ob: ob}
|
||||
|
||||
res := e.testOneTarget(tgt, map[string]bool{"auto": true}, true, time.Now(), groupTestSem())
|
||||
if ob.dials != 0 {
|
||||
t.Fatalf("target dialled %d time(s) while waiting on the board; want 0", ob.dials)
|
||||
}
|
||||
if res.OK || res.Error != groupTestErrNotReached {
|
||||
t.Fatalf("result = %+v, want ok=false with the not-reached explanation", res)
|
||||
}
|
||||
}
|
||||
|
||||
// A fresh DEAD observation is an answer, not a timeout: ok=false with the
|
||||
// path-dead explanation, delay 0, and no exit-address connection for a path
|
||||
// that just failed its probe.
|
||||
func TestTestOneTargetReportsFreshDeath(t *testing.T) {
|
||||
e := New()
|
||||
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "auto"}}
|
||||
tgt := groupTestTarget{name: "auto", ob: ob}
|
||||
|
||||
since := time.Now().Add(-time.Minute)
|
||||
e.URLTestHistory().MarkFailed("auto")
|
||||
|
||||
res := e.testOneTarget(tgt, map[string]bool{"auto": true}, true, since, groupTestSem())
|
||||
if ob.dials != 0 {
|
||||
t.Fatalf("a dead target was dialled %d time(s); the exit lookup is for alive paths only", ob.dials)
|
||||
}
|
||||
if res.OK || res.Error != groupTestErrPathDead {
|
||||
t.Fatalf("result = %+v, want ok=false with the path-dead explanation", res)
|
||||
}
|
||||
if res.DelayMs != 0 || res.ExitIP != "" {
|
||||
t.Fatalf("a dead path must carry no delay and no exit address: %+v", res)
|
||||
}
|
||||
}
|
||||
|
||||
// TestGroupTestSingleton is the "no parallel runs" invariant: while a run holds the
|
||||
// guard, a second TestGroups must be refused rather than starting a concurrent run
|
||||
// (each of these opens real tunnelled connections).
|
||||
|
||||
@@ -100,6 +100,14 @@ type observatoryState struct {
|
||||
cycles uint64
|
||||
// busy guards against a slow tick overlapping the next one.
|
||||
busy bool
|
||||
// force is the ONE-SHOT "check now" flag raised by RefreshObservatory: while
|
||||
// it is up the freshness gate is bypassed, so a manual refresh really
|
||||
// RE-MEASURES everything on the plan instead of skipping whatever happens to
|
||||
// be fresh, and each finished batch immediately nudges the next one so the
|
||||
// pass runs back-to-back rather than on the 10s tick. It is cleared the
|
||||
// moment the forced pass schedules its last batch (cursor reaches the end of
|
||||
// the plan) — one full walk, then back to the polite steady state.
|
||||
force bool
|
||||
}
|
||||
|
||||
// ConfigureObservatory installs (or re-tunes, or stops) the observatory. It is
|
||||
@@ -120,6 +128,7 @@ func (e *Engine) ConfigureObservatory(cfg ObservatoryConfig) {
|
||||
|
||||
if !cfg.Enabled {
|
||||
e.obs.plan, e.obs.used, e.obs.cursor = nil, nil, 0
|
||||
e.obs.force = false
|
||||
if e.obs.stop != nil {
|
||||
close(e.obs.stop)
|
||||
e.obs.stop, e.obs.done, e.obs.nudge = nil, nil, nil
|
||||
@@ -163,6 +172,7 @@ func (e *Engine) StopObservatory() {
|
||||
e.obs.stop, e.obs.done, e.obs.nudge = nil, nil, nil
|
||||
e.obs.enabled = false
|
||||
e.obs.plan, e.obs.used, e.obs.cursor = nil, nil, 0
|
||||
e.obs.force = false
|
||||
if stop != nil {
|
||||
close(stop)
|
||||
}
|
||||
@@ -172,6 +182,38 @@ func (e *Engine) StopObservatory() {
|
||||
}
|
||||
}
|
||||
|
||||
// RefreshObservatory asks for an OUT-OF-TURN full pass of the probe plan: the
|
||||
// cursor is rewound to the top, the one-shot force flag is raised so the pass
|
||||
// bypasses the freshness gate (a "check now" that skipped everything fresh
|
||||
// would measure nothing and answer with yesterday's numbers), and the loop is
|
||||
// nudged so the first batch starts immediately instead of on the next 10s tick.
|
||||
//
|
||||
// This is the ONLY way the panel's "Test" button touches the network now: the
|
||||
// manual group test (grouptest.go) no longer dials anything itself — it calls
|
||||
// this, then watches the health board for observations newer than its start
|
||||
// instant. One prober, one set of dial paths, one truth.
|
||||
//
|
||||
// No-op when the observatory is disabled or has no plan: there is nothing to
|
||||
// walk, and raising force with no loop to consume it would only arm a stale
|
||||
// flag for a future enable. The caller (TestGroups) reports that situation
|
||||
// through its own results; this method stays silent about it on purpose —
|
||||
// it is a request, not a query.
|
||||
func (e *Engine) RefreshObservatory() {
|
||||
e.obsMu.Lock()
|
||||
defer e.obsMu.Unlock()
|
||||
if !e.obs.enabled || len(e.obs.plan) == 0 {
|
||||
return
|
||||
}
|
||||
e.obs.cursor = 0
|
||||
e.obs.force = true
|
||||
if e.obs.nudge != nil {
|
||||
select {
|
||||
case e.obs.nudge <- struct{}{}:
|
||||
default:
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ObservatoryStatus reports whether the observatory is on, how many measurements
|
||||
// the current plan holds, and how many full passes have completed. Cheap; used by
|
||||
// tests and available for diagnostics.
|
||||
@@ -240,17 +282,18 @@ func (e *Engine) observatoryLoop(stop <-chan struct{}, done chan<- struct{}, nud
|
||||
// observatoryTickOnce runs one batch. It is deliberately conservative about when
|
||||
// it does nothing:
|
||||
//
|
||||
// - a manual exit test (grouptest.go) is in flight => SKIP. That run is the
|
||||
// human's explicit request; the observatory defers rather than competing for
|
||||
// the uplink. It checks the guard but never CLAIMS it, so a pressed Test
|
||||
// button can never be answered "already running" because of background work;
|
||||
// - the previous tick has not finished => SKIP, so slow probes never stack;
|
||||
// - the engine is stopped => jobs no longer resolve and are skipped (probeJob).
|
||||
//
|
||||
// There is deliberately NO deferral to a manual group test any more. The guard
|
||||
// that used to sit here ("skip the tick while e.groupTestRunning is up") made
|
||||
// sense when the manual run dialled the uplink itself and the two would have
|
||||
// competed for it; that run no longer probes ANYTHING — it asks for a forced
|
||||
// pass via RefreshObservatory and then reads the board. Keeping the guard would
|
||||
// therefore be worse than pointless: a manual run now DEPENDS on the
|
||||
// observatory ticking, and skipping ticks for its whole duration would deadlock
|
||||
// the very refresh it is waiting on (up to its full 120s deadline).
|
||||
func (e *Engine) observatoryTickOnce() {
|
||||
if e.groupTestRunning.Load() {
|
||||
return
|
||||
}
|
||||
|
||||
e.obsMu.Lock()
|
||||
if e.obs.busy || !e.obs.enabled || len(e.obs.plan) == 0 {
|
||||
e.obsMu.Unlock()
|
||||
@@ -258,7 +301,8 @@ func (e *Engine) observatoryTickOnce() {
|
||||
}
|
||||
if e.obs.cursor >= len(e.obs.plan) {
|
||||
// Cycle complete: count it and start the next pass from the top. This is
|
||||
// the ONE place the cursor is deliberately rewound.
|
||||
// the ONE place the cursor is rewound on the observatory's own schedule
|
||||
// (RefreshObservatory rewinds it too, but that is the caller's request).
|
||||
e.obs.cycles++
|
||||
e.obs.cursor = 0
|
||||
}
|
||||
@@ -270,12 +314,32 @@ func (e *Engine) observatoryTickOnce() {
|
||||
batch := append([]ProbeJob(nil), e.obs.plan[start:end]...)
|
||||
e.obs.cursor = end
|
||||
interval := e.obs.interval
|
||||
// A forced pass (RefreshObservatory) bypasses the freshness gate for its one
|
||||
// walk of the plan. The flag is captured for THIS batch and cleared the
|
||||
// moment the walk's last batch is scheduled: the remaining jobs of this very
|
||||
// batch still run unfiltered off the captured copy, and the next pass is an
|
||||
// ordinary polite one again.
|
||||
force := e.obs.force
|
||||
if force && end >= len(e.obs.plan) {
|
||||
e.obs.force = false
|
||||
}
|
||||
e.obs.busy = true
|
||||
e.obsMu.Unlock()
|
||||
|
||||
defer func() {
|
||||
e.obsMu.Lock()
|
||||
e.obs.busy = false
|
||||
// Mid-forced-pass: chain straight into the next batch instead of waiting
|
||||
// out the 10s tick. A human pressed "check now"; walking a subscription-
|
||||
// sized plan at one batch per tick would stretch the answer past the
|
||||
// manual run's deadline for no reason — the batch size still bounds how
|
||||
// much is in flight at once, which is the limit that actually matters.
|
||||
if e.obs.force && e.obs.nudge != nil {
|
||||
select {
|
||||
case e.obs.nudge <- struct{}{}:
|
||||
default:
|
||||
}
|
||||
}
|
||||
e.obsMu.Unlock()
|
||||
}()
|
||||
|
||||
@@ -287,7 +351,7 @@ func (e *Engine) observatoryTickOnce() {
|
||||
sem := make(chan struct{}, observatoryConcurrency)
|
||||
var wg sync.WaitGroup
|
||||
for _, j := range batch {
|
||||
if !observatoryShouldProbe(hist, j, interval, now) {
|
||||
if !force && !observatoryShouldProbe(hist, j, interval, now) {
|
||||
continue
|
||||
}
|
||||
wg.Add(1)
|
||||
|
||||
@@ -100,28 +100,89 @@ func TestObservatoryTickStoppedEngine(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The exit-test gate: while a manual group/chain test is in flight the tick is a
|
||||
// no-op — the observatory checks the guard but never claims it, so a pressed
|
||||
// Test button is never refused because of background work.
|
||||
func TestObservatoryDefersToExitTest(t *testing.T) {
|
||||
// The INVERSE of the old exit-test gate, pinned so it cannot come back: the
|
||||
// tick must RUN while a manual group test is in flight. The manual run no
|
||||
// longer probes anything — it asks for a forced pass and then waits on the
|
||||
// board — so a tick that deferred to it would deadlock the very refresh the
|
||||
// run is polling for, until its 120s deadline reported "not reached" about
|
||||
// paths nobody attempted.
|
||||
func TestObservatoryTicksDuringManualRun(t *testing.T) {
|
||||
e := New()
|
||||
defer e.StopObservatory()
|
||||
e.ConfigureObservatory(ObservatoryConfig{Enabled: true, Options: obsFixture(), ProbeInterval: time.Minute})
|
||||
|
||||
e.groupTestRunning.Store(true)
|
||||
e.observatoryTickOnce()
|
||||
if _, cursor, _ := obsState(e); cursor != 0 {
|
||||
t.Fatal("tick ran while an exit test was in flight")
|
||||
if _, cursor, _ := obsState(e); cursor == 0 {
|
||||
t.Fatal("tick deferred to an in-flight manual run; the run depends on the tick now")
|
||||
}
|
||||
e.groupTestRunning.Store(false)
|
||||
|
||||
// And the guard is free: a manual run can start immediately.
|
||||
// The observatory never claims the manual-run guard either way round: a
|
||||
// manual run can still start immediately.
|
||||
if !e.TestGroups(nil, "") {
|
||||
t.Fatal("a manual exit test was refused; the observatory must never hold groupTestRunning")
|
||||
t.Fatal("a manual test was refused; the observatory must never hold groupTestRunning")
|
||||
}
|
||||
waitGroupTestIdle(t, e)
|
||||
}
|
||||
|
||||
// RefreshObservatory is the manual run's whole probing story: it rewinds the
|
||||
// cursor, raises the one-shot force flag (bypassing the freshness gate for one
|
||||
// full walk), and the flag clears itself when the forced pass schedules its
|
||||
// last batch. Disabled or plan-less observatories ignore it entirely.
|
||||
func TestRefreshObservatoryForcesOnePass(t *testing.T) {
|
||||
e := New()
|
||||
defer e.StopObservatory()
|
||||
|
||||
// Disabled: a refresh request is a documented no-op.
|
||||
e.RefreshObservatory()
|
||||
e.obsMu.Lock()
|
||||
force := e.obs.force
|
||||
e.obsMu.Unlock()
|
||||
if force {
|
||||
t.Fatal("RefreshObservatory raised force on a disabled observatory")
|
||||
}
|
||||
|
||||
e.ConfigureObservatory(ObservatoryConfig{Enabled: true, Options: obsFixture(), ProbeInterval: time.Minute})
|
||||
|
||||
// Fresh observations on every planned tag: the polite gate would skip them
|
||||
// all, which is exactly what a forced pass must NOT do.
|
||||
hist := e.URLTestHistory()
|
||||
now := time.Now()
|
||||
hist.StoreURLTestHistory("n1", &adapter.URLTestHistory{LastOK: now, Delay: 5})
|
||||
hist.StoreURLTestHistory("n2", &adapter.URLTestHistory{LastOK: now, Delay: 5})
|
||||
|
||||
// Pretend a walk was mid-plan; the refresh must rewind it.
|
||||
e.obsMu.Lock()
|
||||
e.obs.cursor = 1
|
||||
e.obsMu.Unlock()
|
||||
e.RefreshObservatory()
|
||||
e.obsMu.Lock()
|
||||
cursor, force := e.obs.cursor, e.obs.force
|
||||
e.obsMu.Unlock()
|
||||
if cursor != 0 || !force {
|
||||
t.Fatalf("after refresh: cursor=%d force=%v, want 0/true", cursor, force)
|
||||
}
|
||||
|
||||
// One tick covers the whole 2-job plan (batch is 24), so the forced pass
|
||||
// completes and the flag clears itself — one full walk, not a permanent mode.
|
||||
e.observatoryTickOnce()
|
||||
e.obsMu.Lock()
|
||||
cursor, force = e.obs.cursor, e.obs.force
|
||||
e.obsMu.Unlock()
|
||||
if force {
|
||||
t.Fatal("force flag survived the forced pass; it must be one-shot")
|
||||
}
|
||||
if cursor != 2 {
|
||||
t.Fatalf("cursor after forced tick = %d, want 2 (the pass really walked the plan)", cursor)
|
||||
}
|
||||
// The stopped engine resolves no outbounds, so the probes were skipped —
|
||||
// but the fresh history proves the GATE was bypassed only if the jobs were
|
||||
// attempted; attempt-tracking lives in probeJob, which needs a box. What is
|
||||
// pinned here is the flag lifecycle and the rewind, the two halves
|
||||
// TestGroups depends on.
|
||||
}
|
||||
|
||||
// The freshness gate: a job whose every covered tag has an observation younger
|
||||
// than the global probe interval is skipped — that set is exactly what an ACTIVE
|
||||
// group is measuring itself, and the observatory must not duplicate or suppress
|
||||
|
||||
+39
-20
@@ -27,8 +27,12 @@ import (
|
||||
// - a CHAIN entry tag "chain-<n>-hN" (the exit wrapper, generate/chain.go) is
|
||||
// probed ITSELF: dialling the copy pulls the whole L1→…→Ln path through its
|
||||
// Detour links, so one probe is the chain's end-to-end health. Intermediate
|
||||
// NODE hops are not probed separately — their death is visible in the exit
|
||||
// probe and no selection depends on them;
|
||||
// NODE hop wrappers ("chain-<n>-h<i>" that are plain outbounds) are probed
|
||||
// TOO — not because selection needs them (it does not), but because an
|
||||
// operator does: dialling "chain-<n>-h1" measures exactly the prefix of the
|
||||
// path up to and including hop 1, so when the exit probe dies the per-hop
|
||||
// probes say WHICH hop died instead of only that the chain did (the
|
||||
// ChainHealth hops surface, grouphealth.go);
|
||||
// - a GROUP hop inside a chain ("chain-<n>-h<i>", a wrapper selector reached by
|
||||
// walking the exit's Detour links) additionally has its member copies
|
||||
// "chain-<n>-h<i>-<member>" probed: they are the wrapper's selection
|
||||
@@ -291,8 +295,20 @@ func (w *planWalk) visitMember(group, member string) {
|
||||
// walkDetour follows a probed tag's Detour links toward the router. A group met
|
||||
// on the way is a chain's GROUP HOP wrapper: its member copies are selection
|
||||
// candidates and are probed, each through its own prefix of the path. A plain
|
||||
// outbound met on the way is an intermediate node hop — marked used, never probed
|
||||
// on its own — and the walk continues through it.
|
||||
// outbound met on the way falls in two halves:
|
||||
//
|
||||
// - a chain NODE HOP wrapper ("chain-<n>-h<i>", parseChainExitTag) is probed
|
||||
// as a measurement of its own. Dialling it pulls exactly the prefix of the
|
||||
// chain up to and including hop i through the Detour links, so its verdict
|
||||
// LOCALISES a failure: the exit probe says the chain died, the hop probes
|
||||
// say where. It used to be skipped here on the argument that its death
|
||||
// shows up in the exit probe anyway — true for end-to-end health, useless
|
||||
// for an operator staring at a dead 4-hop chain. The measurement is stored
|
||||
// under the wrapper tag alone: a chain prefix is nobody else's dial path,
|
||||
// so no alias may borrow its verdict;
|
||||
// - anything else on the detour path — an "egress-…" interface outbound, an
|
||||
// ordinary node's egress detour — keeps today's behaviour: marked used,
|
||||
// never probed on its own, and the walk continues through it.
|
||||
func (w *planWalk) walkDetour(tag string) {
|
||||
if tag == "" || w.walked[tag] {
|
||||
return
|
||||
@@ -303,25 +319,28 @@ func (w *planWalk) walkDetour(tag string) {
|
||||
return
|
||||
}
|
||||
w.used[tag] = true
|
||||
if ent.group {
|
||||
next := ""
|
||||
for _, m := range ent.members {
|
||||
ment, ok := w.byTag[m]
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
w.used[m] = true
|
||||
if !ment.group && probeablePlanTag(ment.typ, m) {
|
||||
w.addTarget(dialKeyTag(m), m)
|
||||
}
|
||||
if next == "" {
|
||||
next = ment.detour // every member detours into the same previous hop
|
||||
}
|
||||
if !ent.group {
|
||||
if _, isHopWrapper := parseChainExitTag(tag); isHopWrapper && probeablePlanTag(ent.typ, tag) {
|
||||
w.addTarget(dialKeyTag(tag), tag)
|
||||
}
|
||||
w.walkDetour(next)
|
||||
w.walkDetour(ent.detour)
|
||||
return
|
||||
}
|
||||
w.walkDetour(ent.detour)
|
||||
next := ""
|
||||
for _, m := range ent.members {
|
||||
ment, ok := w.byTag[m]
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
w.used[m] = true
|
||||
if !ment.group && probeablePlanTag(ment.typ, m) {
|
||||
w.addTarget(dialKeyTag(m), m)
|
||||
}
|
||||
if next == "" {
|
||||
next = ment.detour // every member detours into the same previous hop
|
||||
}
|
||||
}
|
||||
w.walkDetour(next)
|
||||
}
|
||||
|
||||
// dialKeyTag / dialKeyCopy build the dedup key for one measurement target.
|
||||
|
||||
@@ -218,6 +218,64 @@ func TestObservatoryPlanChain(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// A chain with a plain NODE hop wrapper on the way to the exit: the wrapper IS
|
||||
// a measurement of its own now (walkDetour) — dialling "chain-<c>-h1" measures
|
||||
// exactly the path prefix up to hop 1, which is what localises a dead hop. And
|
||||
// the other half of the honesty contract: a base node that appears ONLY as a
|
||||
// chain group-hop member is never measured under its BASE tag — the member
|
||||
// copy's observation stays on the copy, because the copy's dial path (through
|
||||
// the chain prefix) is not the base outbound's dial path, and no job may write
|
||||
// a direct-looking verdict onto a tag it did not dial.
|
||||
func TestObservatoryPlanChainHopWrappersProbed(t *testing.T) {
|
||||
opts := option.Options{
|
||||
Outbounds: []option.Outbound{
|
||||
// L1: a plain node hop wrapper.
|
||||
fixNode("chain-c2-h1", ""),
|
||||
// L2: a group hop over one member copy of base node N.
|
||||
fixNode("chain-c2-h2-N", "chain-c2-h1"),
|
||||
fixGroup(C.TypeSelector, "chain-c2-h2", "chain-c2-h2-N"),
|
||||
// L3: the exit, a plain node copy.
|
||||
fixNode("chain-c2-h3", "chain-c2-h2"),
|
||||
// The base node behind the L2 member copy: emitted, referenced by
|
||||
// NOTHING but the copy's name.
|
||||
fixNode("N", ""),
|
||||
},
|
||||
Route: fixRoute("", "chain-c2-h3"),
|
||||
}
|
||||
byDial, used := planOf(t, opts)
|
||||
|
||||
// The node hop wrapper is a job of its own, stored only under itself.
|
||||
j := assertPlanHas(t, byDial, "chain-c2-h1")
|
||||
if len(j.Store) != 1 || j.Store[0] != "chain-c2-h1" {
|
||||
t.Errorf("node hop wrapper store = %v, want only itself (a chain prefix is nobody else's path)", j.Store)
|
||||
}
|
||||
// The exit and the group-hop member copy are jobs, as before.
|
||||
assertPlanHas(t, byDial, "chain-c2-h3")
|
||||
assertPlanHas(t, byDial, "chain-c2-h2-N")
|
||||
|
||||
// NO job may record anything under the base tag N: not as a dial, not as a
|
||||
// store alias. A direct measurement can never land on it from this config.
|
||||
if _, ok := byDial["N"]; ok {
|
||||
t.Error("plan dials the base node N, which no rule reaches")
|
||||
}
|
||||
for dial, job := range byDial {
|
||||
for _, tag := range job.Store {
|
||||
if tag == "N" {
|
||||
t.Errorf("job %q stores under the base tag N; a chain member copy must never alias its base", dial)
|
||||
}
|
||||
}
|
||||
}
|
||||
if used["N"] {
|
||||
t.Error("used-set claims the base node N; only the chain copies are on the path")
|
||||
}
|
||||
// The wrappers themselves are used (they are the path).
|
||||
for _, u := range []string{"chain-c2-h1", "chain-c2-h2", "chain-c2-h3", "chain-c2-h2-N"} {
|
||||
if !used[u] {
|
||||
t.Errorf("used-set is missing %q", u)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// DNS server detours and route.Final are roots too — a group referenced only as a
|
||||
// resolver's detour is still a used, probed path.
|
||||
func TestObservatoryPlanDNSDetourRoot(t *testing.T) {
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
package engine
|
||||
|
||||
import (
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
)
|
||||
|
||||
// Standing down the self-check of unused urltest groups (health board §5.C).
|
||||
//
|
||||
// # Why the engine does this, and not the generator
|
||||
//
|
||||
// A urltest group probes its own members: a warm-up sweep at PostStart and a
|
||||
// ticker for as long as traffic touches it (protocol/group/urltest.go). For a
|
||||
// group the routing rules actually use, that probing travels the same path the
|
||||
// traffic does and is welcome. For a group NO rule reaches, it dials the
|
||||
// members' base outbounds directly from the router — a path nothing uses — and
|
||||
// stores the results under the base tags, where every health consumer reads
|
||||
// them as "the node's health". A node that only works behind a tunnel then
|
||||
// reads dead on the shared board because an idle group measured it over the
|
||||
// blocked direct WAN. The observatory's plan (probeplan.go) is the single
|
||||
// statement of what gets probed and along which paths; this file makes the
|
||||
// config that reaches box.New SAY so, by flipping SelfCheck off on every
|
||||
// urltest group outside the plan's used-set.
|
||||
//
|
||||
// The generator cannot do this: whether a group is used is a property of the
|
||||
// EMITTED rules and DNS detours as a whole, and BuildObservatoryPlan is the one
|
||||
// place that reachability is computed. Recomputing it in the generator would be
|
||||
// a second copy of the same walk, and second copies drift.
|
||||
|
||||
// standDownUnusedSelfCheck flips SelfCheck to false on every urltest group in
|
||||
// opts that the observatory's used-set does not cover, and returns how many it
|
||||
// stood down (for the apply log and the tests). Selector groups are left alone
|
||||
// — they have no self-check to stand down — and a group any rule, route.Final
|
||||
// or DNS detour reaches keeps its default (nil == on).
|
||||
//
|
||||
// It MUTATES the option structs the caller handed in: opts.Outbounds carries
|
||||
// the generator's option values as pointers behind an `any` (the generator
|
||||
// always emits *option.URLTestOutboundOptions), so writing through them changes
|
||||
// the caller's config. That is deliberate, not an accident to guard against:
|
||||
// the plan is the single source of what gets probed, and the config that
|
||||
// reaches box.New must already say so — a copy-on-write here would build a box
|
||||
// whose groups probe paths the hash and the plan claim nobody probes. It runs
|
||||
// BEFORE hashOptions in applyLocked for the same reason: the hash must describe
|
||||
// what is really built, so a rule change that flips a group used<->unused is a
|
||||
// real config change and triggers a swap.
|
||||
//
|
||||
// A urltest outbound whose Options is not the expected pointer type (a value,
|
||||
// or something foreign) is skipped rather than guessed at: it did not come from
|
||||
// our generator, and silently rebuilding somebody else's option struct is worse
|
||||
// than leaving one group's self-check up.
|
||||
func standDownUnusedSelfCheck(opts option.Options) int {
|
||||
_, used := BuildObservatoryPlan(opts, "")
|
||||
stood := 0
|
||||
for i := range opts.Outbounds {
|
||||
ob := &opts.Outbounds[i]
|
||||
if ob.Type != C.TypeURLTest || used[ob.Tag] {
|
||||
continue
|
||||
}
|
||||
utOpts, ok := ob.Options.(*option.URLTestOutboundOptions)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
// One fresh pointer per group, so no two option structs alias a shared
|
||||
// bool that a later caller could flip for both at once.
|
||||
off := false
|
||||
utOpts.SelfCheck = &off
|
||||
stood++
|
||||
}
|
||||
return stood
|
||||
}
|
||||
@@ -0,0 +1,204 @@
|
||||
package engine
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
)
|
||||
|
||||
// standDownUnusedSelfCheck is the enforcement half of the one-probe rule: a
|
||||
// urltest group no rule reaches gets SelfCheck=false written into its option
|
||||
// struct (in place — the struct the box will be built from), a rule-reachable
|
||||
// one keeps its nil default, and selectors are never touched at all (they have
|
||||
// no self-check to stand down). The count it returns is what applyLocked logs.
|
||||
func TestStandDownUnusedSelfCheck(t *testing.T) {
|
||||
opts := option.Options{
|
||||
Outbounds: []option.Outbound{
|
||||
fixNode("n1", ""),
|
||||
fixNode("idle1", ""),
|
||||
fixNode("pin1", ""),
|
||||
// Reached by a rule: must keep probing itself.
|
||||
fixGroup(C.TypeURLTest, "auto", "n1"),
|
||||
// Reached by nothing: its self-check dials a path nobody uses and
|
||||
// must be stood down.
|
||||
fixGroup(C.TypeURLTest, "idle", "idle1"),
|
||||
// A selector reached by nothing: left alone — no self-check exists.
|
||||
fixGroup(C.TypeSelector, "pins", "pin1"),
|
||||
},
|
||||
Route: fixRoute("auto"),
|
||||
}
|
||||
|
||||
stood := standDownUnusedSelfCheck(opts)
|
||||
if stood != 1 {
|
||||
t.Fatalf("stood down %d group(s), want exactly 1 (idle)", stood)
|
||||
}
|
||||
|
||||
find := func(tag string) option.Outbound {
|
||||
t.Helper()
|
||||
for _, ob := range opts.Outbounds {
|
||||
if ob.Tag == tag {
|
||||
return ob
|
||||
}
|
||||
}
|
||||
t.Fatalf("outbound %q missing from opts", tag)
|
||||
return option.Outbound{}
|
||||
}
|
||||
|
||||
// The unused urltest group: SelfCheck is now an explicit false in the very
|
||||
// struct the caller handed in (the mutation is the point — the config that
|
||||
// reaches box.New must say what the plan says).
|
||||
idle := find("idle").Options.(*option.URLTestOutboundOptions)
|
||||
if idle.SelfCheck == nil || *idle.SelfCheck {
|
||||
t.Fatalf("idle group SelfCheck = %v, want an explicit false", idle.SelfCheck)
|
||||
}
|
||||
|
||||
// The used urltest group keeps the nil default (self-check on): its probing
|
||||
// travels the path traffic actually takes and is welcome.
|
||||
auto := find("auto").Options.(*option.URLTestOutboundOptions)
|
||||
if auto.SelfCheck != nil {
|
||||
t.Fatalf("used group SelfCheck = %v, want nil (untouched default)", *auto.SelfCheck)
|
||||
}
|
||||
|
||||
// The selector's option struct has no SelfCheck field at all; what is
|
||||
// pinned here is that stand-down neither counted it nor mangled its type.
|
||||
if _, ok := find("pins").Options.(*option.SelectorOutboundOptions); !ok {
|
||||
t.Fatal("selector options were rebuilt; stand-down must leave selectors alone")
|
||||
}
|
||||
|
||||
// Idempotence: a second pass finds the same one group (already-false is
|
||||
// still counted — the function reports plan membership, not novelty) and
|
||||
// changes nothing further. This is what keeps the apply hash stable across
|
||||
// the every-minute reconcile.
|
||||
if again := standDownUnusedSelfCheck(opts); again != 1 {
|
||||
t.Fatalf("second pass stood down %d, want 1 (deterministic on identical opts)", again)
|
||||
}
|
||||
if idle.SelfCheck == nil || *idle.SelfCheck {
|
||||
t.Fatal("second pass flipped the idle group's SelfCheck back")
|
||||
}
|
||||
}
|
||||
|
||||
// TestStandDownNeverTouchesChainHopWrappers pins the property the production
|
||||
// config lives or dies by, DIRECTLY rather than transitively through the plan
|
||||
// tests: a CHAIN HOP WRAPPER group ("chain-<name>-h<i>", type urltest) must
|
||||
// NEVER be stood down.
|
||||
//
|
||||
// The failure mode this guards against is subtle and silent. The owner's live
|
||||
// rule routes through egress:ewan -> node:awgout -> group:sub0 -> group:sub1
|
||||
// -> group:sub2 — a chain whose group hops are rebuilt as urltest wrappers
|
||||
// over per-chain member copies. Those wrappers are the ONLY probing that
|
||||
// travels the real path: their own schedule is the on-path measurement, and
|
||||
// the selection inside the chain (which member each hop dials) depends on the
|
||||
// verdicts it writes. Reachability of a wrapper is established by walkDetour
|
||||
// (probeplan.go) walking the exit's Detour links — a DIFFERENT code path from
|
||||
// the root/member expansion the other tests exercise. If a future change to
|
||||
// that walk dropped the wrappers from the used-set, standDownUnusedSelfCheck
|
||||
// would obediently flip their SelfCheck off, the chain would stop measuring
|
||||
// itself, and nothing else would ever measure it: we would have replaced the
|
||||
// false reading this rework removed with NO reading at all. Both existing
|
||||
// tests could stay green while that happened — the plan test asserts the
|
||||
// used-set, not the stand-down; the stand-down test above never mentions
|
||||
// chains. This one closes that gap by asserting the stand-down's OUTPUT on
|
||||
// the production topology.
|
||||
//
|
||||
// The fixture mirrors what the generator actually emits for that config, not
|
||||
// a toy: an exit wrapper chain-p-h4 (urltest over member copies) whose copies
|
||||
// detour into chain-p-h3 (urltest), whose copies detour into chain-p-h2
|
||||
// (urltest), whose copies detour into chain-p-h1 (the plain AmneziaWG hop
|
||||
// copy), which detours into egress-ewan — the detour lives on the member
|
||||
// COPIES, not on the wrappers, exactly as generate/chain.go builds it. The
|
||||
// base groups sub0/sub1/sub2 are ALSO emitted, over the base node tags,
|
||||
// reached by nothing — which is precisely what buildGroups does today and
|
||||
// precisely the false direct prober the stand-down exists to silence.
|
||||
func TestStandDownNeverTouchesChainHopWrappers(t *testing.T) {
|
||||
opts := option.Options{
|
||||
Outbounds: []option.Outbound{
|
||||
// The egress the whole chain leaves through.
|
||||
fixUtility(C.TypeDirect, "egress-ewan"),
|
||||
// L1: the plain AWG hop copy (in production an endpoint; a plain
|
||||
// outbound here exercises the same walk — walkDetour treats both as
|
||||
// non-group entries).
|
||||
fixNode("chain-p-h1", "egress-ewan"),
|
||||
// L2..L4: urltest group hop wrappers over per-chain member copies;
|
||||
// each COPY detours into the previous hop.
|
||||
fixNode("chain-p-h2-a", "chain-p-h1"),
|
||||
fixNode("chain-p-h2-b", "chain-p-h1"),
|
||||
fixGroup(C.TypeURLTest, "chain-p-h2", "chain-p-h2-a", "chain-p-h2-b"),
|
||||
fixNode("chain-p-h3-a", "chain-p-h2"),
|
||||
fixNode("chain-p-h3-b", "chain-p-h2"),
|
||||
fixGroup(C.TypeURLTest, "chain-p-h3", "chain-p-h3-a", "chain-p-h3-b"),
|
||||
fixNode("chain-p-h4-a", "chain-p-h3"),
|
||||
fixNode("chain-p-h4-b", "chain-p-h3"),
|
||||
fixGroup(C.TypeURLTest, "chain-p-h4", "chain-p-h4-a", "chain-p-h4-b"),
|
||||
// The base nodes and the base groups over them: emitted alongside
|
||||
// the chain, reached by no rule — the generator keeps emitting them
|
||||
// today, and they are exactly the direct-WAN probers that poisoned
|
||||
// the board.
|
||||
fixNode("a", ""),
|
||||
fixNode("b", ""),
|
||||
fixGroup(C.TypeURLTest, "sub0", "a", "b"),
|
||||
fixGroup(C.TypeURLTest, "sub1", "a", "b"),
|
||||
fixGroup(C.TypeURLTest, "sub2", "a", "b"),
|
||||
},
|
||||
// One rule, targeting the chain's exit wrapper — the production shape.
|
||||
Route: fixRoute("", "chain-p-h4"),
|
||||
}
|
||||
|
||||
stood := standDownUnusedSelfCheck(opts)
|
||||
if stood != 3 {
|
||||
t.Fatalf("stood down %d group(s), want exactly the 3 base groups sub0/sub1/sub2", stood)
|
||||
}
|
||||
|
||||
utOptions := func(tag string) *option.URLTestOutboundOptions {
|
||||
t.Helper()
|
||||
for _, ob := range opts.Outbounds {
|
||||
if ob.Tag == tag {
|
||||
o, ok := ob.Options.(*option.URLTestOutboundOptions)
|
||||
if !ok {
|
||||
t.Fatalf("outbound %q is not a urltest group", tag)
|
||||
}
|
||||
return o
|
||||
}
|
||||
}
|
||||
t.Fatalf("outbound %q missing from opts", tag)
|
||||
return nil
|
||||
}
|
||||
|
||||
// Every hop wrapper keeps its self-check: nil, the untouched default. An
|
||||
// explicit false on ANY of these is the chain going blind.
|
||||
for _, wrapper := range []string{"chain-p-h2", "chain-p-h3", "chain-p-h4"} {
|
||||
if sc := utOptions(wrapper).SelfCheck; sc != nil {
|
||||
t.Errorf("hop wrapper %q SelfCheck = %v, want nil — its own schedule IS the on-path measurement", wrapper, *sc)
|
||||
}
|
||||
}
|
||||
// Every base group is stood down: nothing routes through them, and their
|
||||
// probing would dial the base nodes over the direct WAN.
|
||||
for _, base := range []string{"sub0", "sub1", "sub2"} {
|
||||
if sc := utOptions(base).SelfCheck; sc == nil || *sc {
|
||||
t.Errorf("base group %q SelfCheck = %v, want an explicit false", base, sc)
|
||||
}
|
||||
}
|
||||
|
||||
// And the reason the stand-down of sub0/sub1/sub2 is SAFE in this config:
|
||||
// with the base groups silenced, the base node tags a/b have no measurement
|
||||
// path at all — no job dials them and no job stores under them (the chain's
|
||||
// member copies record only under themselves). Their board entries simply
|
||||
// stop being written, honestly untested, instead of carrying a false
|
||||
// direct-WAN verdict.
|
||||
jobs, used := BuildObservatoryPlan(opts, "")
|
||||
for _, baseTag := range []string{"a", "b"} {
|
||||
if used[baseTag] {
|
||||
t.Errorf("used-set claims base node %q; only the chain copies are on the path", baseTag)
|
||||
}
|
||||
for _, j := range jobs {
|
||||
if j.Dial == baseTag {
|
||||
t.Errorf("plan dials base node %q, which no rule reaches", baseTag)
|
||||
}
|
||||
for _, s := range j.Store {
|
||||
if s == baseTag {
|
||||
t.Errorf("job %q stores under base node %q; a chain copy must never alias its base", j.Dial, baseTag)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,344 @@
|
||||
package engine
|
||||
|
||||
// Bounded, deterministic teardown of a SUPERSEDED sing-box instance.
|
||||
//
|
||||
// # The defect this file exists for
|
||||
//
|
||||
// sing-box's Box has no live reload, so every applied config change builds a
|
||||
// fresh Box and retires the old one (see engine.go). Retiring it used to be one
|
||||
// line — `_ = old.Close()` — and that line had three separate problems, all of
|
||||
// which had to be true at once for the observed failure:
|
||||
//
|
||||
// 1. NO CANCELLATION. Upstream's own runner gives every Box its OWN cancellable
|
||||
// context and calls cancel() BEFORE Close (cmd/sing-box/cmd_run.go:136 and
|
||||
// :188-190). This engine handed EVERY box.New the one shared, never-cancelled
|
||||
// context built in New (engine.go), and box.New does not derive a cancellable
|
||||
// child of what it is given (box.go: `ctx := options.Context` and nothing
|
||||
// else). So every goroutine inside a retired box that would have stopped on
|
||||
// context cancellation simply never stopped, and only what each adapter's
|
||||
// Close() explicitly tears down actually went away.
|
||||
//
|
||||
// 2. NO TIME BUDGET. Box.Close walks its subsystems SEQUENTIALLY and waits for
|
||||
// each one forever; the only thing watching the clock is a taskmonitor that
|
||||
// prints "close endpoint/wireguard[...] take too much time to finish!" after
|
||||
// C.StopTimeout and then keeps waiting anyway (common/taskmonitor/monitor.go).
|
||||
// A wireguard endpoint that will not come down therefore stalls the rest of
|
||||
// the close list, and every subsystem AFTER it in the walk is never reached.
|
||||
//
|
||||
// 3. NO EVIDENCE. The error was discarded (`_ =`) and the pointer to the old box
|
||||
// was overwritten in the same breath, so after the swap nothing in the process
|
||||
// could tell — or even ask — whether the previous generation had actually
|
||||
// died. On the router this accumulated: up to four generations logging side by
|
||||
// side inside one shaterd process, each with its own outbound connections and,
|
||||
// worse, its own WireGuard devices. Two devices sharing a private key evict
|
||||
// each other at the peer, which is exactly the fault generate/wgdedup.go
|
||||
// removes WITHIN a config — reintroduced here BETWEEN generations.
|
||||
//
|
||||
// # The contract now
|
||||
//
|
||||
// - Every box gets its own cancellable context, cancelled before Close.
|
||||
// - Close runs against a HARD budget (closeBudget). When it expires the apply
|
||||
// moves on: the new box is already started and serving, and blocking the
|
||||
// control plane on a shutdown that is not going to finish would only add an
|
||||
// unresponsive panel to the problem.
|
||||
// - An overrun is LOUD, not swallowed: an ERROR line naming the generation, and
|
||||
// a critical entry in the operator-facing warning set for as long as the
|
||||
// abandoned generation is still running (see PendingCloses).
|
||||
// - Generations do not stack: before an apply builds another instance it gives
|
||||
// any abandoned teardown one more budget to finish, and if one is still alive
|
||||
// the swap is forced onto the close-old-then-start-new path so the process
|
||||
// never holds two LIVE boxes on top of an abandoned one.
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"io"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
box "github.com/sagernet/sing-box"
|
||||
)
|
||||
|
||||
// ErrCloseTimeout reports that a superseded instance did not finish shutting
|
||||
// down inside the close budget and has been ABANDONED (it may still be running).
|
||||
//
|
||||
// It is deliberately NOT propagated out of Apply. The swap it accompanies
|
||||
// succeeded — the new box is built, started and carrying traffic — and returning
|
||||
// a failure there would make the caller abort the netplane stage and, with no
|
||||
// engine fault to point at, leave a stale ruleset loaded over a healthy engine.
|
||||
// The fact is surfaced through PendingCloses() instead, which reaches
|
||||
// `shaterd status` and the panel and clears itself the moment the shutdown
|
||||
// finally completes. Close()/Teardown DO return it: there the failure to stop is
|
||||
// the whole point of the operation.
|
||||
var ErrCloseTimeout = errors.New("engine instance did not shut down within the close budget")
|
||||
|
||||
// defaultCloseBudget is the HARD wall-clock budget for retiring one superseded
|
||||
// box.
|
||||
//
|
||||
// Five seconds, for three reasons that all point at the same number:
|
||||
//
|
||||
// - it is exactly sing-box's own C.StopTimeout — the threshold at which the
|
||||
// engine itself declares a single lifecycle stop to be taking too long. A
|
||||
// close that blows past the budget upstream considers excessive is by
|
||||
// definition not a slow close, it is a stuck one;
|
||||
// - it stays under C.FatalStopTimeout (10s), which is when upstream's CLI gives
|
||||
// up and calls the process unclosable. We want to have already reacted by
|
||||
// then;
|
||||
// - it is paid AFTER the replacement box is started and serving, and only on an
|
||||
// apply that really changed something (the hash gate makes a no-op reconcile
|
||||
// free), so the worst case is five seconds added to one apply — not to the
|
||||
// every-minute cron reconcile, and never to a status poll.
|
||||
const defaultCloseBudget = 5 * time.Second
|
||||
|
||||
// closeBudget/closeBoxFn are package-wide seams. Production never touches them;
|
||||
// tests in this package and in shater/apply install a short budget and a
|
||||
// deliberately hung closer to exercise the abandoned-generation path without
|
||||
// standing up a box that really refuses to die.
|
||||
var (
|
||||
seamMu sync.RWMutex
|
||||
closeBudget = defaultCloseBudget
|
||||
closeBoxFn = func(c io.Closer) error { return c.Close() }
|
||||
)
|
||||
|
||||
// SetCloseBudget overrides the close budget and returns a function restoring the
|
||||
// previous value. TEST SEAM — production uses defaultCloseBudget.
|
||||
func SetCloseBudget(d time.Duration) (restore func()) {
|
||||
seamMu.Lock()
|
||||
prev := closeBudget
|
||||
closeBudget = d
|
||||
seamMu.Unlock()
|
||||
return func() {
|
||||
seamMu.Lock()
|
||||
closeBudget = prev
|
||||
seamMu.Unlock()
|
||||
}
|
||||
}
|
||||
|
||||
// SetBoxCloser overrides HOW a superseded instance is closed and returns a
|
||||
// function restoring the previous closer. TEST SEAM — production calls
|
||||
// Box.Close. A closer that never returns is how the abandoned-generation path is
|
||||
// tested.
|
||||
func SetBoxCloser(fn func(io.Closer) error) (restore func()) {
|
||||
seamMu.Lock()
|
||||
prev := closeBoxFn
|
||||
closeBoxFn = fn
|
||||
seamMu.Unlock()
|
||||
return func() {
|
||||
seamMu.Lock()
|
||||
closeBoxFn = prev
|
||||
seamMu.Unlock()
|
||||
}
|
||||
}
|
||||
|
||||
func currentCloseBudget() time.Duration {
|
||||
seamMu.RLock()
|
||||
defer seamMu.RUnlock()
|
||||
return closeBudget
|
||||
}
|
||||
|
||||
func currentBoxCloser() func(io.Closer) error {
|
||||
seamMu.RLock()
|
||||
defer seamMu.RUnlock()
|
||||
return closeBoxFn
|
||||
}
|
||||
|
||||
// pendingClose tracks ONE retirement in flight. It lives in Engine.pending from
|
||||
// the moment the retirement starts until Close returns, and `abandoned` records
|
||||
// whether the budget expired while it was still running — i.e. whether this
|
||||
// generation is a leak the operator must be told about, or merely a close that
|
||||
// is in progress under a lock the caller is holding anyway.
|
||||
type pendingClose struct {
|
||||
gen uint64
|
||||
hash string // config hash the retired box was built from
|
||||
since time.Time
|
||||
done chan struct{} // closed when the underlying Close finally returns
|
||||
abandoned bool // guarded by Engine.pendingMu
|
||||
finished bool // guarded by Engine.pendingMu
|
||||
}
|
||||
|
||||
// StuckClose is the read-side view of a superseded generation that overran the
|
||||
// close budget and is STILL running inside this process.
|
||||
type StuckClose struct {
|
||||
// Generation is the 1-based sequence number of the box that will not die.
|
||||
// It matches nothing in the sing-box log by itself, but it lets two status
|
||||
// reads tell "the same old generation" from "another one just leaked".
|
||||
Generation uint64
|
||||
// Hash is the config hash the leaked box was built from — the same value
|
||||
// `shaterd status` reported as `hash` while that config was the running one.
|
||||
// It is what connects "an engine is stuck" to WHICH configuration is stuck.
|
||||
Hash string
|
||||
// Since is when its shutdown was started.
|
||||
Since time.Time
|
||||
// Elapsed is how long it has been shutting down, as of the read.
|
||||
Elapsed time.Duration
|
||||
}
|
||||
|
||||
// retireLocked tears down a superseded box under the close budget. The caller
|
||||
// holds e.mu and must clear e.instance itself.
|
||||
//
|
||||
// Order matters and mirrors upstream: cancel the box's context FIRST so every
|
||||
// goroutine keyed on it unwinds, and only then call Close, which is what actually
|
||||
// releases the listeners, the WireGuard devices and the cache-file lock.
|
||||
//
|
||||
// Returns nil on a clean shutdown, ErrCloseTimeout when the budget expired (the
|
||||
// close keeps running in the background and is reaped when it finishes), or the
|
||||
// close's own error.
|
||||
func (e *Engine) retireLocked(b *box.Box, cancel func(), gen uint64, hash string) error {
|
||||
if cancel != nil {
|
||||
cancel()
|
||||
}
|
||||
if b == nil {
|
||||
return nil
|
||||
}
|
||||
|
||||
pc := &pendingClose{gen: gen, hash: hash, since: time.Now(), done: make(chan struct{})}
|
||||
e.pendingMu.Lock()
|
||||
e.pending = append(e.pending, pc)
|
||||
e.pendingMu.Unlock()
|
||||
|
||||
closeFn := currentBoxCloser()
|
||||
errc := make(chan error, 1)
|
||||
go func() {
|
||||
err := closeFn(b)
|
||||
e.pendingMu.Lock()
|
||||
pc.finished = true
|
||||
wasAbandoned := pc.abandoned
|
||||
e.removePendingLocked(pc)
|
||||
e.pendingMu.Unlock()
|
||||
if wasAbandoned && e.log != nil {
|
||||
// The counterpart of the ERROR below: an operator who saw the leak
|
||||
// reported must be able to see it resolve without restarting anything.
|
||||
e.log.Warn("engine generation ", gen, " finally finished shutting down after ",
|
||||
time.Since(pc.since).Round(time.Millisecond), "; it is no longer running")
|
||||
}
|
||||
close(pc.done)
|
||||
errc <- err
|
||||
}()
|
||||
|
||||
budget := currentCloseBudget()
|
||||
timer := time.NewTimer(budget)
|
||||
defer timer.Stop()
|
||||
select {
|
||||
case err := <-errc:
|
||||
return err
|
||||
case <-timer.C:
|
||||
}
|
||||
|
||||
// The budget expired. Claim the generation as abandoned — unless the close
|
||||
// happened to land in the same instant, in which case there is nothing to
|
||||
// report and we take its real result.
|
||||
e.pendingMu.Lock()
|
||||
abandoned := !pc.finished
|
||||
if abandoned {
|
||||
pc.abandoned = true
|
||||
}
|
||||
e.pendingMu.Unlock()
|
||||
if !abandoned {
|
||||
return <-errc
|
||||
}
|
||||
|
||||
if e.log != nil {
|
||||
e.log.Error("engine generation ", gen, " (config ", shortHash(hash), ") did NOT shut down within ", budget,
|
||||
" and has been ABANDONED: it may still hold its outbound connections and its ",
|
||||
"WireGuard devices (two devices with the same private key evict each other at ",
|
||||
"the peer), and it keeps writing to this log. The new configuration is running; ",
|
||||
"restart shaterd if this generation never clears.")
|
||||
}
|
||||
return ErrCloseTimeout
|
||||
}
|
||||
|
||||
// shortHash renders a config hash the way the log and the panel both need it:
|
||||
// enough to identify the configuration, short enough to read in a syslog line.
|
||||
func shortHash(h string) string {
|
||||
if len(h) < 12 {
|
||||
return "unknown"
|
||||
}
|
||||
return h[:12]
|
||||
}
|
||||
|
||||
// removePendingLocked drops pc from the pending list. Caller holds pendingMu.
|
||||
func (e *Engine) removePendingLocked(pc *pendingClose) {
|
||||
for i, p := range e.pending {
|
||||
if p == pc {
|
||||
e.pending = append(e.pending[:i], e.pending[i+1:]...)
|
||||
return
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// pendingSnapshot copies the in-flight retirements. Leaf lock only.
|
||||
func (e *Engine) pendingSnapshot() []*pendingClose {
|
||||
e.pendingMu.Lock()
|
||||
defer e.pendingMu.Unlock()
|
||||
return append([]*pendingClose(nil), e.pending...)
|
||||
}
|
||||
|
||||
// PendingCloses lists the superseded generations that overran their close budget
|
||||
// and are still running. Empty is the healthy answer.
|
||||
//
|
||||
// It takes ONLY the pending leaf lock — never e.mu — so the panel and
|
||||
// `shaterd status` can report a leaking teardown even while the apply that
|
||||
// produced it is still holding the engine mutex. That is deliberate: the one
|
||||
// moment this information matters most is while an apply is stalled behind a
|
||||
// shutdown that will not finish.
|
||||
func (e *Engine) PendingCloses() []StuckClose {
|
||||
now := time.Now()
|
||||
e.pendingMu.Lock()
|
||||
defer e.pendingMu.Unlock()
|
||||
out := make([]StuckClose, 0, len(e.pending))
|
||||
for _, p := range e.pending {
|
||||
if !p.abandoned {
|
||||
continue
|
||||
}
|
||||
out = append(out, StuckClose{
|
||||
Generation: p.gen,
|
||||
Hash: p.hash,
|
||||
Since: p.since,
|
||||
Elapsed: now.Sub(p.since),
|
||||
})
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// Generations reports how many sing-box instances this process is carrying: the
|
||||
// running one (0 or 1) plus every superseded generation whose shutdown has not
|
||||
// finished. ONE is the healthy answer for a started engine; anything above it is
|
||||
// the leak this file exists to make impossible to hide.
|
||||
func (e *Engine) Generations() int {
|
||||
e.mu.Lock()
|
||||
live := 0
|
||||
if e.instance != nil {
|
||||
live = 1
|
||||
}
|
||||
e.mu.Unlock()
|
||||
e.pendingMu.Lock()
|
||||
defer e.pendingMu.Unlock()
|
||||
return live + len(e.pending)
|
||||
}
|
||||
|
||||
// awaitAbandonedLocked gives every already-abandoned generation ONE more close
|
||||
// budget to finish, and returns how many are still running afterwards. The caller
|
||||
// holds e.mu and is about to build another instance.
|
||||
//
|
||||
// This is what keeps generations from stacking. Without it, a box that will not
|
||||
// come down means the NEXT apply quietly adds a third instance to the process,
|
||||
// and the one after that a fourth — which is precisely the accumulation observed
|
||||
// on the router. The wait is bounded by one budget for the whole set (not per
|
||||
// entry), so a permanently stuck generation costs a single extra budget on an
|
||||
// apply that actually changes the config, and nothing at all on the every-minute
|
||||
// no-op reconcile, which never gets past the hash gate.
|
||||
func (e *Engine) awaitAbandonedLocked() int {
|
||||
pending := e.pendingSnapshot()
|
||||
if len(pending) == 0 {
|
||||
return 0
|
||||
}
|
||||
deadline := time.NewTimer(currentCloseBudget())
|
||||
defer deadline.Stop()
|
||||
for _, p := range pending {
|
||||
select {
|
||||
case <-p.done:
|
||||
case <-deadline.C:
|
||||
return len(e.PendingCloses())
|
||||
}
|
||||
}
|
||||
return len(e.PendingCloses())
|
||||
}
|
||||
@@ -0,0 +1,247 @@
|
||||
// lx: pins the invariant that a chain copy of an AmneziaWG node carries the
|
||||
// SAME obfuscation parameters as the base node, field for field.
|
||||
//
|
||||
// Context: on the router a node placed behind an egress hop
|
||||
// (egress:ewan -> node:awgout, tag chain-<name>-h1) handshook forever and never
|
||||
// passed a transport packet, while the same node standalone was healthy. One
|
||||
// candidate explanation was that rebuildNode loses AWG params on the copy — it
|
||||
// does not (both the base and the copy go through the same
|
||||
// builder.wireguardEndpoint / amneziaOptions, the copy differing only in Tag and
|
||||
// DialerOptions.Detour), and this test holds that line. The real cause was the
|
||||
// bind swap the detour triggers: a detour makes Endpoint.Start pick ClientBind
|
||||
// instead of conn.StdNetBind, and ClientBind unconditionally overwrote bytes 1-3
|
||||
// of every datagram, shredding the h4 magic header of transport packets. See
|
||||
// transport/wireguard/client_bind.go (hasReserved) and its regression tests.
|
||||
//
|
||||
// Honest note: this test passes both before and after that fix — it is a pin on
|
||||
// a path that was never broken, not the reproducer for the bug.
|
||||
//
|
||||
// Unlike generate_test.go this file is NOT Linux-gated: it stops at
|
||||
// GenerateWithWarnings and never calls engine.Apply / box.New, so it needs no
|
||||
// routing_mark validation and runs on every platform.
|
||||
package generate
|
||||
|
||||
import (
|
||||
"encoding/base64"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"reflect"
|
||||
"testing"
|
||||
|
||||
"github.com/sagernet/sing-box/option"
|
||||
"github.com/sagernet/sing-box/shater/model"
|
||||
)
|
||||
|
||||
// awgKey returns a valid 32-byte base64 WireGuard key seeded by fill. Local to
|
||||
// this file so it does not depend on the Linux-only suite's helpers.
|
||||
func awgKey(fill byte) string {
|
||||
b := make([]byte, 32)
|
||||
for i := range b {
|
||||
b[i] = fill + byte(i)
|
||||
}
|
||||
return base64.StdEncoding.EncodeToString(b)
|
||||
}
|
||||
|
||||
// TestChainCopyPreservesAmneziaWGOptions builds a chain whose terminal hop is an
|
||||
// AmneziaWG node and asserts the chain copy's AmneziaWGOptions equals the base
|
||||
// endpoint's, comparing every field of the struct (so a field added later
|
||||
// without being mapped in amneziaOptions is caught here too).
|
||||
func TestChainCopyPreservesAmneziaWGOptions(t *testing.T) {
|
||||
priv := awgKey(1)
|
||||
pub := awgKey(9)
|
||||
// Every knob the option struct carries that the share-link can express:
|
||||
// jc/jmin/jmax, s1-s3, ranged h1-h4 (AWG 2.0), i1-i5. s4 is deliberately left
|
||||
// unset (0) — that is the real-world shape in which the transport magic lands
|
||||
// in bytes 0-3 and the ClientBind bug bit.
|
||||
uri := fmt.Sprintf(
|
||||
"awg://%s@203.0.113.10:51820?publickey=%s&address=10.13.13.2/32&allowedips=0.0.0.0/0"+
|
||||
"&jc=4&jmin=40&jmax=70&s1=86&s2=57&s3=13"+
|
||||
"&h1=1618116899-1618116949&h2=1795397486-1795397536&h3=3333333333&h4=1618116899-1618116949"+
|
||||
"&i1=%s&i2=%s&i3=%s&i4=%s&i5=%s#awgout",
|
||||
url.QueryEscape(priv), url.QueryEscape(pub),
|
||||
url.QueryEscape("<b 0xf0>"), url.QueryEscape("<c>"), url.QueryEscape("<t>"),
|
||||
url.QueryEscape("<r 10>"), url.QueryEscape("<b 0xab>"),
|
||||
)
|
||||
|
||||
node := model.Node{Name: "awgout", Enabled: true, URI: uri}
|
||||
|
||||
// Two SEPARATE configs, not one. wgdedup refuses to materialise a WireGuard
|
||||
// node twice from one private key (two devices sharing a key evict each other's
|
||||
// session), so a single config can hold either the base endpoint or the chain
|
||||
// copy — never both. That is also why on the router this node exists only as
|
||||
// "chain-<name>-h1". The invariant under test is therefore cross-config: the
|
||||
// same node, referenced directly vs referenced through a chain, must yield the
|
||||
// same AmneziaWG parameters.
|
||||
direct := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{node},
|
||||
Rules: []model.Rule{
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "node:awgout"},
|
||||
},
|
||||
}
|
||||
chained := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{node},
|
||||
Egresses: []model.Egress{
|
||||
{Name: "ewan", Type: "interface", Interface: "wan"},
|
||||
},
|
||||
Chains: []model.Chain{
|
||||
{Name: "viaewan", Hops: []string{"egress:ewan", "node:awgout"}},
|
||||
},
|
||||
Rules: []model.Rule{
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "chain:viaewan"},
|
||||
},
|
||||
}
|
||||
|
||||
directOpts, warns, err := GenerateWithWarnings(direct)
|
||||
if err != nil {
|
||||
t.Fatalf("GenerateWithWarnings(direct): %v (warnings: %v)", err, warns)
|
||||
}
|
||||
chainedOpts, chainWarns, err := GenerateWithWarnings(chained)
|
||||
if err != nil {
|
||||
t.Fatalf("GenerateWithWarnings(chained): %v (warnings: %v)", err, chainWarns)
|
||||
}
|
||||
|
||||
base, ok := awgOptionsForTag(directOpts, "awgout")
|
||||
if !ok {
|
||||
t.Fatalf("base endpoint %q not found; tags = %v (warnings: %v)",
|
||||
"awgout", endpointTags(directOpts), warns)
|
||||
}
|
||||
if !base.IsSet() {
|
||||
t.Fatalf("base endpoint has no AmneziaWG params: %+v", base)
|
||||
}
|
||||
|
||||
// Absolute expectation, not just parity. A "copy == base" assertion alone is
|
||||
// satisfied when BOTH lose a field (they share amneziaOptions), so pin the
|
||||
// literal values the share-link carries. Adding a field to
|
||||
// option.AmneziaWGOptions without mapping it in amneziaOptions fails the
|
||||
// exhaustiveness check below.
|
||||
want := option.AmneziaWGOptions{
|
||||
Jc: 4, Jmin: 40, Jmax: 70,
|
||||
S1: 86, S2: 57, S3: 13, S4: 0,
|
||||
H1: "1618116899-1618116949",
|
||||
H2: "1795397486-1795397536",
|
||||
H3: "3333333333",
|
||||
H4: "1618116899-1618116949",
|
||||
I1: "<b 0xf0>", I2: "<c>", I3: "<t>", I4: "<r 10>", I5: "<b 0xab>",
|
||||
}
|
||||
if base != want {
|
||||
t.Fatalf("base AmneziaWG params lost/garbled in parsing:\n got %+v\nwant %+v", base, want)
|
||||
}
|
||||
// Exhaustiveness: every field of the struct must be exercised above, so a
|
||||
// newly added knob cannot slip through unmapped and unnoticed.
|
||||
assertAWGFieldsCovered(t, want)
|
||||
|
||||
// The chain hop copy: buildHopWrapper tags it chain-<chain>-h<idx>. Find it as
|
||||
// "the wireguard endpoint in the chained config" rather than hardcoding the
|
||||
// index, so a change in hop numbering does not silently no-op this test.
|
||||
var (
|
||||
copyTag string
|
||||
copyOptions option.AmneziaWGOptions
|
||||
found bool
|
||||
)
|
||||
for _, endpoint := range chainedOpts.Endpoints {
|
||||
wg, isWG := endpoint.Options.(*option.WireGuardEndpointOptions)
|
||||
if !isWG {
|
||||
continue
|
||||
}
|
||||
copyTag, copyOptions, found = endpoint.Tag, wg.AmneziaWGOptions, true
|
||||
break
|
||||
}
|
||||
if !found {
|
||||
t.Fatalf("no chain copy endpoint emitted; endpoint tags = %v (warnings: %v)",
|
||||
endpointTags(chainedOpts), chainWarns)
|
||||
}
|
||||
if copyTag == "awgout" {
|
||||
t.Fatalf("expected a per-hop chain copy tag, got the base tag %q", copyTag)
|
||||
}
|
||||
|
||||
// Field-by-field, via reflection: any field of AmneziaWGOptions the copy fails
|
||||
// to carry is named explicitly rather than hidden behind one struct diff.
|
||||
baseValue := reflect.ValueOf(base)
|
||||
copyValue := reflect.ValueOf(copyOptions)
|
||||
for i := 0; i < baseValue.NumField(); i++ {
|
||||
field := baseValue.Type().Field(i)
|
||||
want := baseValue.Field(i).Interface()
|
||||
got := copyValue.Field(i).Interface()
|
||||
if !reflect.DeepEqual(want, got) {
|
||||
t.Errorf("chain copy %q lost AmneziaWG field %s: got %#v, want %#v",
|
||||
copyTag, field.Name, got, want)
|
||||
}
|
||||
}
|
||||
if !t.Failed() && base != copyOptions {
|
||||
t.Fatalf("chain copy %q AmneziaWGOptions differ from base: %+v vs %+v",
|
||||
copyTag, copyOptions, base)
|
||||
}
|
||||
|
||||
// The copy must additionally differ from the base in exactly the way the chain
|
||||
// intends: it detours through the egress hop, the base does not.
|
||||
baseDialer, okBase := dialerForTag(directOpts, "awgout")
|
||||
copyDialer, okCopy := dialerForTag(chainedOpts, copyTag)
|
||||
if !okBase || !okCopy {
|
||||
t.Fatalf("dialer options missing: base=%v copy=%v", okBase, okCopy)
|
||||
}
|
||||
if baseDialer.Detour != "" {
|
||||
t.Errorf("base endpoint must dial directly, got Detour=%q", baseDialer.Detour)
|
||||
}
|
||||
if copyDialer.Detour == "" {
|
||||
t.Errorf("chain copy %q must carry the egress hop as Detour, got empty", copyTag)
|
||||
}
|
||||
}
|
||||
|
||||
// assertAWGFieldsCovered fails if any field of option.AmneziaWGOptions is left
|
||||
// at its zero value in the expectation, other than the ones deliberately unset:
|
||||
//
|
||||
// S4 — kept 0 on purpose; that is the shape in which the transport magic
|
||||
// lands in bytes 0-3, i.e. the configuration that broke on the router.
|
||||
// Id/Ip/Ib — WireSock masquerade sugar, mutually exclusive with an explicit I1
|
||||
// (device_awg.go rejects the combination), and I1 is set here.
|
||||
//
|
||||
// The point is that adding a knob to the struct without extending this test
|
||||
// turns into a failure here rather than silent non-coverage.
|
||||
func assertAWGFieldsCovered(t *testing.T, want option.AmneziaWGOptions) {
|
||||
t.Helper()
|
||||
deliberatelyUnset := map[string]bool{"S4": true, "Id": true, "Ip": true, "Ib": true}
|
||||
value := reflect.ValueOf(want)
|
||||
for i := 0; i < value.NumField(); i++ {
|
||||
name := value.Type().Field(i).Name
|
||||
if deliberatelyUnset[name] {
|
||||
continue
|
||||
}
|
||||
if value.Field(i).IsZero() {
|
||||
t.Errorf("AmneziaWGOptions.%s is not exercised by this test "+
|
||||
"(add it to the share-link and to `want`, or to deliberatelyUnset)", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// awgOptionsForTag returns the AmneziaWGOptions of the wireguard endpoint tagged
|
||||
// tag.
|
||||
func awgOptionsForTag(opts option.Options, tag string) (option.AmneziaWGOptions, bool) {
|
||||
for _, endpoint := range opts.Endpoints {
|
||||
if endpoint.Tag != tag {
|
||||
continue
|
||||
}
|
||||
wg, ok := endpoint.Options.(*option.WireGuardEndpointOptions)
|
||||
if !ok {
|
||||
return option.AmneziaWGOptions{}, false
|
||||
}
|
||||
return wg.AmneziaWGOptions, true
|
||||
}
|
||||
return option.AmneziaWGOptions{}, false
|
||||
}
|
||||
|
||||
// dialerForTag returns the DialerOptions of the wireguard endpoint tagged tag.
|
||||
func dialerForTag(opts option.Options, tag string) (option.DialerOptions, bool) {
|
||||
for _, endpoint := range opts.Endpoints {
|
||||
if endpoint.Tag != tag {
|
||||
continue
|
||||
}
|
||||
wg, ok := endpoint.Options.(*option.WireGuardEndpointOptions)
|
||||
if !ok {
|
||||
return option.DialerOptions{}, false
|
||||
}
|
||||
return wg.DialerOptions, true
|
||||
}
|
||||
return option.DialerOptions{}, false
|
||||
}
|
||||
@@ -185,7 +185,15 @@ func TestDNSFilterRemoteBlocklistHTTPClient(t *testing.T) {
|
||||
{Name: "cf", Type: "doh", Address: "https://1.1.1.1/dns-query", Detour: "direct"},
|
||||
},
|
||||
Blocklists: []model.Blocklist{
|
||||
{Name: "remote-ads", Enabled: true, Source: "url", URL: srv.URL, Response: "nxdomain", UpdateInterval: "24h"},
|
||||
// The ".srs" suffix is LOAD-BEARING, not decoration: ruleSetURLIsEngineNative
|
||||
// (ruleset.go) decides remote-vs-compiled-local by URL EXTENSION alone, and
|
||||
// this test is about the REMOTE path — the engine fetching the compiled set
|
||||
// itself through the direct outbound. httptest.NewServer's bare
|
||||
// "http://127.0.0.1:<port>" has no extension, so it fell into the TEXT-list
|
||||
// path instead: the list was downloaded by generate's own listFetcher, parsed
|
||||
// as a hosts file and compiled into a LOCAL rule-set, which every assertion
|
||||
// below then contradicted. Do not trim the suffix.
|
||||
{Name: "remote-ads", Enabled: true, Source: "url", URL: srv.URL + "/blocklist.srs", Response: "nxdomain", UpdateInterval: "24h"},
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@@ -310,6 +310,19 @@ func GenerateWithWarningsAt(m *model.Model, now time.Time) (option.Options, []st
|
||||
Route: route,
|
||||
DNS: dns,
|
||||
}
|
||||
|
||||
// LAST, on the finished config: fold away duplicate WireGuard devices.
|
||||
//
|
||||
// A WG/AWG node may be materialised several times over — the always-emitted
|
||||
// base endpoint (outbound.go), a per-chain hop copy and a chain group hop's
|
||||
// per-member copy (chain.go) — and unlike a TCP proxy copy, each of those is a
|
||||
// real device holding the SAME private key. A WireGuard peer keeps one session
|
||||
// per key, so the copies evict each other and none of them passes traffic:
|
||||
// every chain containing a WG node was dead for exactly this reason. Running
|
||||
// on the assembled opts (rather than at each producer) is what makes the
|
||||
// guarantee hold for all three paths and any future fourth. See wgdedup.go.
|
||||
b.dedupWireGuardEndpoints(&opts)
|
||||
|
||||
return opts, b.warnings, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -227,6 +227,12 @@ func TestAllReachableProtocols(t *testing.T) {
|
||||
{Name: "vmess-ws", Enabled: true, URI: vmessLink},
|
||||
{Name: "trojan1", Enabled: true, URI: "trojan://password@example.com:443?sni=example.com#trojan1"},
|
||||
{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"},
|
||||
// QUIC protocols travel the same road: share-link -> parse.Proxy ->
|
||||
// outbound -> box.New. Before parse grew these two branches the links
|
||||
// were dropped at the parser and the outbounds below never existed.
|
||||
{Name: "hy2-1", Enabled: true, URI: "hysteria2://secret@example.com:8443/?sni=example.com&alpn=h3#hy2-1"},
|
||||
{Name: "hy2-alias", Enabled: true, URI: "hy2://secret@example.org#hy2-alias"},
|
||||
{Name: "tuic1", Enabled: true, URI: "tuic://22222222-2222-2222-2222-222222222222:secret@example.com:443/?sni=example.com&alpn=h3&congestion_control=cubic&udp_relay_mode=native#tuic1"},
|
||||
},
|
||||
}
|
||||
opts, warns, changed := applyAndClose(t, m)
|
||||
@@ -236,16 +242,34 @@ func TestAllReachableProtocols(t *testing.T) {
|
||||
if len(warns) != 0 {
|
||||
t.Fatalf("unexpected warnings: %v", warns)
|
||||
}
|
||||
for _, tag := range []string{"vless-ws", "vless-grpc", "vmess-ws", "trojan1", "ss1"} {
|
||||
for _, tag := range []string{"vless-ws", "vless-grpc", "vmess-ws", "trojan1", "ss1", "hy2-1", "hy2-alias", "tuic1"} {
|
||||
if findOutbound(opts, tag) == nil {
|
||||
t.Fatalf("outbound %q not emitted", tag)
|
||||
}
|
||||
}
|
||||
// The hy2 alias link carries no port and no sni: parse must have supplied the
|
||||
// scheme default (443) and the server name, or the node would dial nowhere.
|
||||
if ob := findOutbound(opts, "hy2-alias"); ob != nil {
|
||||
o, ok := ob.Options.(*option.Hysteria2OutboundOptions)
|
||||
if !ok {
|
||||
t.Fatalf("hy2-alias options type = %T", ob.Options)
|
||||
}
|
||||
if o.ServerPort != 443 {
|
||||
t.Fatalf("hy2-alias port = %d, want the scheme default 443", o.ServerPort)
|
||||
}
|
||||
if o.TLS == nil || o.TLS.ServerName != "example.org" {
|
||||
t.Fatalf("hy2-alias tls = %+v", o.TLS)
|
||||
}
|
||||
if o.TLS.UTLS != nil {
|
||||
t.Fatalf("uTLS must never reach a QUIC outbound: %+v", o.TLS.UTLS)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- White-box: hysteria2 / tuic / shadowtls option mapping validates. -------
|
||||
// These protocols are not yet produced by parse.ParseShareLink, so we drive the
|
||||
// mapping directly with synthetic parse.Proxy values and validate via box.New.
|
||||
// shadowtls is not produced by parse.ParseShareLink (hysteria2 and tuic now
|
||||
// are, see TestAllReachableProtocols), so the mapping is driven directly with
|
||||
// synthetic parse.Proxy values and validated via box.New.
|
||||
|
||||
func TestQUICAndShadowTLSMappingValidates(t *testing.T) {
|
||||
b := newBuilder(&model.Model{Globals: model.DefaultGlobals()})
|
||||
|
||||
@@ -162,14 +162,23 @@ func TestObservatoryPlanFromGeneratedConfig(t *testing.T) {
|
||||
t.Errorf("chain exit %q store = %v, want only itself (a chain prefix is nobody else's path)", chainExit, store)
|
||||
}
|
||||
}
|
||||
// The chain's intermediate node hop n1 is used (the exit detours through it) but
|
||||
// NOT probed separately — its death is visible in the exit probe, and no
|
||||
// selection depends on it. n1 IS probed via the plain "auto" group, so the
|
||||
// assertion is "no SEPARATE chain-hop-h1 job", not "n1 never dialled".
|
||||
if jobDials(jobs, "chain-hop-h1") {
|
||||
t.Errorf("plan dials intermediate chain hop chain-hop-h1 — only the exit is probed end-to-end")
|
||||
// The chain's intermediate NODE hop wrapper IS probed now, as a measurement
|
||||
// of its own: dialling chain-hop-h1 measures exactly the prefix of the path
|
||||
// up to and including hop 1, which is what lets an operator see WHICH hop
|
||||
// died instead of only that the chain did (engine/probeplan.go walkDetour,
|
||||
// surfaced through ChainHealth.Hops). Its store is only itself — a chain
|
||||
// prefix is nobody else's dial path, so the base node n1 must never inherit
|
||||
// a verdict measured through the chain.
|
||||
chainHop := "chain-hop-h1"
|
||||
if !jobDials(jobs, chainHop) {
|
||||
t.Errorf("plan does not probe intermediate chain hop %q; jobs=%v", chainHop, dialsOfJobs(jobs))
|
||||
} else {
|
||||
store := jobStore(jobs, chainHop)
|
||||
if len(store) != 1 || store[0] != chainHop {
|
||||
t.Errorf("chain hop %q store = %v, want only itself (no base alias for a chain prefix)", chainHop, store)
|
||||
}
|
||||
}
|
||||
if !used[chainExit] || !used["chain-hop-h1"] {
|
||||
if !used[chainExit] || !used[chainHop] {
|
||||
t.Errorf("used-set is missing chain hop tags; used=%v", used)
|
||||
}
|
||||
// The used groups themselves are in the used-set but never dialled (a group is
|
||||
|
||||
@@ -0,0 +1,236 @@
|
||||
//go:build linux
|
||||
|
||||
// The behavioural half of the build-tag contract (D23).
|
||||
//
|
||||
// shater/buildtags's test proves, statically and on any host, that the shipped
|
||||
// tag set (scripts/router-tags.sh) still NAMES every tag a declared feature
|
||||
// needs. That is necessary but not sufficient: a tag can be present and still
|
||||
// insufficient, and a tag list is only a hypothesis until something is built
|
||||
// with it. This file is the experiment — one node of every declared protocol,
|
||||
// driven through engine.Apply (box.New + Start) under WHATEVER tags the test
|
||||
// binary was compiled with.
|
||||
//
|
||||
// Run it with the shipped set via scripts/check-router-tags.sh (CI does, before
|
||||
// the artifact is built). Then the two halves compose:
|
||||
//
|
||||
// tag set covers the declared features (buildtags test, tag-less)
|
||||
// + everything compiled in constructs (this test, run WITH the shipped set)
|
||||
// = the binary we ship supports what we say it does.
|
||||
//
|
||||
// The 2026-07-25 WireGuard outage — `with_gvisor` trimmed while `with_wireguard`
|
||||
// stayed, so every shipped binary died with "gVisor is not included in this
|
||||
// build" on the first WireGuard node — is caught here at "wg"/"awg", because
|
||||
// box.New initialises endpoints and the WireGuard device constructor is the
|
||||
// gVisor stub without the tag.
|
||||
package generate
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"net/url"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
|
||||
"github.com/sagernet/sing-box/shater/buildtags"
|
||||
"github.com/sagernet/sing-box/shater/model"
|
||||
)
|
||||
|
||||
// protoCase is one declared protocol, expressed the way a user would add it: a
|
||||
// share link. feature keys it to a row of buildtags.Features (substring match on
|
||||
// the feature name) — that row supplies the build tags the case needs. An empty
|
||||
// feature means "compiled in unconditionally".
|
||||
type protoCase struct {
|
||||
tag string
|
||||
uri string
|
||||
feature string
|
||||
endpoint bool // arrives as an option.Endpoint, not an option.Outbound
|
||||
}
|
||||
|
||||
func shippedProtoCases(t *testing.T) []protoCase {
|
||||
t.Helper()
|
||||
const uuid = "11111111-1111-1111-1111-111111111111"
|
||||
// A vmess ws+tls link (same fixture the parse suite uses).
|
||||
const vmessLink = "vmess://eyJ2IjoiMiIsInBzIjoidm1lc3MtdyIsImFkZCI6ImV4YW1wbGUubmV0IiwicG9ydCI6IjQ0MyIsImlkIjoiMzMzMzMzMzMtMzMzMy0zMzMzLTMzMzMtMzMzMzMzMzMzMzMzIiwiYWlkIjoiMCIsInNjeSI6ImF1dG8iLCJuZXQiOiJ3cyIsImhvc3QiOiJleGFtcGxlLm5ldCIsInBhdGgiOiIvd3MiLCJ0bHMiOiJ0bHMifQ=="
|
||||
|
||||
priv, pub := validKey(1), validKey(9)
|
||||
wgURI := fmt.Sprintf(
|
||||
"wireguard://%s@203.0.113.10:51820?publickey=%s&address=10.13.13.2/32&allowedips=0.0.0.0/0#wg",
|
||||
url.QueryEscape(priv), url.QueryEscape(pub),
|
||||
)
|
||||
awgURI := fmt.Sprintf(
|
||||
"awg://%s@203.0.113.11:51820?publickey=%s&address=10.13.13.3/32&allowedips=0.0.0.0/0&jc=4&jmin=40&jmax=70&s1=30&s2=40&h1=1111111111&h2=2222222222&h3=3333333333&h4=444444444#awg",
|
||||
url.QueryEscape(priv), url.QueryEscape(pub),
|
||||
)
|
||||
|
||||
return []protoCase{
|
||||
// --- always compiled in -------------------------------------------
|
||||
{tag: "ss", uri: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss"},
|
||||
{tag: "vmess-ws", uri: vmessLink},
|
||||
{tag: "trojan", uri: "trojan://password@example.com:443?sni=example.com#trojan"},
|
||||
{tag: "vless-ws", uri: "vless://" + uuid + "@example.com:443?type=ws&security=tls&path=/vl&host=cdn.example.com&sni=cdn.example.com#vless-ws"},
|
||||
{tag: "vless-grpc", uri: "vless://" + uuid + "@example.com:443?type=grpc&security=tls&serviceName=gsvc&sni=example.com#vless-grpc"},
|
||||
{tag: "vless-httpupgrade", uri: "vless://" + uuid + "@example.com:443?type=httpupgrade&security=tls&path=/hu&host=cdn.example.com&sni=cdn.example.com#vless-httpupgrade"},
|
||||
|
||||
// --- tag-gated ------------------------------------------------------
|
||||
{tag: "vless-reality", feature: "REALITY",
|
||||
uri: "vless://" + uuid + "@example.com:443?type=tcp&security=reality&pbk=" + url.QueryEscape("jNXHt1yRo0vDuchQlIP6Z0ZvjT3KtzVI-T4E7RoLJS0") + "&sid=0123456789abcdef&sni=example.com&fp=chrome&flow=xtls-rprx-vision#vless-reality"},
|
||||
{tag: "vless-utls", feature: "uTLS ClientHello",
|
||||
uri: "vless://" + uuid + "@example.com:443?type=ws&security=tls&path=/u&sni=example.com&fp=firefox#vless-utls"},
|
||||
{tag: "vless-quic", feature: "QUIC v2ray transport",
|
||||
uri: "vless://" + uuid + "@example.com:443?type=quic&security=tls&sni=example.com#vless-quic"},
|
||||
{tag: "vless-xhttp", feature: "XHTTP",
|
||||
uri: "vless://" + uuid + "@example.com:443?type=xhttp&security=tls&path=/xh&host=cdn.example.com&sni=cdn.example.com#vless-xhttp"},
|
||||
{tag: "hy2", feature: "Hysteria2",
|
||||
uri: "hysteria2://secret@example.com:8443/?sni=example.com&alpn=h3#hy2"},
|
||||
{tag: "tuic", feature: "TUIC",
|
||||
uri: "tuic://22222222-2222-2222-2222-222222222222:secret@example.com:443/?sni=example.com&alpn=h3&congestion_control=cubic&udp_relay_mode=native#tuic"},
|
||||
{tag: "wg", feature: "WireGuard nodes", endpoint: true, uri: wgURI},
|
||||
{tag: "awg", feature: "AmneziaWG", endpoint: true, uri: awgURI},
|
||||
}
|
||||
}
|
||||
|
||||
// featuresWithoutProbe are declared features this file cannot express as a node,
|
||||
// with the reason. Anything NOT listed here must be exercised by a protoCase —
|
||||
// TestEveryTagGatedFeatureIsProbed enforces that, so a new tag-gated feature
|
||||
// cannot be added without either a probe or a conscious exemption.
|
||||
var featuresWithoutProbe = map[string]string{
|
||||
"QUIC / HTTP3 DNS transports (quic://, h3://)": "shater's resolver types are udp/tcp/doh/dot/local/fakeip — the model cannot express a quic:// resolver; the tag is already proven by the hysteria2/tuic/quic-transport cases",
|
||||
"badtls fast path (zero-copy TLS read-wait / ktls, used by every TLS outbound)": "a link-time/performance path, not a constructible config object; its absence is caught at link time (-checklinkname=0) and by the buildtags test",
|
||||
}
|
||||
|
||||
// featureFor resolves a protoCase's feature key to its buildtags row.
|
||||
func featureFor(t *testing.T, key string) buildtags.Feature {
|
||||
t.Helper()
|
||||
for _, f := range buildtags.Features {
|
||||
if strings.Contains(f.Name, key) {
|
||||
return f
|
||||
}
|
||||
}
|
||||
t.Fatalf("protoCase names feature %q, which is in no buildtags.Features row — keep the two lists linked", key)
|
||||
return buildtags.Feature{}
|
||||
}
|
||||
|
||||
// TestShippedTagSetConstructsDeclaredProtocols builds one node per declared
|
||||
// protocol and proves box.New+Start accepts every one of them under the tags
|
||||
// this binary was compiled with. Protocols whose tags are genuinely absent are
|
||||
// SKIPPED loudly (with the missing tags), never silently dropped — and the
|
||||
// shipped set is guaranteed to contain those tags by
|
||||
// shater/buildtags.TestRouterTagSetCoversDeclaredFeatures, so under
|
||||
// scripts/check-router-tags.sh nothing is skipped.
|
||||
func TestShippedTagSetConstructsDeclaredProtocols(t *testing.T) {
|
||||
t.Logf("compiled build tags: %v", buildtags.Compiled())
|
||||
|
||||
var (
|
||||
nodes []model.Node
|
||||
wantOut []string
|
||||
wantEnd []string
|
||||
skipped []string
|
||||
)
|
||||
for _, c := range shippedProtoCases(t) {
|
||||
if c.feature != "" {
|
||||
f := featureFor(t, c.feature)
|
||||
if missing := buildtags.MissingTags(f); len(missing) > 0 {
|
||||
skipped = append(skipped, fmt.Sprintf("%s (missing %s)", c.tag, strings.Join(missing, ",")))
|
||||
continue
|
||||
}
|
||||
}
|
||||
nodes = append(nodes, model.Node{Name: c.tag, Enabled: true, URI: c.uri})
|
||||
if c.endpoint {
|
||||
wantEnd = append(wantEnd, c.tag)
|
||||
} else {
|
||||
wantOut = append(wantOut, c.tag)
|
||||
}
|
||||
}
|
||||
if len(skipped) > 0 {
|
||||
// Under a plain `go test` (no tags) a protocol that is not compiled in
|
||||
// cannot be constructed — skipping is the only honest thing to do, and
|
||||
// shater/buildtags is what guards the tag list itself.
|
||||
//
|
||||
// Under scripts/check-router-tags.sh we were compiled with the SHIPPED
|
||||
// set, which buildtags has already proven covers every declared feature.
|
||||
// So a skip here means the two disagree — a trimmed tag, or a check run
|
||||
// with the wrong -tags. Either way the run must be red, not "PASS (2
|
||||
// protocols skipped)".
|
||||
if os.Getenv("SHATER_ROUTER_TAG_CHECK") == "1" {
|
||||
t.Fatalf("the SHIPPED tag set does not compile in %d declared protocol(s): %s\n"+
|
||||
"Nothing may be skipped in a router-tag-set run — add the missing tags to scripts/router-tags.sh.",
|
||||
len(skipped), strings.Join(skipped, "; "))
|
||||
}
|
||||
t.Logf("NOT compiled in, skipped: %s", strings.Join(skipped, "; "))
|
||||
}
|
||||
|
||||
// No inbound on purpose: this test is about protocol construction, and a
|
||||
// tproxy listener would demand CAP_NET_ADMIN from every runner. The tproxy
|
||||
// path is covered by the rest of the suite.
|
||||
m := &model.Model{Globals: model.DefaultGlobals(), Nodes: nodes}
|
||||
|
||||
opts, warns, changed := applyAndClose(t, m)
|
||||
if !changed {
|
||||
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
|
||||
}
|
||||
if len(warns) != 0 {
|
||||
t.Fatalf("a declared protocol produced generator warnings — it is not fully supported: %v", warns)
|
||||
}
|
||||
|
||||
for _, tag := range wantOut {
|
||||
if findOutbound(opts, tag) == nil {
|
||||
t.Errorf("outbound %q was not emitted (warnings: %v)", tag, warns)
|
||||
}
|
||||
}
|
||||
for _, tag := range wantEnd {
|
||||
var found *option.Endpoint
|
||||
for i := range opts.Endpoints {
|
||||
if opts.Endpoints[i].Tag == tag {
|
||||
found = &opts.Endpoints[i]
|
||||
}
|
||||
}
|
||||
if found == nil {
|
||||
t.Errorf("endpoint %q was not emitted (warnings: %v)", tag, warns)
|
||||
continue
|
||||
}
|
||||
if found.Type != C.TypeWireGuard {
|
||||
t.Errorf("endpoint %q type = %s, want wireguard", tag, found.Type)
|
||||
}
|
||||
}
|
||||
|
||||
// AmneziaWG is the driving requirement: if it is compiled in, the obfuscation
|
||||
// params must actually be on the endpoint, not parsed-and-dropped.
|
||||
if buildtags.Has("with_awg") {
|
||||
for i := range opts.Endpoints {
|
||||
if opts.Endpoints[i].Tag != "awg" {
|
||||
continue
|
||||
}
|
||||
wg, ok := opts.Endpoints[i].Options.(*option.WireGuardEndpointOptions)
|
||||
if !ok {
|
||||
t.Fatalf("awg endpoint options type = %T", opts.Endpoints[i].Options)
|
||||
}
|
||||
if !wg.AmneziaWGOptions.IsSet() || wg.Jc != 4 {
|
||||
t.Errorf("AmneziaWG params did not reach the endpoint: %+v", wg.AmneziaWGOptions)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestEveryTagGatedFeatureIsProbed keeps this file from rotting: a feature added
|
||||
// to buildtags.Features must be constructed here, or explicitly exempted with a
|
||||
// reason. Otherwise the next trimmed tag is invisible again.
|
||||
func TestEveryTagGatedFeatureIsProbed(t *testing.T) {
|
||||
probed := map[string]bool{}
|
||||
for _, c := range shippedProtoCases(t) {
|
||||
if c.feature != "" {
|
||||
probed[featureFor(t, c.feature).Name] = true
|
||||
}
|
||||
}
|
||||
for _, f := range buildtags.Features {
|
||||
if probed[f.Name] {
|
||||
continue
|
||||
}
|
||||
if _, ok := featuresWithoutProbe[f.Name]; ok {
|
||||
continue
|
||||
}
|
||||
t.Errorf("declared feature %q has no protoCase and no featuresWithoutProbe exemption: add one, or the tag it needs can be trimmed unnoticed", f.Name)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,252 @@
|
||||
package generate
|
||||
|
||||
// Where the router's traffic actually ENDS UP, read off the engine config that
|
||||
// was generated for it.
|
||||
//
|
||||
// WHY THIS EXISTS. The panel's headline readout was derived from apply.Status's
|
||||
// `plane` field, and `plane` answers a different question than the one the
|
||||
// readout asked. `plane` says how much of the DATA PLANE is installed — is the
|
||||
// `inet shater` table loaded, is policy routing in place, is the engine up. It
|
||||
// says nothing about where the diverted packets go once the engine has them.
|
||||
//
|
||||
// A router in the field ran with a single enabled rule, `default -> direct`, no
|
||||
// groups and no rule-sets. Every piece of the plane was installed, so
|
||||
// plane == "full", so the panel said "Protected — traffic from your network is
|
||||
// going through the tunnel". There was no tunnel: route.Final was `direct` and
|
||||
// the whole LAN went out the plain WAN with its real address, under a green LED.
|
||||
// That is the product's worst defect class — a silent lie in the reassuring
|
||||
// direction — so the verdict below is computed from the thing that actually
|
||||
// decides the answer.
|
||||
//
|
||||
// THE SOURCE IS THE GENERATED CONFIG, NOT THE DESIRED STATE. TrafficOf reads
|
||||
// option.Options: route.Final, the emitted route rules, and the outbound table
|
||||
// they name. That is the config the engine was handed, so the verdict already
|
||||
// accounts for everything that happens between "what the operator wrote" and
|
||||
// "what runs": a schedule outside its window (the rule is simply not emitted), a
|
||||
// catch-all shadowed by a later one (only the winner reached Final), a rule whose
|
||||
// target did not resolve and fell back to direct/block (ruleKillFallback), a rule
|
||||
// with no engine-evaluable matcher (dropped with a warning). Re-deriving any of
|
||||
// that from *model.Model would be a second implementation of buildRoute, and the
|
||||
// two would drift — which is exactly the failure mode model/reachability.go was
|
||||
// written to stop.
|
||||
|
||||
import (
|
||||
"net/netip"
|
||||
"strings"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
)
|
||||
|
||||
// Traffic verdicts. Four states, because collapsing them is how the readout
|
||||
// started lying in the first place.
|
||||
const (
|
||||
// VerdictTunnel — route.Final points into a tunnel, so everything that is not
|
||||
// matched by a more specific rule is proxied.
|
||||
VerdictTunnel = "tunnel"
|
||||
// VerdictSplit — the default leaves the router directly, but at least one rule
|
||||
// does send its traffic into a tunnel. Selective protection: legitimate and
|
||||
// common, but NOT "protected".
|
||||
VerdictSplit = "split"
|
||||
// VerdictDirect — the default leaves directly and nothing is tunnelled at all.
|
||||
// The engine is running and carrying traffic straight out. This is the field
|
||||
// case above.
|
||||
VerdictDirect = "direct"
|
||||
// VerdictBlocked — the default is `block`, the kill-switch backstop: unmatched
|
||||
// traffic is DROPPED, not let out. Nothing leaks; whether anything works at all
|
||||
// depends on TunnelRules.
|
||||
VerdictBlocked = "blocked"
|
||||
)
|
||||
|
||||
// Traffic is the verdict for one generated engine config.
|
||||
//
|
||||
// The zero value (Verdict == "") means "not known" — no config has been
|
||||
// generated/applied by this daemon yet. A consumer must not read it as any of the
|
||||
// four verdicts; in particular it is NOT "tunnel".
|
||||
type Traffic struct {
|
||||
// Verdict is one of the four constants above, or "" when unknown.
|
||||
Verdict string `json:"verdict"`
|
||||
// Default is the outbound tag route.Final names — the engine's own vocabulary
|
||||
// ("direct", "block", a node/group tag, a chain's entry-hop tag). Published for
|
||||
// `shaterd status` and debugging; the wording in the UI is driven by Verdict,
|
||||
// not by parsing this.
|
||||
Default string `json:"default"`
|
||||
// TunnelRules counts the emitted route rules whose matched traffic goes into a
|
||||
// tunnel. It is what separates "selective protection" from "none", and with
|
||||
// VerdictBlocked it separates "only the listed traffic gets out" from "nothing
|
||||
// gets out at all".
|
||||
TunnelRules int `json:"tunnel_rules"`
|
||||
}
|
||||
|
||||
// TrafficOf computes the verdict for a generated engine config.
|
||||
//
|
||||
// A nil Route yields the zero Traffic (unknown) rather than a guess: every config
|
||||
// this package produces has one, so a missing Route means the caller handed us
|
||||
// something we did not build.
|
||||
func TrafficOf(opts option.Options) Traffic {
|
||||
if opts.Route == nil {
|
||||
return Traffic{}
|
||||
}
|
||||
idx := indexOutbounds(opts)
|
||||
final := opts.Route.Final
|
||||
t := Traffic{Default: final}
|
||||
for _, r := range opts.Route.Rules {
|
||||
tag, ok := ruleRouteOutbound(r)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if idx.tunnels(tag) {
|
||||
t.TunnelRules++
|
||||
}
|
||||
}
|
||||
switch {
|
||||
case idx.tunnels(final):
|
||||
t.Verdict = VerdictTunnel
|
||||
case idx.kind[final] == kindBlock:
|
||||
t.Verdict = VerdictBlocked
|
||||
case t.TunnelRules > 0:
|
||||
t.Verdict = VerdictSplit
|
||||
default:
|
||||
// Includes an unknown/empty Final: an unrecognised default is reported as the
|
||||
// unprotected one. Erring toward the alarm is allowed here; erring toward
|
||||
// reassurance is the bug.
|
||||
t.Verdict = VerdictDirect
|
||||
}
|
||||
return t
|
||||
}
|
||||
|
||||
// Outbound kinds, from the point of view of "does traffic that leaves through
|
||||
// this tag leave the router by some path other than the plain WAN".
|
||||
const (
|
||||
kindDirect = iota // direct outbound, incl. every interface/plain egress
|
||||
kindBlock // block outbound
|
||||
kindGroup // selector/urltest: a tunnel iff one of its members is
|
||||
kindProxy // a real remote proxy outbound or endpoint
|
||||
)
|
||||
|
||||
// outboundIndex maps every outbound/endpoint tag to its kind, plus the member
|
||||
// lists of the group outbounds so a group can be resolved to what it dials.
|
||||
type outboundIndex struct {
|
||||
kind map[string]int
|
||||
members map[string][]string
|
||||
}
|
||||
|
||||
func indexOutbounds(opts option.Options) outboundIndex {
|
||||
idx := outboundIndex{
|
||||
kind: make(map[string]int, len(opts.Outbounds)+len(opts.Endpoints)),
|
||||
members: make(map[string][]string),
|
||||
}
|
||||
for _, ob := range opts.Outbounds {
|
||||
switch ob.Type {
|
||||
case C.TypeDirect:
|
||||
// Covers the baseline `direct` AND every interface/plain `egress-*`
|
||||
// outbound: an egress picks WHICH uplink the packet leaves by, never
|
||||
// whether it leaves in the clear.
|
||||
idx.kind[ob.Tag] = kindDirect
|
||||
case C.TypeBlock:
|
||||
idx.kind[ob.Tag] = kindBlock
|
||||
case C.TypeSelector:
|
||||
idx.kind[ob.Tag] = kindGroup
|
||||
if o, ok := ob.Options.(*option.SelectorOutboundOptions); ok && o != nil {
|
||||
idx.members[ob.Tag] = o.Outbounds
|
||||
}
|
||||
case C.TypeURLTest:
|
||||
idx.kind[ob.Tag] = kindGroup
|
||||
if o, ok := ob.Options.(*option.URLTestOutboundOptions); ok && o != nil {
|
||||
idx.members[ob.Tag] = o.Outbounds
|
||||
}
|
||||
default:
|
||||
if dialsLoopback(ob.Options) {
|
||||
// A proxy that dials this very router — the `byedpi` egress is a SOCKS5
|
||||
// outbound to 127.0.0.1, where the local ciadpi desync helper listens.
|
||||
// It reshapes the packets, which is worth having, but the connection
|
||||
// still leaves over the plain WAN with this router's own address. Calling
|
||||
// that "going through the tunnel" would be a smaller version of the same
|
||||
// lie this file exists to remove.
|
||||
idx.kind[ob.Tag] = kindDirect
|
||||
} else {
|
||||
idx.kind[ob.Tag] = kindProxy
|
||||
}
|
||||
}
|
||||
}
|
||||
// Endpoints share the outbound tag namespace and are all real tunnels today
|
||||
// (wireguard / AmneziaWG).
|
||||
for _, ep := range opts.Endpoints {
|
||||
idx.kind[ep.Tag] = kindProxy
|
||||
}
|
||||
return idx
|
||||
}
|
||||
|
||||
// tunnels reports whether traffic sent to tag leaves the router through a tunnel.
|
||||
// An unknown tag is NOT a tunnel: the honest answer to "we have never heard of
|
||||
// this outbound" is "we cannot claim it protects anything".
|
||||
func (idx outboundIndex) tunnels(tag string) bool {
|
||||
return idx.tunnelsSeen(tag, make(map[string]bool, 4))
|
||||
}
|
||||
|
||||
func (idx outboundIndex) tunnelsSeen(tag string, seen map[string]bool) bool {
|
||||
if seen[tag] {
|
||||
// A group cycle cannot survive box.New, but this function also runs against
|
||||
// hand-built options in tests; refusing to recurse forever is free.
|
||||
return false
|
||||
}
|
||||
seen[tag] = true
|
||||
switch idx.kind[tag] {
|
||||
case kindProxy:
|
||||
return true
|
||||
case kindGroup:
|
||||
// A group is a tunnel when ANY member is one. A group of nodes is the normal
|
||||
// case; a group that also holds `direct` still tunnels some of the time, and
|
||||
// "sometimes tunnelled" must not be reported as "never".
|
||||
for _, m := range idx.members[tag] {
|
||||
if idx.tunnelsSeen(m, seen) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
// ruleRouteOutbound returns the outbound tag a route rule sends matched traffic
|
||||
// to, and false for every rule that routes nothing — the leading sniff and
|
||||
// hijack-dns rules, and the DoH-block reject rules.
|
||||
func ruleRouteOutbound(r option.Rule) (string, bool) {
|
||||
var action option.RuleAction
|
||||
switch r.Type {
|
||||
case C.RuleTypeDefault:
|
||||
action = r.DefaultOptions.RuleAction
|
||||
case C.RuleTypeLogical:
|
||||
action = r.LogicalOptions.RuleAction
|
||||
default:
|
||||
return "", false
|
||||
}
|
||||
if action.Action != C.RuleActionTypeRoute {
|
||||
return "", false
|
||||
}
|
||||
tag := action.RouteOptions.Outbound
|
||||
return tag, tag != ""
|
||||
}
|
||||
|
||||
// dialsLoopback reports whether an outbound's server address is this machine.
|
||||
// Only the two protocols a local helper is ever wired as are inspected (SOCKS is
|
||||
// what `byedpi` emits; HTTP is here so the answer does not depend on which of the
|
||||
// two a future helper picks).
|
||||
func dialsLoopback(o any) bool {
|
||||
var server string
|
||||
switch v := o.(type) {
|
||||
case *option.SOCKSOutboundOptions:
|
||||
server = v.Server
|
||||
case *option.HTTPOutboundOptions:
|
||||
server = v.Server
|
||||
default:
|
||||
return false
|
||||
}
|
||||
server = strings.TrimSpace(server)
|
||||
if strings.EqualFold(server, "localhost") {
|
||||
return true
|
||||
}
|
||||
addr, err := netip.ParseAddr(server)
|
||||
return err == nil && addr.IsLoopback()
|
||||
}
|
||||
@@ -0,0 +1,233 @@
|
||||
// Where the traffic actually ends up (TrafficOf). These are pure option-struct
|
||||
// assertions over a generated config — no box.New — so they build everywhere.
|
||||
//
|
||||
// The case that matters most is TestTrafficProductionDefaultDirectIsNotProtected:
|
||||
// it is the config found on a live router (one enabled rule, `default -> direct`,
|
||||
// no groups, no rule-sets) that the panel painted green and called "Protected".
|
||||
package generate
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
|
||||
"github.com/sagernet/sing-box/shater/model"
|
||||
)
|
||||
|
||||
// trafficOf generates m and returns the verdict, failing on a generate error.
|
||||
func trafficOf(t *testing.T, m *model.Model) Traffic {
|
||||
t.Helper()
|
||||
opts, _, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("generate: %v", err)
|
||||
}
|
||||
return TrafficOf(opts)
|
||||
}
|
||||
|
||||
// ssNode is a parseable shadowsocks node — enough for a real proxy outbound.
|
||||
func ssNode(name string) model.Node {
|
||||
return model.Node{Name: name, Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#" + name}
|
||||
}
|
||||
|
||||
// --- the field case ---------------------------------------------------------
|
||||
|
||||
// TestTrafficProductionDefaultDirectIsNotProtected is the regression for the
|
||||
// defect this file exists for. mini_router v0.2.10-r1 ran with exactly this
|
||||
// config: ONE enabled rule named `default`, order 100, target `direct`, zero
|
||||
// rule-sets and zero groups. Every part of the data plane was installed, so
|
||||
// status reported plane="full", and the panel's headline read
|
||||
// "Protected — traffic from your network is going through the tunnel".
|
||||
//
|
||||
// route.Final was `direct`. Nothing was tunnelled at all.
|
||||
func TestTrafficProductionDefaultDirectIsNotProtected(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Rules: []model.Rule{
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "direct"},
|
||||
},
|
||||
}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictDirect {
|
||||
t.Fatalf("Verdict = %q, want %q — a router whose only rule is `default -> direct` proxies nothing", got.Verdict, VerdictDirect)
|
||||
}
|
||||
if got.Default != tagDirect {
|
||||
t.Errorf("Default = %q, want %q", got.Default, tagDirect)
|
||||
}
|
||||
if got.TunnelRules != 0 {
|
||||
t.Errorf("TunnelRules = %d, want 0", got.TunnelRules)
|
||||
}
|
||||
}
|
||||
|
||||
// --- the three healthy-ish shapes -------------------------------------------
|
||||
|
||||
// A catch-all pointing at a group: everything unmatched is proxied.
|
||||
func TestTrafficDefaultToGroupIsTunnel(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{ssNode("ss1")},
|
||||
Groups: []model.Group{{Name: "auto", Nodes: []string{"ss1"}}},
|
||||
Rules: []model.Rule{
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "group:auto"},
|
||||
},
|
||||
}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictTunnel {
|
||||
t.Fatalf("Verdict = %q, want %q; default=%q", got.Verdict, VerdictTunnel, got.Default)
|
||||
}
|
||||
}
|
||||
|
||||
// A conditional rule into a group with a `direct` default: selective protection.
|
||||
// The default still leaves in the clear, and that is what must be said.
|
||||
func TestTrafficDirectDefaultWithProxyRuleIsSplit(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{ssNode("ss1")},
|
||||
Groups: []model.Group{{Name: "auto", Nodes: []string{"ss1"}}},
|
||||
Rules: []model.Rule{
|
||||
{Name: "lan-via-proxy", Enabled: true, Order: 10, Src: []string{"192.168.1.0/24"}, Target: "group:auto"},
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "direct"},
|
||||
},
|
||||
}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictSplit {
|
||||
t.Fatalf("Verdict = %q, want %q; default=%q rules=%d", got.Verdict, VerdictSplit, got.Default, got.TunnelRules)
|
||||
}
|
||||
if got.TunnelRules != 1 {
|
||||
t.Errorf("TunnelRules = %d, want 1", got.TunnelRules)
|
||||
}
|
||||
}
|
||||
|
||||
// No catch-all rule at all with the kill-switch closed: Final is the fail-closed
|
||||
// backstop, so unmatched traffic is DROPPED. Nothing leaks — reporting this as
|
||||
// "going out directly" would be a lie in the other direction.
|
||||
func TestTrafficNoDefaultRuleIsBlocked(t *testing.T) {
|
||||
m := &model.Model{Globals: model.DefaultGlobals()}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictBlocked {
|
||||
t.Fatalf("Verdict = %q, want %q; default=%q", got.Verdict, VerdictBlocked, got.Default)
|
||||
}
|
||||
if got.TunnelRules != 0 {
|
||||
t.Errorf("TunnelRules = %d, want 0", got.TunnelRules)
|
||||
}
|
||||
}
|
||||
|
||||
// Kill-switch open with no catch-all: Final is `direct`, so everything unmatched
|
||||
// leaves in the clear.
|
||||
func TestTrafficNoDefaultRuleOpenKillSwitchIsDirect(t *testing.T) {
|
||||
g := model.DefaultGlobals()
|
||||
g.KillSwitch = "open"
|
||||
got := trafficOf(t, &model.Model{Globals: g})
|
||||
if got.Verdict != VerdictDirect {
|
||||
t.Fatalf("Verdict = %q, want %q; default=%q", got.Verdict, VerdictDirect, got.Default)
|
||||
}
|
||||
}
|
||||
|
||||
// --- the shapes only the GENERATED config can reveal ------------------------
|
||||
|
||||
// Two condition-less rules: the LAST one owns route.Final (model.RuleReachability).
|
||||
// A `default -> group` at order 20 followed by a `default -> direct` at order 100
|
||||
// looks protected in the rule list and proxies nothing. Reading the model instead
|
||||
// of the generated config is how a verdict gets this wrong.
|
||||
func TestTrafficShadowedCatchAllFollowsTheWinner(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{ssNode("ss1")},
|
||||
Groups: []model.Group{{Name: "auto", Nodes: []string{"ss1"}}},
|
||||
Rules: []model.Rule{
|
||||
{Name: "default-proxy", Enabled: true, Order: 20, Target: "group:auto"},
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "direct"},
|
||||
},
|
||||
}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictDirect {
|
||||
t.Fatalf("Verdict = %q, want %q — the later condition-less rule owns Final", got.Verdict, VerdictDirect)
|
||||
}
|
||||
}
|
||||
|
||||
// An interface egress is a choice of UPLINK, never a tunnel: `default -> egress:wan2`
|
||||
// sends everything out a second WAN, in the clear.
|
||||
func TestTrafficEgressDefaultIsNotATunnel(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Egresses: []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}},
|
||||
Rules: []model.Rule{
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "egress:wan2"},
|
||||
},
|
||||
}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictDirect {
|
||||
t.Fatalf("Verdict = %q, want %q; default=%q", got.Verdict, VerdictDirect, got.Default)
|
||||
}
|
||||
}
|
||||
|
||||
// A byedpi egress is a SOCKS5 outbound to 127.0.0.1 — the local desync helper.
|
||||
// It reshapes the packets but the connection still leaves over the plain WAN with
|
||||
// this router's address, so it must not be reported as a tunnel.
|
||||
func TestTrafficByedpiEgressIsNotATunnel(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Egresses: []model.Egress{{Name: "bd", Type: "byedpi", Port: 1080}},
|
||||
Rules: []model.Rule{
|
||||
{Name: "default", Enabled: true, Order: 100, Target: "egress:bd"},
|
||||
},
|
||||
}
|
||||
got := trafficOf(t, m)
|
||||
if got.Verdict != VerdictDirect {
|
||||
t.Fatalf("Verdict = %q, want %q; default=%q", got.Verdict, VerdictDirect, got.Default)
|
||||
}
|
||||
}
|
||||
|
||||
// --- unit-level: the outbound index ------------------------------------------
|
||||
|
||||
func TestTrafficOfNilRouteIsUnknown(t *testing.T) {
|
||||
if got := (TrafficOf(option.Options{})); got.Verdict != "" {
|
||||
t.Fatalf("Verdict = %q, want \"\" (unknown) for options we did not build", got.Verdict)
|
||||
}
|
||||
}
|
||||
|
||||
// A group is a tunnel when any member is, and only then. A selector holding
|
||||
// nothing but `direct` routes nothing anywhere special.
|
||||
func TestTrafficGroupMembershipDecidesTunnel(t *testing.T) {
|
||||
base := []option.Outbound{
|
||||
{Type: C.TypeDirect, Tag: tagDirect, Options: &option.DirectOutboundOptions{}},
|
||||
{Type: C.TypeBlock, Tag: tagBlock, Options: &option.StubOptions{}},
|
||||
{Type: C.TypeShadowsocks, Tag: "ss1", Options: &option.ShadowsocksOutboundOptions{}},
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
members []string
|
||||
want string
|
||||
}{
|
||||
{"proxy member", []string{"ss1", tagDirect}, VerdictTunnel},
|
||||
{"direct only", []string{tagDirect}, VerdictDirect},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
opts := option.Options{
|
||||
Outbounds: append(append([]option.Outbound{}, base...), option.Outbound{
|
||||
Type: C.TypeSelector, Tag: "grp",
|
||||
Options: &option.SelectorOutboundOptions{Outbounds: tc.members},
|
||||
}),
|
||||
Route: &option.RouteOptions{Final: "grp"},
|
||||
}
|
||||
if got := TrafficOf(opts); got.Verdict != tc.want {
|
||||
t.Fatalf("Verdict = %q, want %q", got.Verdict, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A group that references itself must not hang the status poll.
|
||||
func TestTrafficGroupCycleTerminates(t *testing.T) {
|
||||
opts := option.Options{
|
||||
Outbounds: []option.Outbound{
|
||||
{Type: C.TypeSelector, Tag: "a", Options: &option.SelectorOutboundOptions{Outbounds: []string{"b"}}},
|
||||
{Type: C.TypeSelector, Tag: "b", Options: &option.SelectorOutboundOptions{Outbounds: []string{"a"}}},
|
||||
},
|
||||
Route: &option.RouteOptions{Final: "a"},
|
||||
}
|
||||
if got := TrafficOf(opts); got.Verdict != VerdictDirect {
|
||||
t.Fatalf("Verdict = %q, want %q for a cyclic group with no real member", got.Verdict, VerdictDirect)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,596 @@
|
||||
package generate
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
|
||||
"github.com/sagernet/sing-box/shater/netplane"
|
||||
"github.com/sagernet/sing-box/shater/parse"
|
||||
)
|
||||
|
||||
// WireGuard endpoint de-duplication — one private key, one device.
|
||||
//
|
||||
// # The defect this exists to prevent
|
||||
//
|
||||
// Everywhere else in this package a node may be COPIED freely: a per-chain hop
|
||||
// copy (chain.go buildHopWrapper/rebuildNode) and a per-group egress copy
|
||||
// (group.go) are rebuilt fresh from the share-link so each can carry its own
|
||||
// Detour without aliasing the shared base outbound. For a vless/ss/trojan node
|
||||
// that is exactly right — a copy is just another TCP client, and two of them cost
|
||||
// two connections.
|
||||
//
|
||||
// For a WireGuard/AmneziaWG node it is NOT. Each emitted endpoint is a real
|
||||
// device holding the node's private key, and a WireGuard PEER keeps exactly ONE
|
||||
// session per public key: the peer's endpoint address is rewritten by whichever
|
||||
// device most recently authenticated. Two devices built from the same key
|
||||
// therefore evict each other continuously — with persistent keepalive on both,
|
||||
// the eviction loop never settles, and NEITHER of them passes traffic. Observed
|
||||
// on the box as two UDP sockets from shaterd to the same peer port, the server's
|
||||
// peer endpoint flapping between them, zero bytes through the tunnel and every
|
||||
// hop of the chain failing with "context deadline exceeded".
|
||||
//
|
||||
// The multiplier is that buildOutboundsAndEndpoints (outbound.go) emits the BASE
|
||||
// endpoint for every enabled node unconditionally — whether or not anything
|
||||
// references it. So a WG node used only as a chain hop always produced two
|
||||
// devices: the unused base one and the chain copy. That is not an exotic
|
||||
// configuration, it is what any chain containing a WG node looks like, which made
|
||||
// EVERY such chain permanently dead.
|
||||
//
|
||||
// # Why this is one pass over the assembled option.Options
|
||||
//
|
||||
// There are three independent code paths that can materialise a WG device (the
|
||||
// base node, a chain hop copy, a chain GROUP hop's per-member copy) and nothing
|
||||
// stops a fourth from being added. Deduplicating at each producer would have to
|
||||
// be remembered at each producer. Doing it once, on the finished option.Options,
|
||||
// catches all of them by construction and needs no cooperation from the callers:
|
||||
// whatever put a second device in the config, it is gone before the config
|
||||
// reaches box.New.
|
||||
//
|
||||
// # What is left alone
|
||||
//
|
||||
// Non-wireguard outbounds are untouched (their copies are legitimate), groups
|
||||
// keep their semantics, and a node that appears exactly once — the overwhelmingly
|
||||
// common case, including "used only inside a chain" — produces no warning at all.
|
||||
|
||||
// wgDeviceKey returns the PHYSICAL DEVICE identity of an endpoint: two endpoints
|
||||
// with equal keys would be two devices fighting over one session at the peer.
|
||||
//
|
||||
// The identity is deliberately narrow — the private key plus the peer set (public
|
||||
// key + address:port), peers sorted so ordering cannot split a pair — because
|
||||
// collapsing too much is a routing change and collapsing too little leaves the
|
||||
// bug in place. Two nodes with DIFFERENT private keys are different devices at
|
||||
// the peer and must never be merged, however similar the rest of their config is;
|
||||
// two copies of the SAME key are the same device however different their MTU,
|
||||
// AllowedIPs or Detour happen to be.
|
||||
//
|
||||
// ok=false for anything that is not a wireguard endpoint, and for one with no
|
||||
// private key at all (it cannot be identified, and it would not come up anyway).
|
||||
func wgDeviceKey(ep option.Endpoint) (string, bool) {
|
||||
if ep.Type != C.TypeWireGuard {
|
||||
return "", false
|
||||
}
|
||||
wg, ok := ep.Options.(*option.WireGuardEndpointOptions)
|
||||
if !ok {
|
||||
return "", false
|
||||
}
|
||||
priv := strings.TrimSpace(wg.PrivateKey)
|
||||
if priv == "" {
|
||||
return "", false
|
||||
}
|
||||
peers := make([]string, 0, len(wg.Peers))
|
||||
for _, p := range wg.Peers {
|
||||
peers = append(peers, fmt.Sprintf("%s@%s:%d",
|
||||
strings.TrimSpace(p.PublicKey), strings.TrimSpace(p.Address), p.Port))
|
||||
}
|
||||
sort.Strings(peers)
|
||||
return priv + "|" + strings.Join(peers, ","), true
|
||||
}
|
||||
|
||||
// wgPrivateKeyOf recovers the private-key half of a wgDeviceKey.
|
||||
func wgPrivateKeyOf(deviceKey string) string {
|
||||
priv, _, _ := strings.Cut(deviceKey, "|")
|
||||
return priv
|
||||
}
|
||||
|
||||
// dedupWireGuardEndpoints enforces "one private key = at most one live device" on
|
||||
// the fully assembled config. Called from GenerateWithWarningsAt once opts is
|
||||
// complete (chain outbounds/endpoints folded in, route built), because only then
|
||||
// is every producer's output visible and the reachability walk meaningful.
|
||||
//
|
||||
// Per group of endpoints sharing a device identity:
|
||||
//
|
||||
// - endpoints nothing can route to are DELETED, silently. This is the normal
|
||||
// case (the always-emitted base endpoint of a node used only in a chain) and
|
||||
// warning about it would train the operator to ignore the warning list.
|
||||
// - if more than one REACHABLE copy remains, the config asks for something a
|
||||
// single key cannot express — e.g. the same WG node entered over two different
|
||||
// WANs in two chains, which is two devices by definition. The first copy in
|
||||
// tag order is kept (deterministic across runs), the rest are deleted and every
|
||||
// reference to them is rewritten to `block`, and each is reported as critical.
|
||||
// - exactly one left: nothing to say.
|
||||
//
|
||||
// Deleting rather than merely re-pointing is essential: box.New starts EVERY
|
||||
// endpoint in the config regardless of whether anything routes to it, so a
|
||||
// duplicate left in the list would still bring its device up and still fight for
|
||||
// the session. Re-pointing alone would have fixed nothing.
|
||||
func (b *builder) dedupWireGuardEndpoints(opts *option.Options) {
|
||||
if opts == nil || len(opts.Endpoints) < 2 {
|
||||
return
|
||||
}
|
||||
|
||||
tagsByKey := map[string][]string{}
|
||||
var keyOrder []string
|
||||
for i := range opts.Endpoints {
|
||||
key, ok := wgDeviceKey(opts.Endpoints[i])
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if _, seen := tagsByKey[key]; !seen {
|
||||
keyOrder = append(keyOrder, key)
|
||||
}
|
||||
tagsByKey[key] = append(tagsByKey[key], opts.Endpoints[i].Tag)
|
||||
}
|
||||
var dupKeys []string
|
||||
for _, key := range keyOrder {
|
||||
if len(tagsByKey[key]) > 1 {
|
||||
dupKeys = append(dupKeys, key)
|
||||
}
|
||||
}
|
||||
if len(dupKeys) == 0 {
|
||||
return // the common case: no key is materialised twice, nothing to do
|
||||
}
|
||||
|
||||
reachable := reachableOptionTags(opts, b.subscriptionDetourSeeds()...)
|
||||
var names map[string]string // private key -> node name, built only if we must name one
|
||||
drop := map[string]bool{}
|
||||
for _, key := range dupKeys {
|
||||
tags := append([]string(nil), tagsByKey[key]...)
|
||||
sort.Strings(tags) // deterministic survivor, independent of emission order
|
||||
var live []string
|
||||
for _, tag := range tags {
|
||||
if reachable[tag] {
|
||||
live = append(live, tag)
|
||||
continue
|
||||
}
|
||||
drop[tag] = true // dead weight: a device nothing routes to
|
||||
}
|
||||
if len(live) < 2 {
|
||||
continue
|
||||
}
|
||||
if names == nil {
|
||||
names = b.wgNodeNamesByKey()
|
||||
}
|
||||
name := names[wgPrivateKeyOf(key)]
|
||||
if name == "" {
|
||||
name = live[0] // unknown to the model (defensive): name it by its tag
|
||||
}
|
||||
keep := live[0]
|
||||
for _, tag := range live[1:] {
|
||||
drop[tag] = true
|
||||
b.warnf("node %q: this WireGuard node is materialised twice in the engine config — as %q and as %q — and traffic can reach both. A WireGuard peer keeps ONE session per public key, so two devices built from one private key evict each other continuously and NEITHER tunnel passes traffic. Only %q is kept; everything that routed through %q is fail-closed (blocked) instead of leaving over the plain WAN. Give the second path its own WireGuard node (its own key), or route both paths through the same one",
|
||||
name, keep, tag, keep, tag)
|
||||
}
|
||||
}
|
||||
if len(drop) == 0 {
|
||||
return
|
||||
}
|
||||
|
||||
// Rewrite first, delete second: the rewrite must still see the tags it is
|
||||
// replacing. Every dangling reference is pointed at `block`, never at `direct`
|
||||
// — a consumer whose tunnel just disappeared must stop, not fall out onto the
|
||||
// plain WAN with the router's real address.
|
||||
retargetDroppedTags(opts, drop, tagBlock)
|
||||
removeDroppedEndpoints(opts, drop)
|
||||
}
|
||||
|
||||
// wgNodeNamesByKey maps a WireGuard private key back to the model node name that
|
||||
// carries it, so a warning can name the NODE the operator configured rather than
|
||||
// the generated tag of a copy ("chain-ewan-wg-subs-h2" means nothing on the Nodes
|
||||
// page). Built by re-parsing the share-links, which keeps this file self-contained:
|
||||
// no producer has to remember to register its copies here.
|
||||
//
|
||||
// Built lazily (only when a critical duplicate is actually reported) because a
|
||||
// subscription can hold hundreds of nodes and this parses all of them.
|
||||
// First-wins on a key shared by two nodes: they are the same device anyway, so
|
||||
// either name identifies it for the operator.
|
||||
func (b *builder) wgNodeNamesByKey() map[string]string {
|
||||
names := map[string]string{}
|
||||
for i := range b.m.Nodes {
|
||||
n := b.m.Nodes[i]
|
||||
if !n.Enabled {
|
||||
continue
|
||||
}
|
||||
p, err := parse.ParseShareLink(n.URI)
|
||||
if err != nil || p.WG == nil {
|
||||
continue
|
||||
}
|
||||
key := strings.TrimSpace(p.WG.PrivateKey)
|
||||
if key == "" {
|
||||
continue
|
||||
}
|
||||
if _, seen := names[key]; !seen {
|
||||
names[key] = n.Name
|
||||
}
|
||||
}
|
||||
return names
|
||||
}
|
||||
|
||||
// --- reachability over option structures ------------------------------------
|
||||
|
||||
// optionNode is one outbound/endpoint reduced to what the reachability walk needs.
|
||||
type optionNode struct {
|
||||
group bool // selector/urltest: routes through its members
|
||||
members []string // group members (+ its default, if any)
|
||||
detour string // ordinary outbound/endpoint: its DialerOptions.Detour
|
||||
}
|
||||
|
||||
// reachableOptionTags is the option-layer twin of route.walkReachable
|
||||
// (route/reachability_lx.go): the set of outbound/endpoint tags traffic can
|
||||
// currently reach. That one walks live adapters inside a running box, which does
|
||||
// not exist yet at generate time, so the same question is answered over the
|
||||
// option structures.
|
||||
//
|
||||
// One deliberate difference: a group contributes ALL of its members, not just the
|
||||
// one it would select right now. The runtime walk can ask a selector what it has
|
||||
// chosen; here nothing has been chosen yet, and treating the unselected members as
|
||||
// unreachable would delete an endpoint the group is free to switch to a second
|
||||
// later. Over-approximating is the safe direction — the worst it costs is a
|
||||
// duplicate reported as critical instead of being removed silently.
|
||||
// extraSeeds are references that exist OUTSIDE option.Options — today the
|
||||
// subscription fetch detours (see subscriptionDetourSeeds). They are entry points
|
||||
// exactly like a route rule, so they are walked identically.
|
||||
func reachableOptionTags(opts *option.Options, extraSeeds ...string) map[string]bool {
|
||||
graph := optionGraph(opts)
|
||||
reachable := map[string]bool{}
|
||||
for _, seed := range optionSeedTags(opts) {
|
||||
walkOptionReachable(seed, graph, reachable)
|
||||
}
|
||||
for _, seed := range extraSeeds {
|
||||
walkOptionReachable(seed, graph, reachable)
|
||||
}
|
||||
return reachable
|
||||
}
|
||||
|
||||
// subscriptionDetourSeeds returns the outbound tags the SUBSCRIPTION fetcher dials
|
||||
// through — references that are just as real as a route rule's, but that live in
|
||||
// the model rather than in option.Options and so are invisible to the walk above.
|
||||
//
|
||||
// A subscription with fetch_via=proxy is pulled through the engine outbound named
|
||||
// by its fetch_detour: apply.UpdateSubscription (shater/apply/apply.go) hands that
|
||||
// string to engine.HTTPClient, which resolves it with engine.ViaToTag and looks the
|
||||
// tag up in the RUNNING box. So `fetch_detour=node:awg` is a direct, load-bearing
|
||||
// use of that node's BASE endpoint. Without this seed the base endpoint looked
|
||||
// unreferenced, was deleted as a duplicate of the node's chain copy, and updating
|
||||
// the subscription failed with "unknown outbound tag" — a working setup broken by
|
||||
// a pass that is supposed to fix one.
|
||||
//
|
||||
// Enabled is deliberately NOT consulted: UpdateSubscription looks a subscription up
|
||||
// by name and never checks it, so the panel can fetch a disabled one and the detour
|
||||
// must resolve when it does.
|
||||
//
|
||||
// Non-node forms come along for free rather than being filtered out: one mapping
|
||||
// (viaOutboundTag) covers them all, and seeding `egress-<x>` or a group tag costs
|
||||
// nothing — a tag that does not exist is ignored by the walk. Filtering would be
|
||||
// extra code that could only make the result less correct (a WG node that is a
|
||||
// member of a group used ONLY as a fetch detour would lose its endpoint).
|
||||
func (b *builder) subscriptionDetourSeeds() []string {
|
||||
var seeds []string
|
||||
for i := range b.m.Subscriptions {
|
||||
sub := b.m.Subscriptions[i]
|
||||
if !strings.EqualFold(strings.TrimSpace(sub.FetchVia), "proxy") {
|
||||
continue // a direct fetch dials no outbound at all
|
||||
}
|
||||
seeds = append(seeds, viaOutboundTag(sub.FetchDetour))
|
||||
}
|
||||
return seeds
|
||||
}
|
||||
|
||||
// viaOutboundTag mirrors engine.ViaToTag (shater/engine/httpclient.go) EXACTLY:
|
||||
// the shater `via` selector -> the box-internal outbound tag it names.
|
||||
//
|
||||
// ""/"direct" -> "direct"
|
||||
// "group:<X>" -> "<X>"
|
||||
// "node:<X>" -> "<X>"
|
||||
// "egress:<X>" -> "egress-<X>"
|
||||
// "chain:<X>" -> "<X>"
|
||||
// "<X>" -> "<X>"
|
||||
//
|
||||
// Duplicated rather than imported, following the same rule the engine side already
|
||||
// applies to the generator's tag formats (see engine/grouphealth_test.go copyTag):
|
||||
// shater/generate is the layer BELOW the runtime and must not pull the whole engine
|
||||
// in for one string function. TestViaOutboundTagMatchesEngine is the tripwire that
|
||||
// fails the moment the two drift apart.
|
||||
//
|
||||
// Note this is NOT b.resolveTarget: that one validates a target against the emitted
|
||||
// tag sets and MATERIALISES a chain as a side effect. This is a pure spelling of
|
||||
// what the runtime will look up, which is the only question the walk is asking.
|
||||
func viaOutboundTag(via string) string {
|
||||
via = strings.TrimSpace(via)
|
||||
if via == "" || strings.EqualFold(via, tagDirect) {
|
||||
return tagDirect
|
||||
}
|
||||
for _, prefix := range []string{"group:", "node:", "chain:"} {
|
||||
if rest, ok := cutPrefixFold(via, prefix); ok {
|
||||
return strings.TrimSpace(rest)
|
||||
}
|
||||
}
|
||||
if rest, ok := cutPrefixFold(via, "egress:"); ok {
|
||||
return netplane.EgressOutboundTag(strings.TrimSpace(rest))
|
||||
}
|
||||
return via
|
||||
}
|
||||
|
||||
// cutPrefixFold is strings.CutPrefix with a case-insensitive prefix match.
|
||||
func cutPrefixFold(s, prefix string) (string, bool) {
|
||||
if len(s) >= len(prefix) && strings.EqualFold(s[:len(prefix)], prefix) {
|
||||
return s[len(prefix):], true
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// optionGraph indexes every outbound/endpoint tag to its outgoing references.
|
||||
func optionGraph(opts *option.Options) map[string]optionNode {
|
||||
graph := make(map[string]optionNode, len(opts.Outbounds)+len(opts.Endpoints))
|
||||
for i := range opts.Outbounds {
|
||||
ob := opts.Outbounds[i]
|
||||
if members, def, ok := groupMembersOf(ob.Options); ok {
|
||||
all := append([]string(nil), members...)
|
||||
if def != "" {
|
||||
all = append(all, def)
|
||||
}
|
||||
graph[ob.Tag] = optionNode{group: true, members: all}
|
||||
continue
|
||||
}
|
||||
graph[ob.Tag] = optionNode{detour: dialerDetour(ob.Options)}
|
||||
}
|
||||
for i := range opts.Endpoints {
|
||||
graph[opts.Endpoints[i].Tag] = optionNode{detour: dialerDetour(opts.Endpoints[i].Options)}
|
||||
}
|
||||
return graph
|
||||
}
|
||||
|
||||
// optionSeedTags collects every tag traffic can ENTER the outbound graph at: the
|
||||
// route's Final (the kill-switch backstop), each route rule's routed/bypassed
|
||||
// outbound, each DNS server's detour, and the download detours of the remote
|
||||
// rule-sets / geo sources / clash external UI (those dial through an outbound too,
|
||||
// so a tag one of them names is in use even if no user traffic reaches it).
|
||||
func optionSeedTags(opts *option.Options) []string {
|
||||
var seeds []string
|
||||
if rt := opts.Route; rt != nil {
|
||||
seeds = append(seeds, rt.Final)
|
||||
for i := range rt.Rules {
|
||||
seeds = append(seeds, ruleActionOutbound(rt.Rules[i]))
|
||||
}
|
||||
for i := range rt.RuleSet {
|
||||
if rt.RuleSet[i].Type == C.RuleSetTypeRemote {
|
||||
seeds = append(seeds, rt.RuleSet[i].RemoteOptions.DownloadDetour)
|
||||
}
|
||||
}
|
||||
if rt.GeoIP != nil {
|
||||
seeds = append(seeds, rt.GeoIP.DownloadDetour)
|
||||
}
|
||||
if rt.Geosite != nil {
|
||||
seeds = append(seeds, rt.Geosite.DownloadDetour)
|
||||
}
|
||||
}
|
||||
if opts.DNS != nil {
|
||||
for i := range opts.DNS.Servers {
|
||||
seeds = append(seeds, dialerDetour(opts.DNS.Servers[i].Options))
|
||||
}
|
||||
}
|
||||
if opts.Experimental != nil && opts.Experimental.ClashAPI != nil {
|
||||
seeds = append(seeds, opts.Experimental.ClashAPI.ExternalUIDownloadDetour)
|
||||
}
|
||||
return seeds
|
||||
}
|
||||
|
||||
// walkOptionReachable marks tag reachable and descends into everything it routes
|
||||
// through, transitively. The visited set doubles as the cycle guard.
|
||||
func walkOptionReachable(tag string, graph map[string]optionNode, reachable map[string]bool) {
|
||||
if tag == "" || reachable[tag] {
|
||||
return
|
||||
}
|
||||
reachable[tag] = true
|
||||
node, ok := graph[tag]
|
||||
if !ok {
|
||||
return // dangling reference; box.New reports it, this pass does not care
|
||||
}
|
||||
if node.group {
|
||||
for _, member := range node.members {
|
||||
walkOptionReachable(member, graph, reachable)
|
||||
}
|
||||
return
|
||||
}
|
||||
walkOptionReachable(node.detour, graph, reachable)
|
||||
}
|
||||
|
||||
// ruleActionOutbound returns the outbound tag a route rule's action sends traffic
|
||||
// to, or "" for the actions that route nowhere (reject/sniff/dns/hijack-dns/…).
|
||||
// An empty Action is the route action (see option.RuleAction.UnmarshalJSON), so it
|
||||
// is read the same way.
|
||||
func ruleActionOutbound(rule option.Rule) string {
|
||||
action := rule.DefaultOptions.RuleAction
|
||||
if rule.Type == C.RuleTypeLogical {
|
||||
action = rule.LogicalOptions.RuleAction
|
||||
}
|
||||
switch action.Action {
|
||||
case C.RuleActionTypeRoute, "":
|
||||
return action.RouteOptions.Outbound
|
||||
case C.RuleActionTypeBypass:
|
||||
return action.BypassOptions.Outbound
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// --- rewriting + removal ----------------------------------------------------
|
||||
|
||||
// dialerDetour reads DialerOptions.Detour off any typed options that carry a
|
||||
// dialer (every *OutboundOptions, WireGuardEndpointOptions and DNS server
|
||||
// transport embeds option.DialerOptions). "" for options that carry none.
|
||||
func dialerDetour(o any) string {
|
||||
w, ok := o.(option.DialerOptionsWrapper)
|
||||
if !ok {
|
||||
return ""
|
||||
}
|
||||
return w.TakeDialerOptions().Detour
|
||||
}
|
||||
|
||||
// groupMembersOf reports a group outbound's member list and its default member.
|
||||
// ok=false for anything that is not a selector/urltest.
|
||||
func groupMembersOf(o any) (members []string, def string, ok bool) {
|
||||
switch g := o.(type) {
|
||||
case *option.SelectorOutboundOptions:
|
||||
return g.Outbounds, g.Default, true
|
||||
case *option.URLTestOutboundOptions:
|
||||
return g.Outbounds, "", true
|
||||
}
|
||||
return nil, "", false
|
||||
}
|
||||
|
||||
// retargetDroppedTags repoints every reference to a dropped endpoint tag, so the
|
||||
// config stays internally consistent once the endpoint is gone. A dangling tag is
|
||||
// not merely untidy: box.New resolves group members and route targets eagerly and
|
||||
// refuses the whole config over one of them, which with the kill-switch closed is
|
||||
// the entire LAN offline.
|
||||
//
|
||||
// Detours, route targets and the route Final go to `to` (block) — the fail-closed
|
||||
// direction. Group MEMBER lists instead drop the tag and only fall back to a lone
|
||||
// block member when that empties the group: a group that still has other tunnels
|
||||
// to balance over should use them, and a block sitting in a urltest pool would
|
||||
// otherwise be probed and reported dead forever on the group health card.
|
||||
func retargetDroppedTags(opts *option.Options, drop map[string]bool, to string) {
|
||||
for i := range opts.Outbounds {
|
||||
if !retargetGroupMembers(opts.Outbounds[i].Options, drop, to) {
|
||||
retargetDetour(opts.Outbounds[i].Options, drop, to)
|
||||
}
|
||||
}
|
||||
for i := range opts.Endpoints {
|
||||
retargetDetour(opts.Endpoints[i].Options, drop, to)
|
||||
}
|
||||
if opts.DNS != nil {
|
||||
for i := range opts.DNS.Servers {
|
||||
retargetDetour(opts.DNS.Servers[i].Options, drop, to)
|
||||
}
|
||||
}
|
||||
if rt := opts.Route; rt != nil {
|
||||
if drop[rt.Final] {
|
||||
rt.Final = to
|
||||
}
|
||||
for i := range rt.Rules {
|
||||
retargetRuleAction(&rt.Rules[i], drop, to)
|
||||
}
|
||||
for i := range rt.RuleSet {
|
||||
if rt.RuleSet[i].Type == C.RuleSetTypeRemote && drop[rt.RuleSet[i].RemoteOptions.DownloadDetour] {
|
||||
rt.RuleSet[i].RemoteOptions.DownloadDetour = to
|
||||
}
|
||||
}
|
||||
if rt.GeoIP != nil && drop[rt.GeoIP.DownloadDetour] {
|
||||
rt.GeoIP.DownloadDetour = to
|
||||
}
|
||||
if rt.Geosite != nil && drop[rt.Geosite.DownloadDetour] {
|
||||
rt.Geosite.DownloadDetour = to
|
||||
}
|
||||
}
|
||||
if opts.Experimental != nil && opts.Experimental.ClashAPI != nil &&
|
||||
drop[opts.Experimental.ClashAPI.ExternalUIDownloadDetour] {
|
||||
opts.Experimental.ClashAPI.ExternalUIDownloadDetour = to
|
||||
}
|
||||
}
|
||||
|
||||
// retargetRuleAction points a route rule whose target was dropped at `to`. It is
|
||||
// the write half of ruleActionOutbound and must stay in step with it: a rule the
|
||||
// seed walk counted as a reference is a rule this has to be able to repoint.
|
||||
func retargetRuleAction(rule *option.Rule, drop map[string]bool, to string) {
|
||||
action := &rule.DefaultOptions.RuleAction
|
||||
if rule.Type == C.RuleTypeLogical {
|
||||
action = &rule.LogicalOptions.RuleAction
|
||||
}
|
||||
switch action.Action {
|
||||
case C.RuleActionTypeRoute, "":
|
||||
if drop[action.RouteOptions.Outbound] {
|
||||
action.RouteOptions.Outbound = to
|
||||
}
|
||||
case C.RuleActionTypeBypass:
|
||||
if drop[action.BypassOptions.Outbound] {
|
||||
action.BypassOptions.Outbound = to
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// retargetDetour points a dropped Detour at `to`, leaving any other value alone.
|
||||
func retargetDetour(o any, drop map[string]bool, to string) {
|
||||
w, ok := o.(option.DialerOptionsWrapper)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
d := w.TakeDialerOptions()
|
||||
if !drop[d.Detour] {
|
||||
return
|
||||
}
|
||||
d.Detour = to
|
||||
w.ReplaceDialerOptions(d)
|
||||
}
|
||||
|
||||
// retargetGroupMembers prunes dropped members from a group, reporting whether the
|
||||
// options were a group at all (so the caller knows not to look for a detour).
|
||||
// A selector whose Default was dropped falls back to an empty Default, which the
|
||||
// engine reads as "the first member" — always a member that still exists.
|
||||
func retargetGroupMembers(o any, drop map[string]bool, to string) bool {
|
||||
switch g := o.(type) {
|
||||
case *option.SelectorOutboundOptions:
|
||||
g.Outbounds = pruneMembers(g.Outbounds, drop, to)
|
||||
if drop[g.Default] {
|
||||
g.Default = ""
|
||||
}
|
||||
return true
|
||||
case *option.URLTestOutboundOptions:
|
||||
g.Outbounds = pruneMembers(g.Outbounds, drop, to)
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// pruneMembers removes dropped tags from a member list, substituting a single
|
||||
// fallback member when that would leave the group empty (an empty selector/urltest
|
||||
// is refused by box.New, and the fallback is `block` so the emptied group is
|
||||
// fail-closed rather than a hole).
|
||||
func pruneMembers(in []string, drop map[string]bool, fallback string) []string {
|
||||
hit := false
|
||||
for _, tag := range in {
|
||||
if drop[tag] {
|
||||
hit = true
|
||||
break
|
||||
}
|
||||
}
|
||||
if !hit {
|
||||
return in
|
||||
}
|
||||
out := make([]string, 0, len(in))
|
||||
for _, tag := range in {
|
||||
if drop[tag] {
|
||||
continue
|
||||
}
|
||||
out = append(out, tag)
|
||||
}
|
||||
if len(out) == 0 {
|
||||
return []string{fallback}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// removeDroppedEndpoints filters the dropped endpoints out of the config. This is
|
||||
// the step that actually stops the duplicate device from being created.
|
||||
func removeDroppedEndpoints(opts *option.Options, drop map[string]bool) {
|
||||
kept := opts.Endpoints[:0]
|
||||
for _, ep := range opts.Endpoints {
|
||||
if drop[ep.Tag] {
|
||||
continue
|
||||
}
|
||||
kept = append(kept, ep)
|
||||
}
|
||||
opts.Endpoints = kept
|
||||
}
|
||||
@@ -0,0 +1,416 @@
|
||||
package generate
|
||||
|
||||
import (
|
||||
"encoding/base64"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"sort"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
C "github.com/sagernet/sing-box/constant"
|
||||
"github.com/sagernet/sing-box/option"
|
||||
|
||||
"github.com/sagernet/sing-box/shater/engine"
|
||||
"github.com/sagernet/sing-box/shater/model"
|
||||
)
|
||||
|
||||
// These tests pin the one invariant a WireGuard node cannot survive without: the
|
||||
// generated config must never contain two devices built from one private key. See
|
||||
// wgdedup.go for why (a peer keeps a single session per key, so two devices evict
|
||||
// each other and neither passes traffic).
|
||||
//
|
||||
// They are deliberately NOT Linux-gated: nothing here goes through box.New, so the
|
||||
// `routing_mark` platform restriction that gates generate_test.go does not apply.
|
||||
|
||||
// wgDedupKey returns a valid 32-byte base64 WireGuard key seeded by fill. It is a
|
||||
// separate helper from generate_test.go's validKey on purpose: that file is
|
||||
// //go:build linux, and this suite must run on every dev platform.
|
||||
func wgDedupKey(fill byte) string {
|
||||
b := make([]byte, 32)
|
||||
for i := range b {
|
||||
b[i] = fill + byte(i)
|
||||
}
|
||||
return base64.StdEncoding.EncodeToString(b)
|
||||
}
|
||||
|
||||
// wgDedupURI builds a wireguard:// share-link for a node.
|
||||
func wgDedupURI(priv, pub, server string, port int) string {
|
||||
return fmt.Sprintf("wireguard://%s@%s:%d?publickey=%s&address=10.13.13.2/32&allowedips=0.0.0.0/0",
|
||||
url.QueryEscape(priv), server, port, url.QueryEscape(pub))
|
||||
}
|
||||
|
||||
// endpointTags lists the emitted endpoint tags, sorted for stable comparison.
|
||||
func endpointTags(opts option.Options) []string {
|
||||
out := make([]string, 0, len(opts.Endpoints))
|
||||
for i := range opts.Endpoints {
|
||||
out = append(out, opts.Endpoints[i].Tag)
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out
|
||||
}
|
||||
|
||||
// routeRuleOutbounds lists the outbound tag of every route rule that routes.
|
||||
func routeRuleOutbounds(rt *option.RouteOptions) []string {
|
||||
if rt == nil {
|
||||
return nil
|
||||
}
|
||||
var out []string
|
||||
for i := range rt.Rules {
|
||||
if tag := ruleActionOutbound(rt.Rules[i]); tag != "" {
|
||||
out = append(out, tag)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func warningsContaining(warns []string, needle string) []string {
|
||||
var out []string
|
||||
for _, w := range warns {
|
||||
if strings.Contains(w, needle) {
|
||||
out = append(out, w)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// (а) A WG node used ONLY as a chain hop must leave exactly ONE wireguard
|
||||
// endpoint behind, and it must be the chain copy: the base endpoint that
|
||||
// buildOutboundsAndEndpoints emits for every enabled node is the second device
|
||||
// that killed the tunnel, and nothing routes to it. Silent — this is the normal
|
||||
// shape of any chain containing a WG node, not an operator error.
|
||||
//
|
||||
// (в) rides along here: the non-WG hop (ss1) keeps BOTH its base outbound and its
|
||||
// chain copy, because two TCP clients are not a conflict.
|
||||
func TestWGDedupChainOnlyNodeDropsBaseEndpoint(t *testing.T) {
|
||||
priv, pub := wgDedupKey(1), wgDedupKey(9)
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{
|
||||
{Name: "wg1", Enabled: true, URI: wgDedupURI(priv, pub, "203.0.113.10", 51820)},
|
||||
{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.2:8388#ss1"},
|
||||
},
|
||||
Chains: []model.Chain{
|
||||
{Name: "c", Hops: []string{"node:wg1", "node:ss1"}},
|
||||
},
|
||||
Rules: []model.Rule{
|
||||
{Name: "via-chain", Enabled: true, Order: 10, DstPort: "443", Target: "chain:c"},
|
||||
},
|
||||
}
|
||||
|
||||
opts, warns, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "chain-c-h1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [chain-c-h1] (base wg1 removed)", got)
|
||||
}
|
||||
if opts.Endpoints[0].Type != C.TypeWireGuard {
|
||||
t.Fatalf("surviving endpoint type = %q, want %q", opts.Endpoints[0].Type, C.TypeWireGuard)
|
||||
}
|
||||
// The surviving copy must still be the chain's L1: dialled directly from the
|
||||
// router (no detour), with h2 detouring into it.
|
||||
if d := anyDetour(t, opts, "chain-c-h1"); d != "" {
|
||||
t.Fatalf("chain-c-h1 detour = %q, want \"\" (L1 dials directly)", d)
|
||||
}
|
||||
if d := anyDetour(t, opts, "chain-c-h2"); d != "chain-c-h1" {
|
||||
t.Fatalf("chain-c-h2 detour = %q, want chain-c-h1", d)
|
||||
}
|
||||
// (в) the non-WG node keeps its base outbound AND its chain copy.
|
||||
if obByTag(opts, "ss1") == nil {
|
||||
t.Fatalf("base outbound of the non-wireguard node was removed; outbounds=%v", outboundTags(opts))
|
||||
}
|
||||
if obByTag(opts, "chain-c-h2") == nil {
|
||||
t.Fatalf("chain copy of the non-wireguard node missing; outbounds=%v", outboundTags(opts))
|
||||
}
|
||||
// Silent: nothing was misconfigured.
|
||||
if got := warningsContaining(warns, "materialised twice"); len(got) != 0 {
|
||||
t.Fatalf("the ordinary chain case must not warn, got %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// (б) The same WG node entered over two DIFFERENT egresses in two chains asks for
|
||||
// two devices from one key, which the peer cannot give. One copy survives (first
|
||||
// in tag order, deterministic); the loser is deleted and its consumer is pointed
|
||||
// at `block` — never at direct, which would put that traffic on the plain WAN with
|
||||
// the router's real address. The operator is told, critically.
|
||||
func TestWGDedupTwoChainsDifferentEgressFailClosed(t *testing.T) {
|
||||
priv, pub := wgDedupKey(2), wgDedupKey(7)
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{
|
||||
{Name: "wg1", Enabled: true, URI: wgDedupURI(priv, pub, "203.0.113.10", 51820)},
|
||||
},
|
||||
Egresses: []model.Egress{
|
||||
{Name: "w1", Type: "interface", Interface: "eth0"},
|
||||
{Name: "w2", Type: "interface", Interface: "eth1"},
|
||||
},
|
||||
Chains: []model.Chain{
|
||||
{Name: "a", Hops: []string{"egress:w1", "node:wg1"}},
|
||||
{Name: "b", Hops: []string{"egress:w2", "node:wg1"}},
|
||||
},
|
||||
Rules: []model.Rule{
|
||||
{Name: "over-w1", Enabled: true, Order: 10, DstPort: "443", Target: "chain:a"},
|
||||
{Name: "over-w2", Enabled: true, Order: 20, DstPort: "8443", Target: "chain:b"},
|
||||
},
|
||||
}
|
||||
|
||||
opts, warns, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "chain-a-h1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [chain-a-h1] (one device per key)", got)
|
||||
}
|
||||
// The survivor keeps its own entry egress; the loser's rule is fail-closed.
|
||||
if d := anyDetour(t, opts, "chain-a-h1"); d != "egress-w1" {
|
||||
t.Fatalf("chain-a-h1 detour = %q, want egress-w1", d)
|
||||
}
|
||||
targets := routeRuleOutbounds(opts.Route)
|
||||
var sawKept, sawBlocked bool
|
||||
for _, tag := range targets {
|
||||
switch tag {
|
||||
case "chain-a-h1":
|
||||
sawKept = true
|
||||
case tagBlock:
|
||||
sawBlocked = true
|
||||
case "chain-b-h1":
|
||||
t.Fatalf("a rule still routes to the deleted copy chain-b-h1: %v", targets)
|
||||
}
|
||||
}
|
||||
if !sawKept || !sawBlocked {
|
||||
t.Fatalf("route targets = %v, want the survivor kept and the loser fail-closed to %q", targets, tagBlock)
|
||||
}
|
||||
// Critical, and it must name the node plus BOTH copies so the operator can act.
|
||||
found := warningsContaining(warns, "materialised twice")
|
||||
if len(found) != 1 {
|
||||
t.Fatalf("want exactly one duplicate warning, got %v (all: %v)", found, warns)
|
||||
}
|
||||
w := found[0]
|
||||
for _, want := range []string{`node "wg1"`, "chain-a-h1", "chain-b-h1", "fail-closed"} {
|
||||
if !strings.Contains(w, want) {
|
||||
t.Fatalf("warning must contain %q, got: %s", want, w)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// (г) Two DIFFERENT WireGuard nodes are two different devices at two different
|
||||
// peers and must never be folded together, however alike the rest of their config
|
||||
// is. Both chain copies survive; both base endpoints go.
|
||||
func TestWGDedupDistinctNodesNotCollapsed(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{
|
||||
{Name: "wgA", Enabled: true, URI: wgDedupURI(wgDedupKey(3), wgDedupKey(11), "203.0.113.10", 51820)},
|
||||
{Name: "wgB", Enabled: true, URI: wgDedupURI(wgDedupKey(4), wgDedupKey(12), "203.0.113.11", 51820)},
|
||||
},
|
||||
Chains: []model.Chain{
|
||||
{Name: "c", Hops: []string{"node:wgA", "node:wgB"}},
|
||||
},
|
||||
Rules: []model.Rule{
|
||||
{Name: "via-chain", Enabled: true, Order: 10, DstPort: "443", Target: "chain:c"},
|
||||
},
|
||||
}
|
||||
|
||||
opts, warns, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
got := endpointTags(opts)
|
||||
if len(got) != 2 || got[0] != "chain-c-h1" || got[1] != "chain-c-h2" {
|
||||
t.Fatalf("endpoints = %v, want [chain-c-h1 chain-c-h2] (distinct keys kept apart)", got)
|
||||
}
|
||||
keys := map[string]bool{}
|
||||
for i := range opts.Endpoints {
|
||||
k, ok := wgDeviceKey(opts.Endpoints[i])
|
||||
if !ok {
|
||||
t.Fatalf("endpoint %q is not a keyed wireguard endpoint", opts.Endpoints[i].Tag)
|
||||
}
|
||||
keys[k] = true
|
||||
}
|
||||
if len(keys) != 2 {
|
||||
t.Fatalf("expected two distinct device identities, got %d", len(keys))
|
||||
}
|
||||
if got := warningsContaining(warns, "materialised twice"); len(got) != 0 {
|
||||
t.Fatalf("distinct nodes must not warn, got %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// (д) A WG node a rule targets DIRECTLY, with no chain copy anywhere, has exactly
|
||||
// one device already. The pass must leave it completely alone — it only ever
|
||||
// removes a duplicate, never the last copy of a node.
|
||||
func TestWGDedupDirectlyTargetedBaseEndpointKept(t *testing.T) {
|
||||
priv, pub := wgDedupKey(5), wgDedupKey(13)
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{
|
||||
{Name: "wg1", Enabled: true, URI: wgDedupURI(priv, pub, "203.0.113.10", 51820)},
|
||||
},
|
||||
Rules: []model.Rule{
|
||||
{Name: "direct-to-node", Enabled: true, Order: 10, DstPort: "443", Target: "node:wg1"},
|
||||
},
|
||||
}
|
||||
|
||||
opts, warns, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "wg1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [wg1] (base endpoint untouched)", got)
|
||||
}
|
||||
if got := routeRuleOutbounds(opts.Route); len(got) != 1 || got[0] != "wg1" {
|
||||
t.Fatalf("route targets = %v, want [wg1] (rule not rewritten)", got)
|
||||
}
|
||||
if got := warningsContaining(warns, "materialised twice"); len(got) != 0 {
|
||||
t.Fatalf("a singly-used node must not warn, got %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// --- subscription fetch_detour is a real reference ---------------------------
|
||||
|
||||
// fetchDetourModel is the shared fixture for the three fetch_detour cases: one WG
|
||||
// node, optionally used as a chain hop, optionally named as a subscription's fetch
|
||||
// detour. Both switches independent, so the three combinations are exactly the
|
||||
// three outcomes the pass must produce.
|
||||
func fetchDetourModel(inChain bool, fetchDetour string) *model.Model {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{
|
||||
{Name: "wg1", Enabled: true, URI: wgDedupURI(wgDedupKey(21), wgDedupKey(31), "203.0.113.10", 51820)},
|
||||
{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.2:8388#ss1"},
|
||||
},
|
||||
}
|
||||
if inChain {
|
||||
m.Chains = []model.Chain{{Name: "c", Hops: []string{"node:wg1", "node:ss1"}}}
|
||||
m.Rules = []model.Rule{{Name: "via-chain", Enabled: true, Order: 10, DstPort: "443", Target: "chain:c"}}
|
||||
} else {
|
||||
m.Rules = []model.Rule{{Name: "via-ss", Enabled: true, Order: 10, DstPort: "443", Target: "node:ss1"}}
|
||||
}
|
||||
if fetchDetour != "" {
|
||||
m.Subscriptions = []model.Subscription{{
|
||||
Name: "sub0", Enabled: true, URL: "https://example.net/sub",
|
||||
FetchVia: "proxy", FetchDetour: fetchDetour,
|
||||
}}
|
||||
}
|
||||
return m
|
||||
}
|
||||
|
||||
// Case 1 — node only in the chain, no fetch detour on it: the base endpoint is
|
||||
// genuinely unreferenced and goes, the chain works. (The pre-existing behaviour;
|
||||
// pinned here so the new seed cannot accidentally resurrect the base endpoint.)
|
||||
func TestWGDedupFetchDetourAbsentBaseStillDropped(t *testing.T) {
|
||||
opts, warns, err := GenerateWithWarnings(fetchDetourModel(true, ""))
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "chain-c-h1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [chain-c-h1]", got)
|
||||
}
|
||||
if got := warningsContaining(warns, "materialised twice"); len(got) != 0 {
|
||||
t.Fatalf("must be silent, got %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Case 2 — node in the chain AND named as a fetch detour: the config asks for the
|
||||
// node on two different dial paths, which one private key cannot provide. Both are
|
||||
// reachable, so one survives and the other is fail-closed with the critical
|
||||
// warning. Choosing silently for the operator is what we must NOT do.
|
||||
func TestWGDedupFetchDetourAndChainBothReachable(t *testing.T) {
|
||||
opts, warns, err := GenerateWithWarnings(fetchDetourModel(true, "node:wg1"))
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
// Tag order decides: "chain-c-h1" < "wg1", so the chain copy is the survivor.
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "chain-c-h1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [chain-c-h1] (one device per key)", got)
|
||||
}
|
||||
found := warningsContaining(warns, "materialised twice")
|
||||
if len(found) != 1 {
|
||||
t.Fatalf("want exactly one duplicate warning, got %v (all: %v)", found, warns)
|
||||
}
|
||||
for _, want := range []string{`node "wg1"`, "chain-c-h1", `"wg1"`, "fail-closed"} {
|
||||
if !strings.Contains(found[0], want) {
|
||||
t.Fatalf("warning must contain %q, got: %s", want, found[0])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Case 3 — node ONLY named as a fetch detour, no chain copy: a single device that
|
||||
// something really does dial. It must survive untouched, or updating the
|
||||
// subscription fails with "unknown outbound tag" — the regression this seed exists
|
||||
// to prevent.
|
||||
func TestWGDedupFetchDetourOnlyKeepsBaseEndpoint(t *testing.T) {
|
||||
for _, via := range []string{"node:wg1", "wg1", "NODE:wg1", " node:wg1 "} {
|
||||
t.Run(via, func(t *testing.T) {
|
||||
opts, warns, err := GenerateWithWarnings(fetchDetourModel(false, via))
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "wg1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [wg1] (the fetch detour's target)", got)
|
||||
}
|
||||
if got := warningsContaining(warns, "materialised twice"); len(got) != 0 {
|
||||
t.Fatalf("a single device must not warn, got %v", got)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A fetch_via=direct subscription dials no outbound at all, so its (ignored)
|
||||
// fetch_detour must NOT keep a duplicate alive.
|
||||
func TestWGDedupFetchDetourIgnoredWhenFetchViaDirect(t *testing.T) {
|
||||
m := fetchDetourModel(true, "node:wg1")
|
||||
m.Subscriptions[0].FetchVia = "direct"
|
||||
opts, warns, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 1 || got[0] != "chain-c-h1" {
|
||||
t.Fatalf("endpoints = %v, want exactly [chain-c-h1] (direct fetch is not a reference)", got)
|
||||
}
|
||||
if got := warningsContaining(warns, "materialised twice"); len(got) != 0 {
|
||||
t.Fatalf("must be silent, got %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestViaOutboundTagMatchesEngine is the drift tripwire for the duplicated `via`
|
||||
// mapping. viaOutboundTag must agree with engine.ViaToTag on every form, because
|
||||
// the seed is only correct if it names the tag the RUNTIME will look up: if the
|
||||
// engine's mapping changes and this one does not, the pass starts deleting an
|
||||
// endpoint the subscription fetcher still dials.
|
||||
func TestViaOutboundTagMatchesEngine(t *testing.T) {
|
||||
for _, via := range []string{
|
||||
"", " ", "direct", "DIRECT",
|
||||
"node:wg1", "NODE:wg1", " node: wg1 ", "node:",
|
||||
"group:auto", "Group:auto",
|
||||
"egress:wan2", "EGRESS:wan2",
|
||||
"chain:ewan", "Chain:ewan",
|
||||
"wg1", " wg1 ", "weird:value",
|
||||
} {
|
||||
if got, want := viaOutboundTag(via), engine.ViaToTag(via); got != want {
|
||||
t.Errorf("viaOutboundTag(%q) = %q, engine.ViaToTag = %q — the mappings have drifted", via, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// An UNREFERENCED WG node is still a single device and stays: this pass is about
|
||||
// duplicates, not about pruning unused config. (Guards the >1 precondition — an
|
||||
// over-eager version would delete it and silently change what a later `selector`
|
||||
// or a hand-written detour could reach.)
|
||||
func TestWGDedupUnreferencedSingleNodeKept(t *testing.T) {
|
||||
m := &model.Model{
|
||||
Globals: model.DefaultGlobals(),
|
||||
Nodes: []model.Node{
|
||||
{Name: "wg1", Enabled: true, URI: wgDedupURI(wgDedupKey(6), wgDedupKey(14), "203.0.113.10", 51820)},
|
||||
{Name: "wg2", Enabled: true, URI: wgDedupURI(wgDedupKey(8), wgDedupKey(15), "203.0.113.11", 51820)},
|
||||
},
|
||||
}
|
||||
|
||||
opts, _, err := GenerateWithWarnings(m)
|
||||
if err != nil {
|
||||
t.Fatalf("Generate: %v", err)
|
||||
}
|
||||
if got := endpointTags(opts); len(got) != 2 {
|
||||
t.Fatalf("endpoints = %v, want both nodes kept", got)
|
||||
}
|
||||
}
|
||||
+323
-9
@@ -16,6 +16,15 @@
|
||||
// answer "give me the last day" — the prefix is what makes the download
|
||||
// ranges real. UTC only: the binary ships without tzdata, so any local-zone
|
||||
// rendering would be a fiction.
|
||||
// - collapses REPEATS of the same message into "repeated N times: <message>"
|
||||
// (see repeatKey / repeatWindow / maxRepeatKeys): a broken outbound makes
|
||||
// the engine repeat one line hundreds of times a minute, which evicts the
|
||||
// whole rest of the router's syslog ring buffer within minutes. The repeats
|
||||
// do NOT have to be adjacent — a flood usually interleaves a handful of
|
||||
// messages (one per broken chain), so the sink keeps a small table of the
|
||||
// messages seen recently instead of comparing with the previous line only.
|
||||
// Suppression applies to both halves identically; fatal/panic lines are
|
||||
// never suppressed.
|
||||
// - fans it out according to Config: to the REAL os.Stderr (procd relays fd2
|
||||
// to syslog/logread) when ToSyslog, and to a size-capped, 2-segment rotated
|
||||
// file when ToFile. Both off => the line is dropped — that IS the "fully
|
||||
@@ -42,6 +51,7 @@ package logsink
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"container/list"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
@@ -74,6 +84,45 @@ const (
|
||||
// the log's own cap: the guard requires cap+floor available, so even a log
|
||||
// that grows to its full cap leaves the rootfs this much headroom.
|
||||
diskFloorBytes = 4 << 20
|
||||
|
||||
// repeatWindow bounds how long copies of one message may stay silent: the
|
||||
// first occurrence is printed and opens a window; every further copy of THAT
|
||||
// message inside the window is swallowed and reported by a
|
||||
// "repeated N times: <message>" summary when the window ends. A message that
|
||||
// is still flooding keeps its window (one summary per window, counter
|
||||
// restarting); a message that went quiet is forgotten, so it prints in full
|
||||
// the next time it happens.
|
||||
//
|
||||
// Why 5s: the observed flood (dead chains -> "WireGuard is not ready yet")
|
||||
// runs at ~7 lines/s across three chain tags — 436 lines in one minute —
|
||||
// while the router's syslog ring holds ~760 lines total: one faulty
|
||||
// subscription erases every other subsystem's history, including our own
|
||||
// startup lines, in under two minutes. A 5s window turns that into one
|
||||
// summary per message per window (~36 lines/min instead of 436) — short
|
||||
// enough that an operator tailing `logread -f` sees the fault acknowledged
|
||||
// within 5 seconds and keeps seeing it, so a standing problem never looks
|
||||
// like a frozen log. A longer window (30-60s) would hide the ongoing-ness; a
|
||||
// shorter one would not relieve the ring.
|
||||
repeatWindow = 5 * time.Second
|
||||
|
||||
// maxRepeatKeys bounds the table of recently seen messages. Suppression must
|
||||
// work across INTERLEAVED messages, so the sink cannot keep just the last
|
||||
// key — but the keys come from the engine and are as varied as its log, so
|
||||
// the table must not be allowed to grow with them.
|
||||
//
|
||||
// Why 256: the floods worth collapsing are per-outbound or per-chain, and a
|
||||
// pathological config on this box has a few hundred nodes of which only the
|
||||
// broken handful actually log; 256 distinct messages in flight covers that
|
||||
// with room to spare, while a table this size is trivial next to the
|
||||
// process — entries hold the log line itself, ~150 B typically (~40 KiB
|
||||
// total) and 8 KiB at the absolute worst (maxPartialLine), i.e. ~2 MiB even
|
||||
// in the case that cannot really happen.
|
||||
//
|
||||
// Eviction is least-recently-seen: the entry that has gone longest without a
|
||||
// copy is the one least likely to be flooding. Eviction is never silent —
|
||||
// an evicted entry that had swallowed copies prints its summary on the way
|
||||
// out (marked "repeat table full"), so a counter is never simply dropped.
|
||||
maxRepeatKeys = 256
|
||||
)
|
||||
|
||||
// FilePath returns the log-file location for the given persistence choice.
|
||||
@@ -153,9 +202,19 @@ type Sink struct {
|
||||
suspended bool // persistent-path disk guard tripped
|
||||
lastProbe time.Time // last disk-free probe
|
||||
|
||||
// repeat suppression (emitLocked and the repeatSeries helpers below).
|
||||
// series is the table of messages seen inside their window, keyed by
|
||||
// repeatKey; seriesLRU holds the same *repeatSeries values in
|
||||
// most-recently-seen-first order so eviction is O(1).
|
||||
series map[string]*repeatSeries
|
||||
seriesLRU *list.List
|
||||
repeatTimer *time.Timer // armed at the earliest series deadline, if any
|
||||
closed bool // Close ran: the timer must not write any more
|
||||
|
||||
// test seams
|
||||
now func() time.Time
|
||||
free func(string) (uint64, bool)
|
||||
now func() time.Time
|
||||
free func(string) (uint64, bool)
|
||||
window time.Duration // repeatWindow, overridable in tests
|
||||
}
|
||||
|
||||
// New returns a Sink fanning to stderr (the daemon passes os.Stderr) under cfg.
|
||||
@@ -168,10 +227,13 @@ func New(stderr io.Writer, cfg Config) *Sink {
|
||||
purgeLogFiles(cfg.path(), PersistPath, TmpfsPath)
|
||||
}
|
||||
return &Sink{
|
||||
cfg: cfg,
|
||||
stderr: stderr,
|
||||
now: time.Now,
|
||||
free: freeBytes,
|
||||
cfg: cfg,
|
||||
stderr: stderr,
|
||||
series: make(map[string]*repeatSeries),
|
||||
seriesLRU: list.New(),
|
||||
now: time.Now,
|
||||
free: freeBytes,
|
||||
window: repeatWindow,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -210,6 +272,9 @@ func (s *Sink) Reconfigure(cfg Config) {
|
||||
if cfg == s.cfg {
|
||||
return
|
||||
}
|
||||
// Settle every series in progress under the OLD configuration: their
|
||||
// summaries belong to the destination the swallowed lines were headed for.
|
||||
s.flushAllSeriesLocked()
|
||||
if cfg.path() != s.cfg.path() || !cfg.ToFile {
|
||||
s.closeFileLocked()
|
||||
}
|
||||
@@ -222,8 +287,10 @@ func (s *Sink) Reconfigure(cfg Config) {
|
||||
s.cfg = cfg
|
||||
}
|
||||
|
||||
// Close flushes a pending partial line and closes the file segment. The sink
|
||||
// must not be written to afterwards.
|
||||
// Close flushes a pending partial line, emits the summaries of every still-open
|
||||
// series of repeats (no series may be lost at shutdown) and closes the file
|
||||
// segment. The sink must not be written to afterwards; a repeat timer that
|
||||
// fires after Close is a no-op.
|
||||
func (s *Sink) Close() error {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
@@ -231,6 +298,8 @@ func (s *Sink) Close() error {
|
||||
s.emitLocked(s.buf)
|
||||
s.buf = nil
|
||||
}
|
||||
s.flushAllSeriesLocked()
|
||||
s.closed = true
|
||||
if s.file != nil {
|
||||
err := s.file.Close()
|
||||
s.file = nil
|
||||
@@ -242,12 +311,257 @@ func (s *Sink) Close() error {
|
||||
|
||||
// --- internals (caller holds s.mu) -------------------------------------------
|
||||
|
||||
// emitLocked stamps one complete line with the UTC wall clock and fans it out.
|
||||
// emitLocked runs one complete line through repeat suppression and, unless it
|
||||
// is swallowed as a copy of a message already printed inside its window, hands
|
||||
// it to writeLineLocked.
|
||||
func (s *Sink) emitLocked(line []byte) {
|
||||
if !s.cfg.ToSyslog && !s.cfg.ToFile {
|
||||
return // fully off: the line is dropped, nowhere else to go
|
||||
}
|
||||
line = bytes.TrimSuffix(line, []byte{'\r'})
|
||||
key, exempt := repeatKey(line)
|
||||
if exempt {
|
||||
// fatal/panic: always printed, and never opens a series — a dying
|
||||
// daemon must not have its last words counted instead of said. The
|
||||
// pending summaries go out FIRST: the process may not live long enough
|
||||
// to reach Close, and a swallowed count that is never reported is
|
||||
// exactly the failure this whole mechanism exists to avoid.
|
||||
s.flushAllSeriesLocked()
|
||||
s.writeLineLocked(line)
|
||||
return
|
||||
}
|
||||
now := s.now()
|
||||
if ser, ok := s.series[key]; ok {
|
||||
if now.Before(ser.deadline) {
|
||||
ser.count++
|
||||
s.seriesLRU.MoveToFront(ser.el)
|
||||
return // swallowed; the deadline did not move, so no re-arming
|
||||
}
|
||||
// The window elapsed without the timer having got to it (a coarse
|
||||
// timer, a frozen test clock, a burst racing the callback): settle it
|
||||
// here on exactly the same terms the sweep would have used.
|
||||
if s.expireSeriesLocked(ser, now) {
|
||||
// Still flooding: it stays suppressed, this copy opens the count of
|
||||
// the new window.
|
||||
ser.count = 1
|
||||
s.seriesLRU.MoveToFront(ser.el)
|
||||
s.armRepeatTimerLocked()
|
||||
return
|
||||
}
|
||||
}
|
||||
s.openSeriesLocked(key, now)
|
||||
s.writeLineLocked(line)
|
||||
s.armRepeatTimerLocked()
|
||||
}
|
||||
|
||||
// repeatSeries is one message inside its window: the identity that is being
|
||||
// collapsed, how many copies have been swallowed since the head line or the
|
||||
// last summary, and when the current window ends.
|
||||
type repeatSeries struct {
|
||||
key string
|
||||
count int
|
||||
deadline time.Time
|
||||
el *list.Element // this series' node in Sink.seriesLRU
|
||||
}
|
||||
|
||||
// windowLocked is the effective suppression window (tests override s.window).
|
||||
func (s *Sink) windowLocked() time.Duration {
|
||||
if s.window > 0 {
|
||||
return s.window
|
||||
}
|
||||
return repeatWindow
|
||||
}
|
||||
|
||||
// openSeriesLocked starts a window for key, evicting the least recently seen
|
||||
// series first if the table is full. An evicted series that had swallowed
|
||||
// copies reports them on the way out, so the table's size limit can shorten a
|
||||
// window but can never lose a count.
|
||||
func (s *Sink) openSeriesLocked(key string, now time.Time) {
|
||||
for len(s.series) >= maxRepeatKeys {
|
||||
back := s.seriesLRU.Back()
|
||||
if back == nil {
|
||||
break
|
||||
}
|
||||
ev := back.Value.(*repeatSeries)
|
||||
if ev.count > 0 {
|
||||
s.summariseSeriesLocked(ev, " (repeat table full)")
|
||||
}
|
||||
s.dropSeriesLocked(ev)
|
||||
}
|
||||
ser := &repeatSeries{key: key, deadline: now.Add(s.windowLocked())}
|
||||
ser.el = s.seriesLRU.PushFront(ser)
|
||||
s.series[key] = ser
|
||||
}
|
||||
|
||||
// dropSeriesLocked forgets a series entirely (its next copy prints in full).
|
||||
func (s *Sink) dropSeriesLocked(ser *repeatSeries) {
|
||||
s.seriesLRU.Remove(ser.el)
|
||||
delete(s.series, ser.key)
|
||||
}
|
||||
|
||||
// expireSeriesLocked settles a series whose window has ended and reports
|
||||
// whether it stays open. One that swallowed copies prints their summary and
|
||||
// keeps its slot for another window — an ongoing flood must be acknowledged
|
||||
// every window without re-printing its head line. One that swallowed nothing is
|
||||
// forgotten: the message occurs rarely enough that it needs no collapsing at
|
||||
// all, and holding its slot would only push a real flood out of the table.
|
||||
func (s *Sink) expireSeriesLocked(ser *repeatSeries, now time.Time) bool {
|
||||
if ser.count == 0 {
|
||||
s.dropSeriesLocked(ser)
|
||||
return false
|
||||
}
|
||||
s.summariseSeriesLocked(ser, "")
|
||||
ser.count = 0
|
||||
ser.deadline = now.Add(s.windowLocked())
|
||||
return true
|
||||
}
|
||||
|
||||
// summariseSeriesLocked prints one summary line. It NAMES the message it counts
|
||||
// (the key: the level plus the text, i.e. the line minus the uptime field that
|
||||
// repeatKey drops): with several messages collapsed at once, a bare "last
|
||||
// message repeated N times" would leave the reader unable to tell which line
|
||||
// the number belongs to — the very confusion that made the old adjacent-only
|
||||
// suppression useless in the field. note marks a summary that was forced out
|
||||
// early (table full) rather than by its window.
|
||||
func (s *Sink) summariseSeriesLocked(ser *repeatSeries, note string) {
|
||||
unit := "times"
|
||||
if ser.count == 1 {
|
||||
unit = "time"
|
||||
}
|
||||
s.writeLineLocked([]byte(fmt.Sprintf("repeated %d %s%s: %s", ser.count, unit, note, ser.key)))
|
||||
}
|
||||
|
||||
// onRepeatWindow is the timer callback: it settles every series whose window
|
||||
// has ended, so a flood is reported while it happens instead of only when it
|
||||
// stops or when the daemon closes.
|
||||
func (s *Sink) onRepeatWindow() {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
if s.closed {
|
||||
return
|
||||
}
|
||||
now := s.now()
|
||||
for e := s.seriesLRU.Back(); e != nil; {
|
||||
prev := e.Prev() // taken before a possible removal of e
|
||||
ser := e.Value.(*repeatSeries)
|
||||
if !now.Before(ser.deadline) {
|
||||
s.expireSeriesLocked(ser, now)
|
||||
}
|
||||
e = prev
|
||||
}
|
||||
s.armRepeatTimerLocked()
|
||||
}
|
||||
|
||||
// armRepeatTimerLocked points the single repeat timer at the earliest deadline
|
||||
// in the table (and stops it when the table is empty). One timer for all series
|
||||
// keeps the cost at one goroutine wake per window, not one per message.
|
||||
func (s *Sink) armRepeatTimerLocked() {
|
||||
if s.closed {
|
||||
s.stopRepeatTimerLocked()
|
||||
return
|
||||
}
|
||||
var earliest time.Time
|
||||
for _, ser := range s.series {
|
||||
if earliest.IsZero() || ser.deadline.Before(earliest) {
|
||||
earliest = ser.deadline
|
||||
}
|
||||
}
|
||||
if earliest.IsZero() {
|
||||
s.stopRepeatTimerLocked()
|
||||
return
|
||||
}
|
||||
d := earliest.Sub(s.now())
|
||||
if d < 0 {
|
||||
d = 0
|
||||
}
|
||||
if s.repeatTimer == nil {
|
||||
s.repeatTimer = time.AfterFunc(d, s.onRepeatWindow)
|
||||
return
|
||||
}
|
||||
// A callback that already started is harmless: it blocks on s.mu and then
|
||||
// finds nothing expired.
|
||||
s.repeatTimer.Stop()
|
||||
s.repeatTimer.Reset(d)
|
||||
}
|
||||
|
||||
func (s *Sink) stopRepeatTimerLocked() {
|
||||
if s.repeatTimer != nil {
|
||||
s.repeatTimer.Stop()
|
||||
s.repeatTimer = nil
|
||||
}
|
||||
}
|
||||
|
||||
// flushAllSeriesLocked empties the table, printing a summary for every series
|
||||
// that had swallowed copies — oldest-seen first, so the summaries come out in
|
||||
// the order the messages last appeared. Used wherever the sink's state ends:
|
||||
// Close, Reconfigure (the counts belong to the destination they were headed
|
||||
// for) and a fatal/panic line.
|
||||
func (s *Sink) flushAllSeriesLocked() {
|
||||
s.stopRepeatTimerLocked()
|
||||
for e := s.seriesLRU.Back(); e != nil; e = e.Prev() {
|
||||
if ser := e.Value.(*repeatSeries); ser.count > 0 {
|
||||
s.summariseSeriesLocked(ser, "")
|
||||
}
|
||||
}
|
||||
s.seriesLRU.Init()
|
||||
s.series = make(map[string]*repeatSeries)
|
||||
}
|
||||
|
||||
// repeatKey reduces a raw producer line to the identity used for suppression,
|
||||
// and reports whether the line is EXEMPT from it.
|
||||
//
|
||||
// Both of the daemon's log producers (the control-plane formatter built in
|
||||
// cmd/shaterd/logsetup.go and the engine's own, both log.Formatter) render a
|
||||
// line as "<LEVEL>[<seconds since start>] <message>" — see log/format.go. The
|
||||
// bracketed uptime ticks every second, so comparing whole lines would suppress
|
||||
// nothing at all beyond a same-second burst; the key therefore drops exactly
|
||||
// that field and keeps everything else, LEVEL included (the same text at INFO
|
||||
// and at ERROR is not the same event).
|
||||
//
|
||||
// What the key deliberately does NOT drop is the per-connection "[id duration]"
|
||||
// group log.Formatter inserts for context-bound lines: those ids identify
|
||||
// distinct connections, and folding them together would turn "50 connections
|
||||
// failed" into one indistinguishable count. Such lines differ by id, so each
|
||||
// gets its own (single-copy, never summarised) series — which is correct, they
|
||||
// are not repeats. It also means they are the messages most likely to fill the
|
||||
// repeat table; that is what maxRepeatKeys and its least-recently-seen eviction
|
||||
// are for.
|
||||
//
|
||||
// The key doubles as the text a summary names itself with, which is why it
|
||||
// keeps the LEVEL and reads as a message on its own ("ERROR outbound/…: …").
|
||||
//
|
||||
// ANSI colour is stripped first so the key is identical for the coloured
|
||||
// (terminal) and plain (procd) renderings of the same message.
|
||||
//
|
||||
// FATAL/PANIC are exempt: they are emitted at most a handful of times, they are
|
||||
// the reason the operator is reading the log, and a summary line is a worse
|
||||
// thing to find than a duplicate.
|
||||
func repeatKey(line []byte) (string, bool) {
|
||||
b := stripANSI(line)
|
||||
i := 0
|
||||
for i < len(b) && b[i] >= 'A' && b[i] <= 'Z' {
|
||||
i++
|
||||
}
|
||||
if i > 0 && i < len(b) && b[i] == '[' {
|
||||
j := i + 1
|
||||
for j < len(b) && b[j] >= '0' && b[j] <= '9' {
|
||||
j++
|
||||
}
|
||||
if j > i+1 && j < len(b) && b[j] == ']' {
|
||||
level := string(b[:i])
|
||||
return level + string(b[j+1:]), level == "FATAL" || level == "PANIC"
|
||||
}
|
||||
}
|
||||
// Anything not in the daemon's format (a foreign writer, a bare line):
|
||||
// compare it whole.
|
||||
return string(b), false
|
||||
}
|
||||
|
||||
// writeLineLocked stamps one line with the UTC wall clock and fans it out.
|
||||
func (s *Sink) writeLineLocked(line []byte) {
|
||||
if !s.cfg.ToSyslog && !s.cfg.ToFile {
|
||||
return // fully off: the line is dropped, nowhere else to go
|
||||
}
|
||||
ts := s.now().UTC().Format(time.RFC3339)
|
||||
if s.cfg.ToSyslog && s.stderr != nil {
|
||||
out := make([]byte, 0, len(ts)+1+len(line)+1)
|
||||
|
||||
@@ -5,7 +5,10 @@ import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
@@ -379,6 +382,521 @@ func TestReconfigureFileOffDeletesSegments(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// --- repeat suppression -------------------------------------------------------
|
||||
|
||||
// syncBuffer is a bytes.Buffer safe for the concurrent access the repeat timer
|
||||
// goroutine and the test's reader make.
|
||||
type syncBuffer struct {
|
||||
mu sync.Mutex
|
||||
b bytes.Buffer
|
||||
}
|
||||
|
||||
func (s *syncBuffer) Write(p []byte) (int, error) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
return s.b.Write(p)
|
||||
}
|
||||
|
||||
func (s *syncBuffer) String() string {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
return s.b.String()
|
||||
}
|
||||
|
||||
// engLine renders a line exactly as both of the daemon's log.Formatter
|
||||
// producers do: "<LEVEL>[<seconds since start>] <message>".
|
||||
func engLine(sec int, level, msg string) string {
|
||||
return fmt.Sprintf("%s[%04d] %s\n", level, sec, msg)
|
||||
}
|
||||
|
||||
// payloads strips the sink's RFC3339 prefix off every line of a rendered log.
|
||||
func payloads(t *testing.T, body string) []string {
|
||||
t.Helper()
|
||||
if body == "" {
|
||||
return nil
|
||||
}
|
||||
var out []string
|
||||
for _, line := range strings.Split(strings.TrimSuffix(body, "\n"), "\n") {
|
||||
if _, ok := LineTime([]byte(line)); !ok {
|
||||
t.Fatalf("line without RFC3339 prefix: %q", line)
|
||||
}
|
||||
out = append(out, line[strings.IndexByte(line, ' ')+1:])
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// A summary names the message it counts: "repeated N time(s)[ (note)]: <msg>".
|
||||
var repeatSummaryRe = regexp.MustCompile(`^repeated (\d+) times?( \([^)]*\))?: (.*)$`)
|
||||
|
||||
// summaryByMessage returns how many suppressed copies the summaries in payloads
|
||||
// account for PER NAMED MESSAGE, and how many summary lines there were.
|
||||
func summaryByMessage(t *testing.T, lines []string) (map[string]int, int) {
|
||||
t.Helper()
|
||||
per := make(map[string]int)
|
||||
count := 0
|
||||
for _, l := range lines {
|
||||
m := repeatSummaryRe.FindStringSubmatch(l)
|
||||
if m == nil {
|
||||
continue
|
||||
}
|
||||
n, err := strconv.Atoi(m[1])
|
||||
if err != nil {
|
||||
t.Fatalf("unparseable summary %q: %v", l, err)
|
||||
}
|
||||
per[m[3]] += n
|
||||
count++
|
||||
}
|
||||
return per, count
|
||||
}
|
||||
|
||||
// summarySum returns how many suppressed lines the summaries in payloads
|
||||
// account for, and how many summary lines there were.
|
||||
func summarySum(t *testing.T, lines []string) (total, count int) {
|
||||
t.Helper()
|
||||
per, count := summaryByMessage(t, lines)
|
||||
for _, n := range per {
|
||||
total += n
|
||||
}
|
||||
return total, count
|
||||
}
|
||||
|
||||
// TestRepeatRunCollapsed: the simplest shape of the production symptom — one
|
||||
// broken chain repeating the same message back to back — leaves ONE copy of the
|
||||
// line plus ONE summary carrying the right count and naming the message, in
|
||||
// BOTH halves. The uptime field of the producer's prefix differs on every line:
|
||||
// that is exactly what repeatKey must ignore. An unrelated message in between
|
||||
// no longer ENDS the series (a series lives for its window, not until the next
|
||||
// distinct line) — it is simply printed, and the summary follows at Close.
|
||||
func TestRepeatRunCollapsed(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, true, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = time.Hour // no timer flush: this test is about series boundaries
|
||||
|
||||
const msg = "outbound/urltest[chain-ewan-wg-subs-h2]: WireGuard is not ready yet"
|
||||
for sec := 1; sec <= 6; sec++ {
|
||||
if _, err := s.Write([]byte(engLine(sec, "ERROR", msg))); err != nil {
|
||||
t.Fatalf("Write: %v", err)
|
||||
}
|
||||
}
|
||||
_, _ = s.Write([]byte(engLine(7, "INFO", "something else entirely")))
|
||||
if err := s.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
|
||||
for _, src := range []struct{ name, body string }{
|
||||
{"file", readFile(t, cfg.Path)},
|
||||
{"stderr", stderr.String()},
|
||||
} {
|
||||
got := payloads(t, src.body)
|
||||
want := []string{
|
||||
strings.TrimSuffix(engLine(1, "ERROR", msg), "\n"),
|
||||
strings.TrimSuffix(engLine(7, "INFO", "something else entirely"), "\n"),
|
||||
"repeated 5 times: ERROR " + msg,
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("%s: got %d lines %q, want %d %q", src.name, len(got), got, len(want), want)
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("%s line %d = %q, want %q", src.name, i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatAlternatingSuppressedPerMessage is the reworked
|
||||
// TestRepeatAlternatingNotSuppressed. Its old contract — A B A B is four
|
||||
// distinct events, nothing may be swallowed — was the bug: on the router the
|
||||
// flood ALWAYS alternates (one message per broken chain), so adjacency-only
|
||||
// suppression collapsed nothing at all.
|
||||
//
|
||||
// The property that test really guarded, and that this one still guards, is
|
||||
// that distinct messages are never folded into one count: A and B each get
|
||||
// their OWN series, their own summary and their own N. What changed is that the
|
||||
// second copy of each is now suppressed instead of printed.
|
||||
func TestRepeatAlternatingSuppressedPerMessage(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, true, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = time.Hour
|
||||
|
||||
for sec := 1; sec <= 4; sec++ {
|
||||
msg := "alpha happened"
|
||||
if sec%2 == 0 {
|
||||
msg = "beta happened"
|
||||
}
|
||||
_, _ = s.Write([]byte(engLine(sec, "WARN", msg)))
|
||||
}
|
||||
_ = s.Close()
|
||||
|
||||
for _, src := range []struct{ name, body string }{
|
||||
{"file", readFile(t, cfg.Path)},
|
||||
{"stderr", stderr.String()},
|
||||
} {
|
||||
got := payloads(t, src.body)
|
||||
want := []string{
|
||||
strings.TrimSuffix(engLine(1, "WARN", "alpha happened"), "\n"),
|
||||
strings.TrimSuffix(engLine(2, "WARN", "beta happened"), "\n"),
|
||||
"repeated 1 time: WARN alpha happened",
|
||||
"repeated 1 time: WARN beta happened",
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("%s: got %d lines %q, want %d %q", src.name, len(got), got, len(want), want)
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("%s line %d = %q, want %q", src.name, i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
// The counts are per message — never merged into one "repeated 2".
|
||||
per, n := summaryByMessage(t, got)
|
||||
if n != 2 || per["WARN alpha happened"] != 1 || per["WARN beta happened"] != 1 {
|
||||
t.Errorf("%s: summaries %v (%d lines), want one per message with N=1", src.name, per, n)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatInterleavedFloodCollapsed is the field case that motivated the
|
||||
// table: the engine cycles the SAME message over three chain tags, so no two
|
||||
// identical lines are ever adjacent. 300 lines must leave 3 printed heads and
|
||||
// exactly 3 summaries, each naming its own message with its own count.
|
||||
func TestRepeatInterleavedFloodCollapsed(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, true, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = time.Hour // the summaries come from Close, deterministically
|
||||
|
||||
msgs := []string{
|
||||
"outbound/urltest[chain-ewan-wg-subs-h2]: WireGuard is not ready yet",
|
||||
"outbound/urltest[chain-ewan-wg-subs-h3]: WireGuard is not ready yet",
|
||||
"outbound/urltest[chain-ewan-wg-subs-h4]: WireGuard is not ready yet",
|
||||
}
|
||||
const copies = 100
|
||||
for i := 0; i < copies; i++ {
|
||||
for _, m := range msgs {
|
||||
if _, err := s.Write([]byte(engLine(i, "ERROR", m))); err != nil {
|
||||
t.Fatalf("Write: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
if err := s.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
|
||||
for _, src := range []struct{ name, body string }{
|
||||
{"file", readFile(t, cfg.Path)},
|
||||
{"stderr", stderr.String()},
|
||||
} {
|
||||
got := payloads(t, src.body)
|
||||
var want []string
|
||||
for _, m := range msgs { // the heads, in first-seen order
|
||||
want = append(want, strings.TrimSuffix(engLine(0, "ERROR", m), "\n"))
|
||||
}
|
||||
for _, m := range msgs { // the summaries, oldest-seen first
|
||||
want = append(want, fmt.Sprintf("repeated %d times: ERROR %s", copies-1, m))
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("%s: got %d lines %q, want %d %q", src.name, len(got), got, len(want), want)
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("%s line %d = %q, want %q", src.name, i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
per, n := summaryByMessage(t, got)
|
||||
if n != len(msgs) {
|
||||
t.Errorf("%s: %d summary lines, want %d (one per message)", src.name, n, len(msgs))
|
||||
}
|
||||
for _, m := range msgs {
|
||||
if per["ERROR "+m] != copies-1 {
|
||||
t.Errorf("%s: message %q summarised %d copies, want %d", src.name, m, per["ERROR "+m], copies-1)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatTableOverflowReportsEviction: the table is bounded, and hitting the
|
||||
// bound never drops a count silently — the least-recently-seen series is
|
||||
// flushed with a summary that says WHY it was cut short, and the table stays at
|
||||
// its limit.
|
||||
func TestRepeatTableOverflowReportsEviction(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, false, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = time.Hour
|
||||
|
||||
// Fill the table, every key with exactly one swallowed copy to lose.
|
||||
for i := 0; i < maxRepeatKeys; i++ {
|
||||
msg := fmt.Sprintf("chain-%03d is down", i)
|
||||
_, _ = s.Write([]byte(engLine(i, "ERROR", msg)))
|
||||
_, _ = s.Write([]byte(engLine(i, "ERROR", msg)))
|
||||
}
|
||||
// One key too many: the oldest series (chain-000) must be evicted, and its
|
||||
// swallowed copy must be reported on the way out.
|
||||
_, _ = s.Write([]byte(engLine(999, "ERROR", "one key too many")))
|
||||
|
||||
want := "repeated 1 time (repeat table full): ERROR chain-000 is down"
|
||||
got := payloads(t, stderr.String())
|
||||
found := false
|
||||
for _, l := range got {
|
||||
if l == want {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
tail := got
|
||||
if len(tail) > 4 {
|
||||
tail = tail[len(tail)-4:]
|
||||
}
|
||||
t.Fatalf("eviction was silent: no %q (tail of the output: %q)", want, tail)
|
||||
}
|
||||
s.mu.Lock()
|
||||
size, lru := len(s.series), s.seriesLRU.Len()
|
||||
s.mu.Unlock()
|
||||
if size != maxRepeatKeys || lru != maxRepeatKeys {
|
||||
t.Errorf("table holds %d entries (lru %d), want the cap %d", size, lru, maxRepeatKeys)
|
||||
}
|
||||
|
||||
_ = s.Close()
|
||||
// Nothing anywhere was lost: every written line is either printed or counted.
|
||||
got = payloads(t, stderr.String())
|
||||
suppressed, summaries := summarySum(t, got)
|
||||
printed := len(got) - summaries
|
||||
if wrote := 2*maxRepeatKeys + 1; printed+suppressed != wrote {
|
||||
t.Errorf("%d printed + %d suppressed = %d, want %d written", printed, suppressed, printed+suppressed, wrote)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatSeriesSurviveReconfigure: a live Reconfigure settles every open
|
||||
// series into the destination its swallowed copies were headed for, and starts
|
||||
// the next configuration with an empty table.
|
||||
func TestRepeatSeriesSurviveReconfigure(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
cfgA := Config{ToFile: true, Path: filepath.Join(dir, "a.log")}
|
||||
cfgB := Config{ToFile: true, Path: filepath.Join(dir, "b.log")}
|
||||
s := New(nil, cfgA)
|
||||
s.window = time.Hour
|
||||
|
||||
// Two interleaved series: alpha keeps 2 swallowed copies, beta keeps 1.
|
||||
for sec, msg := range []string{"alpha down", "beta down", "alpha down", "beta down", "alpha down"} {
|
||||
_, _ = s.Write([]byte(engLine(sec+1, "ERROR", msg)))
|
||||
}
|
||||
s.Reconfigure(cfgB)
|
||||
_, _ = s.Write([]byte(engLine(6, "ERROR", "alpha down")))
|
||||
if err := s.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
|
||||
a := payloads(t, readFile(t, cfgA.Path))
|
||||
perA, nA := summaryByMessage(t, a)
|
||||
if nA != 2 || perA["ERROR alpha down"] != 2 || perA["ERROR beta down"] != 1 {
|
||||
t.Errorf("a.log summaries %v (%d lines), want alpha=2 beta=1 flushed by Reconfigure: %q", perA, nA, a)
|
||||
}
|
||||
b := payloads(t, readFile(t, cfgB.Path))
|
||||
if _, nB := summaryByMessage(t, b); nB != 0 {
|
||||
t.Errorf("b.log carries summaries of lines written before the swap: %q", b)
|
||||
}
|
||||
// The new configuration starts fresh: the message prints in full again.
|
||||
if len(b) != 1 || b[0] != strings.TrimSuffix(engLine(6, "ERROR", "alpha down"), "\n") {
|
||||
t.Errorf("b.log = %q, want the re-printed head line only", b)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatLastRunSurvivesClose: a run still open at shutdown is summarised by
|
||||
// Close — the last series is never silently lost.
|
||||
func TestRepeatLastRunSurvivesClose(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, true, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = time.Hour
|
||||
|
||||
for sec := 1; sec <= 4; sec++ {
|
||||
_, _ = s.Write([]byte(engLine(sec, "ERROR", "dying in a loop")))
|
||||
}
|
||||
if err := s.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
|
||||
for _, src := range []struct{ name, body string }{
|
||||
{"file", readFile(t, cfg.Path)},
|
||||
{"stderr", stderr.String()},
|
||||
} {
|
||||
got := payloads(t, src.body)
|
||||
if len(got) != 2 {
|
||||
t.Fatalf("%s: got %q, want the line plus one summary", src.name, got)
|
||||
}
|
||||
if want := "repeated 3 times: ERROR dying in a loop"; got[1] != want {
|
||||
t.Errorf("%s: summary = %q, want %q", src.name, got[1], want)
|
||||
}
|
||||
}
|
||||
|
||||
// A run of exactly two lines is summarised with correct grammar.
|
||||
var stderr2 syncBuffer
|
||||
cfg2 := fileCfg(t, true, false, 0)
|
||||
s2 := New(&stderr2, cfg2)
|
||||
s2.window = time.Hour
|
||||
_, _ = s2.Write([]byte(engLine(1, "WARN", "twice only")))
|
||||
_, _ = s2.Write([]byte(engLine(2, "WARN", "twice only")))
|
||||
_ = s2.Close()
|
||||
want := "repeated 1 time: WARN twice only"
|
||||
if got := payloads(t, stderr2.String()); len(got) != 2 || got[1] != want {
|
||||
t.Errorf("two-line run rendered as %q, want the line plus %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatFatalNeverSuppressed: fatal/panic lines are exempt — every copy is
|
||||
// printed, and they do not become the head of a run either.
|
||||
func TestRepeatFatalNeverSuppressed(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, true, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = time.Hour
|
||||
|
||||
for sec := 1; sec <= 3; sec++ {
|
||||
_, _ = s.Write([]byte(engLine(sec, "FATAL", "engine is gone")))
|
||||
}
|
||||
for sec := 4; sec <= 6; sec++ {
|
||||
_, _ = s.Write([]byte(engLine(sec, "PANIC", "engine is gone")))
|
||||
}
|
||||
_ = s.Close()
|
||||
|
||||
for _, src := range []struct{ name, body string }{
|
||||
{"file", readFile(t, cfg.Path)},
|
||||
{"stderr", stderr.String()},
|
||||
} {
|
||||
got := payloads(t, src.body)
|
||||
if len(got) != 6 {
|
||||
t.Fatalf("%s: got %d lines %q, want all 6 fatal/panic copies", src.name, len(got), got)
|
||||
}
|
||||
if _, n := summarySum(t, got); n != 0 {
|
||||
t.Errorf("%s: fatal/panic run produced %d summaries: %q", src.name, n, got)
|
||||
}
|
||||
}
|
||||
|
||||
// A fatal line also settles everything the table was holding BEFORE it
|
||||
// speaks: the process may never reach Close, and a swallowed count that is
|
||||
// never reported is the failure this mechanism exists to prevent.
|
||||
var stderr2 syncBuffer
|
||||
s2 := New(&stderr2, fileCfg(t, true, false, 0))
|
||||
s2.window = time.Hour
|
||||
for sec := 1; sec <= 3; sec++ {
|
||||
_, _ = s2.Write([]byte(engLine(sec, "ERROR", "about to die")))
|
||||
}
|
||||
_, _ = s2.Write([]byte(engLine(4, "FATAL", "engine is gone")))
|
||||
got := payloads(t, stderr2.String())
|
||||
want := []string{
|
||||
strings.TrimSuffix(engLine(1, "ERROR", "about to die"), "\n"),
|
||||
"repeated 2 times: ERROR about to die",
|
||||
strings.TrimSuffix(engLine(4, "FATAL", "engine is gone"), "\n"),
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("got %q, want %q", got, want)
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("line %d = %q, want %q", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
_ = s2.Close()
|
||||
}
|
||||
|
||||
// TestRepeatWindowFlushesOngoingRun: a flood that never stops still reports
|
||||
// itself — the window timer emits a summary without any further input, and the
|
||||
// counter restarts (so the NEXT summary counts only the lines after it).
|
||||
func TestRepeatWindowFlushesOngoingRun(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, false, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = 50 * time.Millisecond
|
||||
|
||||
for sec := 1; sec <= 5; sec++ {
|
||||
_, _ = s.Write([]byte(engLine(sec, "ERROR", "flooding")))
|
||||
}
|
||||
want := "repeated 4 times: ERROR flooding"
|
||||
deadline := time.Now().Add(5 * time.Second)
|
||||
for !strings.Contains(stderr.String(), "repeated ") && time.Now().Before(deadline) {
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
got := payloads(t, stderr.String())
|
||||
if len(got) != 2 || got[1] != want {
|
||||
t.Fatalf("timer flush produced %q, want the line plus %q", got, want)
|
||||
}
|
||||
|
||||
// The flood continues. Whether the second batch lands inside the window the
|
||||
// flush opened (swallowed) or after it lapsed (a fresh head line) is a race
|
||||
// with the timer, so what is asserted is what must hold either way: the
|
||||
// flood is reported AGAIN, every copy is accounted for exactly once, and
|
||||
// every summary names this one message.
|
||||
for sec := 6; sec <= 8; sec++ {
|
||||
_, _ = s.Write([]byte(engLine(sec, "ERROR", "flooding")))
|
||||
}
|
||||
_ = s.Close()
|
||||
got = payloads(t, stderr.String())
|
||||
per, summaries := summaryByMessage(t, got)
|
||||
suppressed, _ := summarySum(t, got)
|
||||
printed := len(got) - summaries
|
||||
if printed+suppressed != 8 {
|
||||
t.Errorf("%d printed + %d suppressed = %d, want the 8 written lines: %q", printed, suppressed, printed+suppressed, got)
|
||||
}
|
||||
if summaries < 2 {
|
||||
t.Errorf("the continuing flood was reported %d time(s), want a second summary: %q", summaries, got)
|
||||
}
|
||||
if len(per) != 1 || per["ERROR flooding"] != suppressed {
|
||||
t.Errorf("summaries %v, want all of them naming %q", per, "ERROR flooding")
|
||||
}
|
||||
}
|
||||
|
||||
// TestRepeatConcurrentWritesKeepCount: the sink is written from every goroutine
|
||||
// of the engine. Under -race, a hammering of identical lines must not lose or
|
||||
// double-count anything: printed copies + summarised copies == lines written.
|
||||
func TestRepeatConcurrentWritesKeepCount(t *testing.T) {
|
||||
var stderr syncBuffer
|
||||
cfg := fileCfg(t, true, true, 0)
|
||||
s := New(&stderr, cfg)
|
||||
s.window = 20 * time.Millisecond // let the timer race the writers on purpose
|
||||
|
||||
const (
|
||||
writers = 8
|
||||
perGo = 100
|
||||
expected = writers * perGo
|
||||
)
|
||||
var wg sync.WaitGroup
|
||||
for g := 0; g < writers; g++ {
|
||||
wg.Add(1)
|
||||
go func(g int) {
|
||||
defer wg.Done()
|
||||
for i := 0; i < perGo; i++ {
|
||||
_, _ = s.Write([]byte(engLine(g*perGo+i, "ERROR", "concurrent flood")))
|
||||
}
|
||||
}(g)
|
||||
}
|
||||
wg.Wait()
|
||||
if err := s.Close(); err != nil {
|
||||
t.Fatalf("Close: %v", err)
|
||||
}
|
||||
|
||||
for _, src := range []struct{ name, body string }{
|
||||
{"file", readFile(t, cfg.Path)},
|
||||
{"stderr", stderr.String()},
|
||||
} {
|
||||
got := payloads(t, src.body)
|
||||
suppressed, summaries := summarySum(t, got)
|
||||
printed := len(got) - summaries
|
||||
if printed+suppressed != expected {
|
||||
t.Errorf("%s: %d printed + %d suppressed = %d, want %d (lines: %q)",
|
||||
src.name, printed, suppressed, printed+suppressed, expected, got)
|
||||
}
|
||||
if printed < 1 {
|
||||
t.Errorf("%s: nothing printed at all — the first line must always show", src.name)
|
||||
}
|
||||
// Suppression really happened (the whole point).
|
||||
if printed > expected/10 {
|
||||
t.Errorf("%s: %d of %d lines printed — suppression did not engage", src.name, printed, expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewFileOffPurgesLeftovers: a sink CREATED with the file off (daemon boot
|
||||
// with the toggle already off) deletes segments left by a previous life.
|
||||
func TestNewFileOffPurgesLeftovers(t *testing.T) {
|
||||
|
||||
@@ -124,8 +124,36 @@ type RuleReach struct {
|
||||
ShadowedByOrder int `json:"shadowed_by_order,omitempty"`
|
||||
// Reason is the operator-facing sentence; "" when Unreachable is false.
|
||||
Reason string `json:"reason,omitempty"`
|
||||
|
||||
// EffectiveEnabled is whether the rule is IN FORCE right now — Rule.Enabled
|
||||
// after the active WAN profile's enable/disable overrides. It is NOT the flag
|
||||
// GET /api/config carries: that one is the desired state the panel PUTs back,
|
||||
// and on a router with profiles the two legitimately disagree.
|
||||
//
|
||||
// This field exists because the panel had no way to tell them apart and so drew
|
||||
// the desired state as if it were the truth: a config with `ewan-default` and
|
||||
// `swan-default` both `enabled '1'` in UCI, under a profile that enables the
|
||||
// first and disables the second, showed BOTH switches on while the engine ran
|
||||
// only one chain. Same defect class as the two-`default` shadowing above — the
|
||||
// interface claiming a setting is in force when it is not.
|
||||
EffectiveEnabled bool `json:"effective_enabled"`
|
||||
// OverriddenBy names the active profile that CHANGED this rule's state, and
|
||||
// Override says which way it went: "enabled" or "disabled". Both are empty
|
||||
// unless the profile actually flipped the outcome — a profile that disables a
|
||||
// rule already switched off in UCI has overridden nothing the operator can see,
|
||||
// and saying otherwise would put a profile's name on every row it merely
|
||||
// mentions.
|
||||
OverriddenBy string `json:"overridden_by,omitempty"`
|
||||
Override string `json:"override,omitempty"`
|
||||
}
|
||||
|
||||
// Override directions carried by RuleReach.Override. Values are part of the
|
||||
// /api/rules/reachability contract; the panel switches on them.
|
||||
const (
|
||||
RuleOverrideEnabled = "enabled"
|
||||
RuleOverrideDisabled = "disabled"
|
||||
)
|
||||
|
||||
// RuleReachability returns one verdict per rule, in the INPUT slice's order.
|
||||
//
|
||||
// The input must already be the EFFECTIVE rule set — profile enable/disable
|
||||
@@ -147,6 +175,10 @@ func RuleReachability(rules []Rule) []RuleReach {
|
||||
for i := range rules {
|
||||
out[i] = RuleReach{
|
||||
Index: i, Name: rules[i].Name, Order: rules[i].Order, ShadowedByIndex: -1,
|
||||
// The input IS the effective set (see the doc comment), so its Enabled flag
|
||||
// is the effective one. Who overrode it — if anyone — is not knowable from
|
||||
// this slice alone; Model.EffectiveRuleReachability fills that in.
|
||||
EffectiveEnabled: rules[i].Enabled,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -208,3 +240,46 @@ func (m *Model) EffectiveRules() ([]Rule, []Warning) {
|
||||
rules, owarns := ApplyProfileRuleOverrides(m.Rules, prof)
|
||||
return rules, append(warns, owarns...)
|
||||
}
|
||||
|
||||
// EffectiveRuleReachability is RuleReachability over m's EFFECTIVE rules, with
|
||||
// each verdict annotated by the active profile's override. It returns one verdict
|
||||
// per rule in m.Rules, at the SAME index — ApplyProfileRuleOverrides returns a
|
||||
// same-length copy in the same order, and RuleReachability preserves its input's
|
||||
// order, so a client can zip the result with GET /api/config's `Rules`.
|
||||
//
|
||||
// The override annotation is a DIFF, not a second reading of the profile's name
|
||||
// lists: a verdict is marked only where the effective flag actually differs from
|
||||
// the desired-state one. That is deliberate on both counts —
|
||||
//
|
||||
// - it cannot drift from ApplyProfileRuleOverrides, because it observes that
|
||||
// function's output rather than re-deciding what it should have done (an
|
||||
// unmigrated rule the profile is forbidden to enable, for instance, produces no
|
||||
// diff and therefore no annotation, with nothing here having to know the rule);
|
||||
// - and it answers the question the operator is actually asking, which is not
|
||||
// "does the profile mention this rule" but "is this row's switch telling me the
|
||||
// truth".
|
||||
func (m *Model) EffectiveRuleReachability() ([]RuleReach, []Warning) {
|
||||
if m == nil {
|
||||
return []RuleReach{}, nil
|
||||
}
|
||||
prof, warns := ResolveActiveProfile(m)
|
||||
eff, owarns := ApplyProfileRuleOverrides(m.Rules, prof)
|
||||
warns = append(warns, owarns...)
|
||||
|
||||
out := RuleReachability(eff)
|
||||
if prof == nil {
|
||||
return out, warns
|
||||
}
|
||||
for i := range out {
|
||||
if i >= len(m.Rules) || m.Rules[i].Enabled == eff[i].Enabled {
|
||||
continue
|
||||
}
|
||||
out[i].OverriddenBy = prof.Name
|
||||
if eff[i].Enabled {
|
||||
out[i].Override = RuleOverrideEnabled
|
||||
} else {
|
||||
out[i].Override = RuleOverrideDisabled
|
||||
}
|
||||
}
|
||||
return out, warns
|
||||
}
|
||||
|
||||
+68
-11
@@ -119,8 +119,16 @@ func (s *Server) handleConfig(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
}
|
||||
|
||||
// configGetRead is handleConfigGet's test seam (same pattern as reachConfigRead /
|
||||
// logConfigRead). It exists so a test can serve GET /api/config and
|
||||
// GET /api/rules/reachability from ONE canned model and compare them: the promise
|
||||
// that the two responses line up by `index` is a promise about a pair of
|
||||
// endpoints, and asserting it against a slice the test holds itself would only
|
||||
// re-check the fixture.
|
||||
var configGetRead = model.ReadUCI
|
||||
|
||||
func (s *Server) handleConfigGet(w http.ResponseWriter, r *http.Request) {
|
||||
m, err := model.ReadUCI()
|
||||
m, err := configGetRead()
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, "read config: "+err.Error())
|
||||
return
|
||||
@@ -153,11 +161,19 @@ var reachConfigRead = model.ReadUCI
|
||||
// it.
|
||||
//
|
||||
// The verdict is computed over the EFFECTIVE rules (active WAN profile's
|
||||
// enable/disable applied via model.EffectiveRules), because a rule the active
|
||||
// profile switched off is not in force and must not be blamed for retiring
|
||||
// enable/disable applied, via model.EffectiveRuleReachability), because a rule the
|
||||
// active profile switched off is not in force and must not be blamed for retiring
|
||||
// anything. Profile-resolution warnings are dropped here: this endpoint answers
|
||||
// one question, and the same warnings already reach the operator through
|
||||
// `shaterd status` on every apply.
|
||||
//
|
||||
// The same call also carries `effective_enabled` (+ `overridden_by`/`override`),
|
||||
// which is this endpoint's answer to "is the switch on that row telling the truth".
|
||||
// It belongs HERE and not on /api/config for the reason above: /api/config is the
|
||||
// desired state, PUT back verbatim, and a rule's Enabled flag there must keep
|
||||
// meaning "what the operator asked for" even while a profile is overriding it.
|
||||
// Both facts come from one read of the config, so they cannot describe different
|
||||
// configs.
|
||||
func (s *Server) handleRulesReachability(w http.ResponseWriter, r *http.Request) {
|
||||
if r.Method != http.MethodGet {
|
||||
writeError(w, http.StatusMethodNotAllowed, "method not allowed")
|
||||
@@ -168,8 +184,7 @@ func (s *Server) handleRulesReachability(w http.ResponseWriter, r *http.Request)
|
||||
writeError(w, http.StatusInternalServerError, "read config: "+err.Error())
|
||||
return
|
||||
}
|
||||
rules, _ := m.EffectiveRules()
|
||||
out := model.RuleReachability(rules)
|
||||
out, _ := m.EffectiveRuleReachability()
|
||||
if out == nil {
|
||||
out = []model.RuleReach{}
|
||||
}
|
||||
@@ -1344,9 +1359,24 @@ type groupTestStartResponse struct {
|
||||
Reason string `json:"reason,omitempty"`
|
||||
}
|
||||
|
||||
// handleGroupsTest → /api/groups/test: measure each target's (group's or chain's)
|
||||
// latency and the public address it currently exits from. A chain is dialled via
|
||||
// its exit wrapper, so the numbers are end-to-end properties of the whole path.
|
||||
// handleGroupsTest → /api/groups/test: report each target's (group's or chain's)
|
||||
// health as the OBSERVATORY just measured it, plus the public address it
|
||||
// currently exits from.
|
||||
//
|
||||
// POST no longer dials anything. It asks the engine's observatory for an
|
||||
// out-of-turn pass of its probe plan and then reports what that pass measured:
|
||||
// each target resolves off the health board the moment its measured path
|
||||
// carries an observation newer than the button press. The consequences the
|
||||
// frontend should expect (the response SHAPE is unchanged):
|
||||
// - delay_ms/ok/tested_unix describe the observatory's measurement of the
|
||||
// target's real routed path — tested_unix is when that observation was
|
||||
// taken, which may be a second or two before the poll that delivered it;
|
||||
// - a target no enabled rule routes through is never measured and comes back
|
||||
// ok=false with an error saying so, immediately — the observatory only
|
||||
// probes paths the rules use, and the panel should render that as a
|
||||
// routing fact, not a failure;
|
||||
// - a run can take up to the engine's wait deadline (120s) when the
|
||||
// observatory has a large plan to walk; poll GET for progress as before.
|
||||
//
|
||||
// Contract for the frontend:
|
||||
// - POST {"name":"auto"} → test that group or chain; an empty/absent name tests
|
||||
@@ -1380,9 +1410,11 @@ func (s *Server) handleGroupsTest(w http.ResponseWriter, r *http.Request) {
|
||||
if n := strings.TrimSpace(req.Name); n != "" {
|
||||
names = []string{n}
|
||||
}
|
||||
// Probe URL from the model (best-effort) — the global one, the only probe
|
||||
// instrument left; a UCI read failure falls back to the engine's built-in
|
||||
// default.
|
||||
// Probe URL from the model (best-effort), passed through for signature
|
||||
// stability only: the engine IGNORES it now. The probe URL is a global
|
||||
// observatory setting installed at apply time, and the manual test reads
|
||||
// the observatory's measurements rather than dialling with a URL of its
|
||||
// own — see engine.TestGroups.
|
||||
var probeURL string
|
||||
if m, err := model.ReadUCI(); err == nil {
|
||||
probeURL = strings.TrimSpace(m.Globals.ProbeURL)
|
||||
@@ -1454,6 +1486,31 @@ type groupHealthResponse struct {
|
||||
// was measured THROUGH the group's egress, so it is not comparable with the global
|
||||
// per-node numbers in /api/stats node_health. That is the entire point of this
|
||||
// endpoint — the same node in two groups with two egresses has two health states.
|
||||
//
|
||||
// chains[] additionally carries the PER-HOP readout for every chain the running
|
||||
// box materialised:
|
||||
//
|
||||
// chains[].hops = [{index,tag,kind,exit,state,delay_ms,age_seconds,selected,
|
||||
// total,tested,alive,dead,untested}...]
|
||||
//
|
||||
// - hops is L1..Ln in wire order (index is 1-based); the entry with exit=true
|
||||
// is the last hop, where traffic leaves to the internet. Each hop's numbers
|
||||
// measure the chain PREFIX up to and including that hop — the observatory
|
||||
// dials the hop wrappers — so a dead hop N with alive hops 1..N-1 localises
|
||||
// the failure to hop N's own leg.
|
||||
// - kind is "node" (one measurement: total=1, selected empty) or "group" (a
|
||||
// roll-up of the hop's member copies, with the same counter invariants as a
|
||||
// group: tested == alive+dead, alive+dead+untested == total; selected is
|
||||
// the node NAME the hop currently picks; delay_ms/age_seconds are the
|
||||
// selected member's observation, or the freshest alive member's when the
|
||||
// selection has none).
|
||||
// - state follows the same closed set and the same honesty rule as groups:
|
||||
// "dead" only on a positive finding, "untested" for no fresh data — and for
|
||||
// a group hop, "alive" as long as ANY member answers.
|
||||
// - an ABSENT/empty hops key means the running box never materialised the
|
||||
// chain (unused, or a 1-hop chain that resolves straight to its target) —
|
||||
// it must NOT be read as "this chain has no hops"; the config, not this
|
||||
// endpoint, knows the configured hop count.
|
||||
func (s *Server) handleGroupsHealth(w http.ResponseWriter, r *http.Request) {
|
||||
if r.Method != http.MethodGet {
|
||||
writeError(w, http.StatusMethodNotAllowed, "method not allowed")
|
||||
|
||||
@@ -95,6 +95,185 @@ func TestRulesReachabilityEmptyIsArray(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// profileModel is the field config that made the effective-state fields
|
||||
// necessary: two rules both `enabled '1'` in UCI, and an active profile that
|
||||
// force-enables one and disables the other. The engine runs one chain; the panel
|
||||
// used to draw two live switches.
|
||||
func profileModel(activeProfile string) *model.Model {
|
||||
return &model.Model{
|
||||
Globals: model.Globals{ActiveProfile: activeProfile},
|
||||
Profiles: []model.Profile{{
|
||||
Name: "ethernet-uplink",
|
||||
Enabled: true,
|
||||
EnableRules: []string{"ewan-default"},
|
||||
DisableRules: []string{"swan-default"},
|
||||
}},
|
||||
Rules: []model.Rule{
|
||||
{Name: "ewan-default", Enabled: true, Order: 10, Target: "direct", DstRuleset: []string{"ru-inside"}},
|
||||
{Name: "swan-default", Enabled: true, Order: 20, Target: "group:auto", DstRuleset: []string{"ru-inside"}},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestRulesReachabilityProfileDisabled: the reported bug. `swan-default` is
|
||||
// `enabled '1'` in UCI, the active profile disables it, and the endpoint must say
|
||||
// so — otherwise the row keeps claiming a rule is in force that the engine never
|
||||
// loaded.
|
||||
func TestRulesReachabilityProfileDisabled(t *testing.T) {
|
||||
s := newTestServer(t)
|
||||
srv := httptest.NewServer(s.Handler())
|
||||
defer srv.Close()
|
||||
cookie := login(t, srv, s)
|
||||
|
||||
orig := reachConfigRead
|
||||
defer func() { reachConfigRead = orig }()
|
||||
reachConfigRead = func() (*model.Model, error) { return profileModel("ethernet-uplink"), nil }
|
||||
|
||||
code, out := getReach(t, srv, cookie)
|
||||
if code != http.StatusOK {
|
||||
t.Fatalf("got %d, want 200", code)
|
||||
}
|
||||
if len(out.Rules) != 2 {
|
||||
t.Fatalf("want one verdict per rule (2), got %d", len(out.Rules))
|
||||
}
|
||||
sw := out.Rules[1]
|
||||
if sw.Name != "swan-default" {
|
||||
t.Fatalf("verdicts must keep the model's rule order: %+v", out.Rules)
|
||||
}
|
||||
if sw.EffectiveEnabled {
|
||||
t.Fatalf("the active profile disables this rule, so it is NOT in force: %+v", sw)
|
||||
}
|
||||
if sw.OverriddenBy != "ethernet-uplink" {
|
||||
t.Fatalf("the row must be able to name who switched it off, got %q: %+v", sw.OverriddenBy, sw)
|
||||
}
|
||||
if sw.Override != model.RuleOverrideDisabled {
|
||||
t.Fatalf("override direction must be %q, got %q", model.RuleOverrideDisabled, sw.Override)
|
||||
}
|
||||
// Existing semantics untouched: neither rule is *unreachable* — both carry a
|
||||
// destination matcher, so the shadowing analysis has nothing to say.
|
||||
if sw.Unreachable || out.Rules[0].Unreachable {
|
||||
t.Fatalf("profile overrides must not be reported as shadowing: %+v", out.Rules)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRulesReachabilityProfileEnabled: the symmetric case — a rule switched OFF in
|
||||
// UCI that the active profile forces on is in force, and says who did it.
|
||||
func TestRulesReachabilityProfileEnabled(t *testing.T) {
|
||||
s := newTestServer(t)
|
||||
srv := httptest.NewServer(s.Handler())
|
||||
defer srv.Close()
|
||||
cookie := login(t, srv, s)
|
||||
|
||||
m := profileModel("ethernet-uplink")
|
||||
m.Rules[0].Enabled = false // desired state says off; the profile says otherwise
|
||||
|
||||
orig := reachConfigRead
|
||||
defer func() { reachConfigRead = orig }()
|
||||
reachConfigRead = func() (*model.Model, error) { return m, nil }
|
||||
|
||||
_, out := getReach(t, srv, cookie)
|
||||
ew := out.Rules[0]
|
||||
if !ew.EffectiveEnabled {
|
||||
t.Fatalf("the active profile force-enables this rule, so it IS in force: %+v", ew)
|
||||
}
|
||||
if ew.OverriddenBy != "ethernet-uplink" || ew.Override != model.RuleOverrideEnabled {
|
||||
t.Fatalf("want enabled-by-ethernet-uplink, got %q/%q: %+v", ew.OverriddenBy, ew.Override, ew)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRulesReachabilityNoProfileUnchanged: with no active profile the effective
|
||||
// flag is just the configured one and NOTHING claims an override — a page that
|
||||
// badges a profile name on a router that uses no profiles is the same lie in the
|
||||
// other direction.
|
||||
func TestRulesReachabilityNoProfileUnchanged(t *testing.T) {
|
||||
s := newTestServer(t)
|
||||
srv := httptest.NewServer(s.Handler())
|
||||
defer srv.Close()
|
||||
cookie := login(t, srv, s)
|
||||
|
||||
m := profileModel("ethernet-uplink")
|
||||
m.Profiles = nil // the pin names nothing that exists…
|
||||
m.Rules[1].Enabled = false // …and one rule is plainly switched off
|
||||
|
||||
orig := reachConfigRead
|
||||
defer func() { reachConfigRead = orig }()
|
||||
reachConfigRead = func() (*model.Model, error) { return m, nil }
|
||||
|
||||
_, out := getReach(t, srv, cookie)
|
||||
if !out.Rules[0].EffectiveEnabled || out.Rules[1].EffectiveEnabled {
|
||||
t.Fatalf("without a profile the effective flag is the configured one: %+v", out.Rules)
|
||||
}
|
||||
for _, r := range out.Rules {
|
||||
if r.OverriddenBy != "" || r.Override != "" {
|
||||
t.Fatalf("no active profile, so no row may name one: %+v", r)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestRulesReachabilityIndexMatchesConfig: the contract the panel zips on — one
|
||||
// verdict per rule, at the SAME position as GET /api/config's `Rules`, with
|
||||
// name/order echoed. Both endpoints are served from ONE canned model here, so this
|
||||
// checks the pair, not the fixture. A profile is active precisely because that is
|
||||
// when the two responses disagree about `Enabled` and the alignment is easiest to
|
||||
// get wrong.
|
||||
func TestRulesReachabilityIndexMatchesConfig(t *testing.T) {
|
||||
s := newTestServer(t)
|
||||
srv := httptest.NewServer(s.Handler())
|
||||
defer srv.Close()
|
||||
cookie := login(t, srv, s)
|
||||
|
||||
m := profileModel("ethernet-uplink")
|
||||
// A third rule, out of Order sequence, so a verdict list accidentally sorted by
|
||||
// Order (the panel's display order) would not line up.
|
||||
m.Rules = append(m.Rules, model.Rule{Name: "zz-first", Enabled: true, Order: 5, Target: "block", DstPort: "25"})
|
||||
|
||||
origReach, origCfg := reachConfigRead, configGetRead
|
||||
defer func() { reachConfigRead, configGetRead = origReach, origCfg }()
|
||||
reachConfigRead = func() (*model.Model, error) { return m, nil }
|
||||
configGetRead = func() (*model.Model, error) { return m, nil }
|
||||
|
||||
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/config", nil)
|
||||
req.AddCookie(cookie)
|
||||
resp, err := http.DefaultClient.Do(req)
|
||||
if err != nil {
|
||||
t.Fatalf("GET /api/config: %v", err)
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
var cfg struct {
|
||||
Rules []struct {
|
||||
Name string
|
||||
Order int
|
||||
Enabled bool
|
||||
}
|
||||
}
|
||||
if err := json.NewDecoder(resp.Body).Decode(&cfg); err != nil {
|
||||
t.Fatalf("decode /api/config: %v", err)
|
||||
}
|
||||
|
||||
_, out := getReach(t, srv, cookie)
|
||||
if len(out.Rules) != len(cfg.Rules) {
|
||||
t.Fatalf("one verdict per configured rule: %d verdicts for %d rules", len(out.Rules), len(cfg.Rules))
|
||||
}
|
||||
for i, v := range out.Rules {
|
||||
if v.Index != i {
|
||||
t.Fatalf("verdict %d carries index %d — the panel keys rows on it", i, v.Index)
|
||||
}
|
||||
if v.Name != cfg.Rules[i].Name || v.Order != cfg.Rules[i].Order {
|
||||
t.Fatalf("verdict %d is about %q/%d, /api/config's row %d is %q/%d",
|
||||
i, v.Name, v.Order, i, cfg.Rules[i].Name, cfg.Rules[i].Order)
|
||||
}
|
||||
}
|
||||
// And the whole point: /api/config still reports the DESIRED state (both
|
||||
// original rules on, as UCI has them) while the verdicts report the effective
|
||||
// one. If these ever agree, the desired-state contract has been broken.
|
||||
if !cfg.Rules[0].Enabled || !cfg.Rules[1].Enabled {
|
||||
t.Fatalf("/api/config must keep serving the desired state verbatim: %+v", cfg.Rules)
|
||||
}
|
||||
if !out.Rules[0].EffectiveEnabled || out.Rules[1].EffectiveEnabled {
|
||||
t.Fatalf("verdicts must report the effective state: %+v", out.Rules)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRulesReachabilityMethodAndAuth: it is a read endpoint behind the session
|
||||
// cookie, like every other /api route except /api/session.
|
||||
func TestRulesReachabilityMethodAndAuth(t *testing.T) {
|
||||
|
||||
@@ -75,6 +75,14 @@ func TestMalformedShareLinksAreRejectedNotPanicking(t *testing.T) {
|
||||
"ss://", "ss://@", "ss://notbase64@", "ss://" + base64.StdEncoding.EncodeToString([]byte("aes-128-gcm:pw")),
|
||||
"ss://" + base64.RawURLEncoding.EncodeToString([]byte("rot13:pw")) + "@h:443", // unknown cipher
|
||||
"wireguard://", "wg://k@host", "awg://@host:51820",
|
||||
"hysteria2://", "hy2://", "hysteria2://@", "hysteria2://pw@", "hy2://pw@:443",
|
||||
"hysteria2://host:443", // auth-less: every real server requires auth
|
||||
"hysteria2://pw@host:0", "hysteria2://pw@host:99999", "hysteria2://pw@host:notaport",
|
||||
"hysteria2://pw@host:-", "hysteria2://pw@host:,", "hysteria2://pw@host:0-0",
|
||||
"hysteria2://pw@[2001:db8::1", "hysteria2://pw@host:443?obfs=salamander",
|
||||
"tuic://", "tuic://@", "tuic://:@host:443", "tuic://uuid@host:443",
|
||||
"tuic://x:y@host:443", "tuic://" + strings.Repeat("a", 4096) + ":pw@host:443",
|
||||
"hysteria2://" + strings.Repeat("p", 4096) + "@host:443?" + strings.Repeat("k=v&", 2000),
|
||||
}
|
||||
for _, uri := range bad {
|
||||
func() {
|
||||
|
||||
+128
-1
@@ -67,6 +67,132 @@ func parsePort(s string) (uint16, error) {
|
||||
return uint16(n), nil
|
||||
}
|
||||
|
||||
// truthyParam reports whether a share-link boolean-ish query value means "yes".
|
||||
// The specs say "1"/"0", but the same flag is written "true", "yes" and "on" by
|
||||
// different clients, and a flag read as false when the link meant true (e.g.
|
||||
// insecure= on a self-signed node) is a silent misconfiguration.
|
||||
func truthyParam(s string) bool {
|
||||
switch strings.ToLower(strings.TrimSpace(s)) {
|
||||
case "1", "true", "yes", "on":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// foldedParams flattens query params into a lookup whose keys are lower-cased
|
||||
// with '-' folded to '_'. Clients disagree on the spelling of the SAME parameter
|
||||
// (pinSHA256/pinsha256, obfs-password/obfs_password, allow_insecure/allowInsecure),
|
||||
// and a parameter missed because of its spelling is a parameter silently dropped.
|
||||
// Only the first value of a repeated key is kept.
|
||||
func foldedParams(q url.Values) map[string]string {
|
||||
out := make(map[string]string, len(q))
|
||||
for k, vs := range q {
|
||||
if len(vs) == 0 {
|
||||
continue
|
||||
}
|
||||
key := strings.ReplaceAll(strings.ToLower(k), "-", "_")
|
||||
if _, seen := out[key]; !seen {
|
||||
out[key] = vs[0]
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// looksLikeUUID reports whether s is a UUID in one of the spellings sing-box's
|
||||
// uuid.FromString accepts: canonical 8-4-4-4-12 hex, bare 32-hex, or either of
|
||||
// those wrapped in braces / prefixed with "urn:uuid:".
|
||||
func looksLikeUUID(s string) bool {
|
||||
s = strings.ToLower(strings.TrimSpace(s))
|
||||
s = strings.TrimPrefix(s, "urn:uuid:")
|
||||
s = strings.TrimSuffix(strings.TrimPrefix(s, "{"), "}")
|
||||
switch len(s) {
|
||||
case 32:
|
||||
case 36:
|
||||
if s[8] != '-' || s[13] != '-' || s[18] != '-' || s[23] != '-' {
|
||||
return false
|
||||
}
|
||||
default:
|
||||
return false
|
||||
}
|
||||
digits := 0
|
||||
for i := 0; i < len(s); i++ {
|
||||
switch c := s[i]; {
|
||||
case c >= '0' && c <= '9', c >= 'a' && c <= 'f':
|
||||
digits++
|
||||
case c == '-':
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
return digits == 32
|
||||
}
|
||||
|
||||
// rewriteHopPorts folds a hysteria2 port-hopping port field ("host:443,5000-6000"
|
||||
// or "host:5000-6000") down to the first concrete port and returns the rewritten
|
||||
// URI plus the ORIGINAL spec (empty when the URI has no hop spec).
|
||||
//
|
||||
// It exists because such an authority is not a legal URI: net/url's port
|
||||
// validator accepts digits only, so url.Parse would reject the entire link and
|
||||
// the node would vanish with a parser error that names nothing the operator can
|
||||
// act on. Folding keeps the node — a hysteria2 server listening with port hopping
|
||||
// redirects the whole range to its single real listener, so any port in the set
|
||||
// reaches it — and the caller turns the returned spec into a warning.
|
||||
func rewriteHopPorts(uri string) (string, string) {
|
||||
i := strings.Index(uri, "://")
|
||||
if i < 0 {
|
||||
return uri, ""
|
||||
}
|
||||
start := i + 3
|
||||
end := len(uri)
|
||||
for j := start; j < len(uri); j++ {
|
||||
if c := uri[j]; c == '/' || c == '?' || c == '#' {
|
||||
end = j
|
||||
break
|
||||
}
|
||||
}
|
||||
auth := uri[start:end]
|
||||
hostStart := 0
|
||||
if at := strings.LastIndexByte(auth, '@'); at >= 0 {
|
||||
hostStart = at + 1
|
||||
}
|
||||
hp := auth[hostStart:]
|
||||
colon := -1
|
||||
if strings.HasPrefix(hp, "[") { // bracketed IPv6 literal
|
||||
if j := strings.IndexByte(hp, ']'); j >= 0 && j+1 < len(hp) && hp[j+1] == ':' {
|
||||
colon = j + 1
|
||||
}
|
||||
} else {
|
||||
colon = strings.LastIndexByte(hp, ':')
|
||||
}
|
||||
if colon < 0 {
|
||||
return uri, ""
|
||||
}
|
||||
spec := hp[colon+1:]
|
||||
if !strings.ContainsAny(spec, ",-") {
|
||||
return uri, ""
|
||||
}
|
||||
first := firstHopPort(spec)
|
||||
if first == "" {
|
||||
return uri, "" // let url.Parse report the malformed authority
|
||||
}
|
||||
return uri[:start] + auth[:hostStart] + hp[:colon+1] + first + uri[end:], spec
|
||||
}
|
||||
|
||||
// firstHopPort returns the first concrete port of a hop spec ("443,5000-6000" ->
|
||||
// "443"; "5000-6000" -> "5000"), or "" when none of the parts is a valid port.
|
||||
func firstHopPort(spec string) string {
|
||||
for _, part := range strings.Split(spec, ",") {
|
||||
part = strings.TrimSpace(part)
|
||||
if d := strings.IndexByte(part, '-'); d >= 0 {
|
||||
part = part[:d]
|
||||
}
|
||||
if _, err := parsePort(part); err == nil {
|
||||
return part
|
||||
}
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// atoiOr parses an int, returning def on failure.
|
||||
func atoiOr(s string, def int) int {
|
||||
if n, err := strconv.Atoi(strings.TrimSpace(s)); err == nil {
|
||||
@@ -255,7 +381,8 @@ func anyMatch(rs []*regexp.Regexp, s string) bool {
|
||||
|
||||
// schemePrefixes is the set of share-link schemes ParseShareLink understands.
|
||||
var schemePrefixes = []string{
|
||||
"vless://", "vmess://", "trojan://", "ss://", "wireguard://", "wg://", "awg://",
|
||||
"vless://", "vmess://", "trojan://", "ss://", "hysteria2://", "hy2://", "tuic://",
|
||||
"wireguard://", "wg://", "awg://",
|
||||
}
|
||||
|
||||
func hasScheme(l string) bool {
|
||||
|
||||
+259
-5
@@ -1,10 +1,10 @@
|
||||
package parse
|
||||
|
||||
// Share-link parsers for vless/vmess/trojan/ss (+ wireguard/wg/awg in
|
||||
// wireguard.go). Ported from v0.1 xrayctl/sharelink.go; the OUTPUT target is the
|
||||
// neutral *Proxy instead of an xray map. Handles modern transports (tcp/raw, ws,
|
||||
// grpc, xhttp, httpupgrade, http/h2, quic) and security none/tls/reality, plus
|
||||
// the base64 quirks and the vmess base64-JSON "ps" name.
|
||||
// Share-link parsers for vless/vmess/trojan/ss/hysteria2/tuic (+ wireguard/wg/awg
|
||||
// in wireguard.go). Ported from v0.1 xrayctl/sharelink.go; the OUTPUT target is
|
||||
// the neutral *Proxy instead of an xray map. Handles modern transports (tcp/raw,
|
||||
// ws, grpc, xhttp, httpupgrade, http/h2, quic) and security none/tls/reality,
|
||||
// plus the base64 quirks and the vmess base64-JSON "ps" name.
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
@@ -27,6 +27,11 @@ func ParseShareLink(uri string) (*Proxy, error) {
|
||||
return parseTrojan(uri)
|
||||
case strings.HasPrefix(uri, "ss://"):
|
||||
return parseSS(uri)
|
||||
case strings.HasPrefix(uri, "hysteria2://"),
|
||||
strings.HasPrefix(uri, "hy2://"):
|
||||
return parseHysteria2(uri)
|
||||
case strings.HasPrefix(uri, "tuic://"):
|
||||
return parseTUIC(uri)
|
||||
case strings.HasPrefix(uri, "wireguard://"),
|
||||
strings.HasPrefix(uri, "wg://"),
|
||||
strings.HasPrefix(uri, "awg://"):
|
||||
@@ -285,3 +290,252 @@ func parseSS(uri string) (*Proxy, error) {
|
||||
Password: password,
|
||||
}, nil
|
||||
}
|
||||
|
||||
// --- QUIC protocols: hysteria2 / tuic ---------------------------------------
|
||||
//
|
||||
// # Which spellings are supported, and why
|
||||
//
|
||||
// hysteria2 follows the OFFICIAL URI scheme
|
||||
// (https://v2.hysteria.network/docs/developers/URI-Scheme/):
|
||||
//
|
||||
// hysteria2://[auth@]hostname[:port]/?key=value&…#name
|
||||
//
|
||||
// `hy2://` is an equally official alias of the same grammar and is accepted
|
||||
// verbatim. The port is OPTIONAL and defaults to 443. The WHOLE userinfo is the
|
||||
// auth string, so an auth containing a raw ':' is stitched back together exactly
|
||||
// like the trojan password (a link that writes it percent-encoded arrives already
|
||||
// decoded by net/url and needs no stitching). Query keys read here: sni,
|
||||
// insecure, obfs, obfs-password, pinSHA256, ech, up, down, fastopen — plus the
|
||||
// two incompatible ways port hopping is written in the wild: the official one
|
||||
// puts the hop set in the port field itself (`host:443,5000-6000`), while
|
||||
// NekoBox/v2rayN keep a single port and add `mport=`/`ports=`. Both are handled.
|
||||
//
|
||||
// tuic has NO official URI scheme. The de-facto form — emitted by
|
||||
// v2rayN/NekoBox/NekoRay and understood by every TUIC v5 client — is:
|
||||
//
|
||||
// tuic://uuid:password@host:port/?key=value&…#name
|
||||
//
|
||||
// uuid and port are mandatory here. The TUIC v4 `token@host:port` form is
|
||||
// REFUSED rather than guessed at: sing-box implements v5 only, so treating a v4
|
||||
// token as a uuid would mint a node that can never authenticate. Query keys read
|
||||
// here: sni, alpn, congestion_control, udp_relay_mode, allow_insecure/insecure,
|
||||
// disable_sni, zero_rtt_handshake/reduce_rtt.
|
||||
//
|
||||
// # What happens to parameters the neutral Proxy cannot carry
|
||||
//
|
||||
// Neither parse.Proxy nor generate/outbound.go has a field for obfs, port
|
||||
// hopping, congestion control, and friends. Swallowing them silently is the one
|
||||
// thing this package must not do, so each such parameter is sorted into one of
|
||||
// two buckets:
|
||||
//
|
||||
// - REFUSE (return an error). Used when ignoring the parameter yields a node
|
||||
// that cannot work, or that works with LESS security than the link asked
|
||||
// for. The error is not a black hole: generate warns per node and skips it
|
||||
// (outbound.go), and `shaterd nodes` prints it verbatim as `parse_error`, so
|
||||
// the operator sees the node and the reason side by side. hysteria2 `obfs`
|
||||
// is the first kind — Salamander obfuscation rewrites every QUIC packet, so
|
||||
// a client that does not perform it never completes a handshake. hysteria2
|
||||
// `pinSHA256` is the second — the link pins one certificate (usually
|
||||
// alongside insecure=1); dropping the pin would leave us accepting ANY
|
||||
// certificate, which is not "slightly different", it is the MITM the pin
|
||||
// exists to stop.
|
||||
//
|
||||
// - FLAG (Proxy.Warnings). Used when the node still connects and still carries
|
||||
// the traffic, but some property of the link did not survive: performance
|
||||
// tuning (congestion control, Brutal up/down), an evasion enhancement (port
|
||||
// hopping, ECH), or a knob sing-box defaults differently. Refusing these
|
||||
// would delete most real-world tuic nodes over a throughput hint, which is
|
||||
// the opposite of honest. `shaterd nodes` reports them as `parse_warnings`.
|
||||
//
|
||||
// Both buckets are per-parameter and explicit: an UNRECOGNISED query key is not
|
||||
// flagged, because feed operators sprinkle tracking junk into links and warning
|
||||
// about it would drown the warnings that mean something.
|
||||
|
||||
// quicTLS builds the TLS block for a QUIC protocol from the query params.
|
||||
//
|
||||
// buildSecurity is reused (sni/host/alpn/insecure handling is identical and must
|
||||
// not drift), but it does NOT fit as-is on two counts, both handled here:
|
||||
//
|
||||
// - its `network` argument would synthesise a v2ray stream transport, which is
|
||||
// meaningless over QUIC — passing "" makes normalizeTransport return "tcp"
|
||||
// and no Transport is built;
|
||||
// - it copies `fp` into TLS.Fingerprint, which generate turns into a uTLS
|
||||
// client. uTLS cannot produce a QUIC TLS config (common/tls/utls_client.go
|
||||
// STDConfig returns "unsupported usage for uTLS"), and the failure surfaces
|
||||
// only when the node is DIALLED, i.e. as an unexplained dead node. The
|
||||
// fingerprint is therefore dropped here, loudly.
|
||||
//
|
||||
// It also widens the insecure flag: the official hysteria2 scheme documents
|
||||
// "1"/"0", but links in the wild spell it "true"/"yes"/"on" and a silently
|
||||
// ignored insecure flag turns a self-signed node into a permanently failing one.
|
||||
func quicTLS(proto string, q url.Values, server string) (*TLS, []string) {
|
||||
tls, _ := buildSecurity("", "tls", q, server)
|
||||
if tls == nil { // unreachable: security is fixed to "tls"
|
||||
tls = &TLS{Enabled: true, SNI: server}
|
||||
}
|
||||
var warns []string
|
||||
if tls.Fingerprint != "" {
|
||||
warns = append(warns, fmt.Sprintf(
|
||||
"%s: fp=%s dropped — a uTLS fingerprint cannot be applied to a QUIC handshake, and keeping it would make the node fail at dial time instead of here",
|
||||
proto, tls.Fingerprint))
|
||||
tls.Fingerprint = ""
|
||||
}
|
||||
if !tls.Insecure {
|
||||
tls.Insecure = truthyParam(q.Get("insecure")) ||
|
||||
truthyParam(q.Get("allow_insecure")) || truthyParam(q.Get("allowInsecure"))
|
||||
}
|
||||
return tls, warns
|
||||
}
|
||||
|
||||
func parseHysteria2(uri string) (*Proxy, error) {
|
||||
// Port hopping in the port field ("host:443,5000-6000") is not a valid URI
|
||||
// authority, so net/url would reject the whole link. Fold it to the first
|
||||
// concrete port BEFORE parsing and remember the original spec for the warning.
|
||||
rewritten, hopSpec := rewriteHopPorts(uri)
|
||||
u, err := url.Parse(rewritten)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("hysteria2: %w", err)
|
||||
}
|
||||
host := u.Hostname()
|
||||
if host == "" {
|
||||
return nil, fmt.Errorf("hysteria2: missing host")
|
||||
}
|
||||
port := uint16(443) // the scheme makes the port optional; 443 is its default
|
||||
if ps := u.Port(); ps != "" {
|
||||
if port, err = parsePort(ps); err != nil {
|
||||
return nil, fmt.Errorf("hysteria2: %w", err)
|
||||
}
|
||||
}
|
||||
// The whole userinfo is the auth string (same rule as the trojan password).
|
||||
auth := ""
|
||||
if u.User != nil {
|
||||
auth = u.User.Username()
|
||||
if pw, ok := u.User.Password(); ok {
|
||||
auth = auth + ":" + pw
|
||||
}
|
||||
}
|
||||
if auth == "" {
|
||||
// The grammar marks auth optional, but every deployed hysteria2 server
|
||||
// configures an auth block, so an authless link is a truncated one.
|
||||
// Accepting it would add a node that can only ever fail to authenticate.
|
||||
return nil, fmt.Errorf("hysteria2: missing auth")
|
||||
}
|
||||
|
||||
q := u.Query()
|
||||
p := foldedParams(q)
|
||||
|
||||
if obfs := strings.ToLower(p["obfs"]); obfs != "" && obfs != "none" {
|
||||
return nil, fmt.Errorf("hysteria2: obfs=%s is not supported — "+
|
||||
"obfuscation rewrites every QUIC packet, so a client that cannot perform it never completes a handshake; "+
|
||||
"the node is refused rather than added as one that can only fail", obfs)
|
||||
}
|
||||
if pin := p["pinsha256"]; pin != "" {
|
||||
return nil, fmt.Errorf("hysteria2: pinSHA256 is not supported — " +
|
||||
"the link pins one certificate and this build cannot honour the pin; " +
|
||||
"ignoring it would mean accepting ANY certificate instead, which is exactly what the pin exists to prevent")
|
||||
}
|
||||
|
||||
tls, warns := quicTLS("hysteria2", q, host)
|
||||
if hopSpec == "" {
|
||||
hopSpec = firstNonEmpty(p["mport"], p["ports"])
|
||||
}
|
||||
if hopSpec != "" {
|
||||
warns = append(warns, fmt.Sprintf(
|
||||
"hysteria2: port hopping %q is not carried into the engine — the node dials %d only. "+
|
||||
"Hopping is a censorship-evasion enhancement layered on top of the real listener "+
|
||||
"(the server listens on one port and redirects the range to it), so the node still works, but it is more fingerprintable",
|
||||
hopSpec, port))
|
||||
}
|
||||
if p["ech"] != "" {
|
||||
warns = append(warns, "hysteria2: ech= is not carried into the engine — "+
|
||||
"the TLS ClientHello, and with it the SNI, travels in the clear")
|
||||
}
|
||||
if p["up"] != "" || p["down"] != "" {
|
||||
warns = append(warns, fmt.Sprintf(
|
||||
"hysteria2: bandwidth hints (up=%q down=%q) are not carried into the engine — "+
|
||||
"the Brutal congestion controller stays off and BBR is used instead; throughput may differ from what the provider sized",
|
||||
p["up"], p["down"]))
|
||||
}
|
||||
if truthyParam(p["fastopen"]) {
|
||||
warns = append(warns, "hysteria2: fastopen is not carried into the engine — "+
|
||||
"the first request waits for the tunnel to be established (latency only, no functional change)")
|
||||
}
|
||||
|
||||
return &Proxy{
|
||||
Name: fragmentName(uri),
|
||||
Protocol: "hysteria2",
|
||||
Server: host,
|
||||
Port: port,
|
||||
Password: auth,
|
||||
TLS: tls,
|
||||
Warnings: warns,
|
||||
}, nil
|
||||
}
|
||||
|
||||
func parseTUIC(uri string) (*Proxy, error) {
|
||||
u, err := url.Parse(uri)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("tuic: %w", err)
|
||||
}
|
||||
host := u.Hostname()
|
||||
if host == "" {
|
||||
return nil, fmt.Errorf("tuic: missing host")
|
||||
}
|
||||
port, err := parsePort(u.Port())
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("tuic: %w", err)
|
||||
}
|
||||
if u.User == nil {
|
||||
return nil, fmt.Errorf("tuic: missing uuid:password")
|
||||
}
|
||||
uuid := u.User.Username()
|
||||
pw, hasPW := u.User.Password()
|
||||
if !hasPW {
|
||||
// A single userinfo component is the TUIC v4 "token" form. sing-box
|
||||
// implements v5 only; passing the token off as a uuid would produce a
|
||||
// node that authenticates against nothing.
|
||||
return nil, fmt.Errorf("tuic: userinfo %q is not uuid:password — "+
|
||||
"only the TUIC v5 form is supported (the v4 token form cannot be dialled by this engine)", uuid)
|
||||
}
|
||||
if !looksLikeUUID(uuid) {
|
||||
// Not cosmetic: the TUIC outbound rejects a non-UUID at CONSTRUCTION,
|
||||
// and that failure aborts the whole generated config — every other node
|
||||
// with it. Refusing here costs one node instead.
|
||||
return nil, fmt.Errorf("tuic: %q is not a uuid", uuid)
|
||||
}
|
||||
|
||||
q := u.Query()
|
||||
p := foldedParams(q)
|
||||
tls, warns := quicTLS("tuic", q, host)
|
||||
|
||||
// sing-box's own defaults are cubic + native; only a DIFFERENT value in the
|
||||
// link is a lost instruction worth reporting.
|
||||
if cc := strings.ToLower(p["congestion_control"]); cc != "" && cc != "cubic" {
|
||||
warns = append(warns, fmt.Sprintf(
|
||||
"tuic: congestion_control=%s is not carried into the engine — the node dials with sing-box's default (cubic); "+
|
||||
"this is a throughput difference, not a connectivity one", cc))
|
||||
}
|
||||
if m := strings.ToLower(p["udp_relay_mode"]); m != "" && m != "native" {
|
||||
warns = append(warns, fmt.Sprintf(
|
||||
"tuic: udp_relay_mode=%s is not carried into the engine — UDP is relayed in native mode (QUIC datagrams)", m))
|
||||
}
|
||||
if truthyParam(p["disable_sni"]) {
|
||||
warns = append(warns, fmt.Sprintf(
|
||||
"tuic: disable_sni is not carried into the engine — the SNI %q is sent in the clear, so the destination is visible to a middlebox", tls.SNI))
|
||||
}
|
||||
if truthyParam(p["zero_rtt_handshake"]) || truthyParam(p["reduce_rtt"]) {
|
||||
warns = append(warns, "tuic: zero_rtt_handshake is not carried into the engine — "+
|
||||
"a full handshake is performed (one extra round trip on connect)")
|
||||
}
|
||||
|
||||
return &Proxy{
|
||||
Name: fragmentName(uri),
|
||||
Protocol: "tuic",
|
||||
Server: host,
|
||||
Port: port,
|
||||
UUID: uuid,
|
||||
Password: pw,
|
||||
TLS: tls,
|
||||
Warnings: warns,
|
||||
}, nil
|
||||
}
|
||||
|
||||
@@ -2,6 +2,7 @@ package parse
|
||||
|
||||
import (
|
||||
"encoding/base64"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
@@ -157,7 +158,328 @@ func TestFingerprintStable(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestUnsupportedScheme(t *testing.T) {
|
||||
if _, err := ParseShareLink("hysteria2://x@h:1"); err == nil {
|
||||
t.Fatal("expected error on unsupported scheme")
|
||||
// hysteria v1 and socks are genuinely not implemented (registry_quic.go
|
||||
// registers hysteria2/tuic only), so they are the honest examples here.
|
||||
for _, uri := range []string{"hysteria://x@h:1", "socks5://u:p@h:1080", "juicity://x@h:1"} {
|
||||
if _, err := ParseShareLink(uri); err == nil {
|
||||
t.Fatalf("expected error on unsupported scheme %q", uri)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- hysteria2 --------------------------------------------------------------
|
||||
|
||||
func TestParseHysteria2(t *testing.T) {
|
||||
uri := "hysteria2://s3cret@example.com:8443/?sni=cdn.example.com&insecure=1&alpn=h3#hy2-1"
|
||||
p, err := ParseShareLink(uri)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.Protocol != "hysteria2" {
|
||||
t.Fatalf("protocol = %v", p.Protocol)
|
||||
}
|
||||
if p.Server != "example.com" || p.Port != 8443 {
|
||||
t.Fatalf("server/port = %v:%v", p.Server, p.Port)
|
||||
}
|
||||
if p.Password != "s3cret" {
|
||||
t.Fatalf("password = %q", p.Password)
|
||||
}
|
||||
if p.UUID != "" || p.Method != "" || p.Transport != nil {
|
||||
t.Fatalf("unexpected non-hysteria2 fields set: %+v", p)
|
||||
}
|
||||
if p.TLS == nil || !p.TLS.Enabled {
|
||||
t.Fatalf("tls = %+v", p.TLS)
|
||||
}
|
||||
if p.TLS.SNI != "cdn.example.com" || !p.TLS.Insecure {
|
||||
t.Fatalf("tls = %+v", p.TLS)
|
||||
}
|
||||
if len(p.TLS.ALPN) != 1 || p.TLS.ALPN[0] != "h3" {
|
||||
t.Fatalf("alpn = %v", p.TLS.ALPN)
|
||||
}
|
||||
if p.TLS.Reality != nil {
|
||||
t.Fatalf("reality must never be built for a QUIC link: %+v", p.TLS.Reality)
|
||||
}
|
||||
if p.Name != "hy2-1" {
|
||||
t.Fatalf("name = %v", p.Name)
|
||||
}
|
||||
if len(p.Warnings) != 0 {
|
||||
t.Fatalf("clean link warned: %v", p.Warnings)
|
||||
}
|
||||
}
|
||||
|
||||
// hy2:// is an official alias, the port defaults to 443, and SNI falls back to
|
||||
// the server address (so generate never has to synthesise one).
|
||||
func TestParseHysteria2AliasDefaultPort(t *testing.T) {
|
||||
p, err := ParseShareLink("hy2://pw@h.example.com#alias")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.Protocol != "hysteria2" || p.Port != 443 {
|
||||
t.Fatalf("hy2 alias/default port: %+v", p)
|
||||
}
|
||||
if p.TLS == nil || p.TLS.SNI != "h.example.com" {
|
||||
t.Fatalf("sni fallback = %+v", p.TLS)
|
||||
}
|
||||
}
|
||||
|
||||
// The WHOLE userinfo is the auth string: a raw ':' must not truncate it, and a
|
||||
// percent-encoded one arrives already decoded.
|
||||
func TestParseHysteria2AuthForms(t *testing.T) {
|
||||
cases := map[string]string{
|
||||
"hysteria2://user:pass@h.example.com:443#a": "user:pass",
|
||||
"hysteria2://p%40ss%23w%3Ax@h.example.com:443": "p@ss#w:x",
|
||||
"hysteria2://plain@h.example.com:443": "plain",
|
||||
}
|
||||
for uri, want := range cases {
|
||||
p, err := ParseShareLink(uri)
|
||||
if err != nil {
|
||||
t.Fatalf("%s: %v", uri, err)
|
||||
}
|
||||
if p.Password != want {
|
||||
t.Errorf("%s: auth = %q, want %q", uri, p.Password, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseHysteria2IPv6AndNoFragment(t *testing.T) {
|
||||
p, err := ParseShareLink("hysteria2://pw@[2001:db8::1]:443?sni=v6.example.com")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.Server != "2001:db8::1" || p.Port != 443 {
|
||||
t.Fatalf("IPv6 host/port mangled: %v/%v", p.Server, p.Port)
|
||||
}
|
||||
if p.Name != "" {
|
||||
t.Fatalf("name from a fragment-less link = %q, want empty", p.Name)
|
||||
}
|
||||
}
|
||||
|
||||
// obfs and pinSHA256 are REFUSED, not flagged: without the first the handshake
|
||||
// can never complete, and ignoring the second silently downgrades the link's
|
||||
// certificate pinning to "accept anything".
|
||||
func TestParseHysteria2RefusesUndeliverableSecurity(t *testing.T) {
|
||||
cases := map[string]string{
|
||||
"hysteria2://pw@h.example.com:443?obfs=salamander&obfs-password=x": "obfs",
|
||||
"hysteria2://pw@h.example.com:443?obfs=gecko": "obfs",
|
||||
"hysteria2://pw@h.example.com:443?pinSHA256=aa:bb:cc": "pinSHA256",
|
||||
"hysteria2://pw@h.example.com:443?pinsha256=aabbcc": "pinSHA256",
|
||||
}
|
||||
for uri, want := range cases {
|
||||
p, err := ParseShareLink(uri)
|
||||
if err == nil {
|
||||
t.Errorf("%s: accepted (%+v), want refusal", uri, p)
|
||||
continue
|
||||
}
|
||||
if !strings.Contains(err.Error(), want) {
|
||||
t.Errorf("%s: error %q does not name %q", uri, err, want)
|
||||
}
|
||||
}
|
||||
// obfs=none is the explicit "no obfuscation" spelling and must pass.
|
||||
if _, err := ParseShareLink("hysteria2://pw@h.example.com:443?obfs=none"); err != nil {
|
||||
t.Errorf("obfs=none refused: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// Port hopping in either spelling keeps the node (the base port is the server's
|
||||
// real listener) but must leave a warning behind.
|
||||
func TestParseHysteria2PortHopping(t *testing.T) {
|
||||
cases := []struct {
|
||||
uri string
|
||||
port uint16
|
||||
}{
|
||||
{"hysteria2://pw@h.example.com:443,5000-6000/?sni=h.example.com#hop", 443},
|
||||
{"hysteria2://pw@h.example.com:5000-6000#hop", 5000},
|
||||
{"hysteria2://pw@h.example.com:443?mport=10000-20000#hop", 443},
|
||||
{"hysteria2://pw@h.example.com:443?ports=10000-20000#hop", 443},
|
||||
{"hysteria2://pw@[2001:db8::1]:443,5000-6000#hop", 443},
|
||||
}
|
||||
for _, c := range cases {
|
||||
p, err := ParseShareLink(c.uri)
|
||||
if err != nil {
|
||||
t.Errorf("%s: %v", c.uri, err)
|
||||
continue
|
||||
}
|
||||
if p.Port != c.port {
|
||||
t.Errorf("%s: port = %d, want %d", c.uri, p.Port, c.port)
|
||||
}
|
||||
if !warnedAbout(p, "port hopping") {
|
||||
t.Errorf("%s: port hopping dropped silently (warnings: %v)", c.uri, p.Warnings)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A uTLS fingerprint cannot be used for a QUIC handshake; keeping it would make
|
||||
// the node fail at dial time with an unrelated-looking message.
|
||||
func TestParseHysteria2DropsUTLSFingerprint(t *testing.T) {
|
||||
p, err := ParseShareLink("hysteria2://pw@h.example.com:443?fp=chrome#fp")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.TLS.Fingerprint != "" {
|
||||
t.Fatalf("fingerprint kept for QUIC: %q", p.TLS.Fingerprint)
|
||||
}
|
||||
if !warnedAbout(p, "fp=chrome") {
|
||||
t.Fatalf("dropped fingerprint without a warning: %v", p.Warnings)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseHysteria2OtherLossesWarn(t *testing.T) {
|
||||
p, err := ParseShareLink("hysteria2://pw@h.example.com:443?up=100&down=500&fastopen=1&ech=AEX+DQ")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, want := range []string{"bandwidth hints", "fastopen", "ech"} {
|
||||
if !warnedAbout(p, want) {
|
||||
t.Errorf("no warning mentioning %q: %v", want, p.Warnings)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// insecure= is spelled several ways in the wild; reading it as false would turn
|
||||
// a working self-signed node into a permanently failing one.
|
||||
func TestParseHysteria2InsecureSpellings(t *testing.T) {
|
||||
for _, q := range []string{"insecure=1", "insecure=true", "insecure=yes", "allow_insecure=true", "allowInsecure=1"} {
|
||||
p, err := ParseShareLink("hysteria2://pw@h.example.com:443?" + q)
|
||||
if err != nil {
|
||||
t.Fatalf("%s: %v", q, err)
|
||||
}
|
||||
if !p.TLS.Insecure {
|
||||
t.Errorf("%s: insecure not honoured", q)
|
||||
}
|
||||
}
|
||||
p, _ := ParseShareLink("hysteria2://pw@h.example.com:443?insecure=0")
|
||||
if p.TLS.Insecure {
|
||||
t.Error("insecure=0 read as true")
|
||||
}
|
||||
}
|
||||
|
||||
// --- tuic -------------------------------------------------------------------
|
||||
|
||||
func TestParseTUIC(t *testing.T) {
|
||||
uri := "tuic://22222222-2222-2222-2222-222222222222:s3cret@example.com:443/?sni=example.com&alpn=h3&congestion_control=cubic&udp_relay_mode=native#tuic-1"
|
||||
p, err := ParseShareLink(uri)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.Protocol != "tuic" {
|
||||
t.Fatalf("protocol = %v", p.Protocol)
|
||||
}
|
||||
if p.Server != "example.com" || p.Port != 443 {
|
||||
t.Fatalf("server/port = %v:%v", p.Server, p.Port)
|
||||
}
|
||||
if p.UUID != "22222222-2222-2222-2222-222222222222" {
|
||||
t.Fatalf("uuid = %q", p.UUID)
|
||||
}
|
||||
if p.Password != "s3cret" {
|
||||
t.Fatalf("password = %q", p.Password)
|
||||
}
|
||||
if p.TLS == nil || p.TLS.SNI != "example.com" || p.TLS.Insecure {
|
||||
t.Fatalf("tls = %+v", p.TLS)
|
||||
}
|
||||
if len(p.TLS.ALPN) != 1 || p.TLS.ALPN[0] != "h3" {
|
||||
t.Fatalf("alpn = %v", p.TLS.ALPN)
|
||||
}
|
||||
if p.Transport != nil {
|
||||
t.Fatalf("a QUIC link must not build a stream transport: %+v", p.Transport)
|
||||
}
|
||||
if p.Name != "tuic-1" {
|
||||
t.Fatalf("name = %v", p.Name)
|
||||
}
|
||||
// congestion_control/udp_relay_mode equal to sing-box's own defaults are not
|
||||
// a loss, so they must not produce noise.
|
||||
if len(p.Warnings) != 0 {
|
||||
t.Fatalf("default-valued knobs warned: %v", p.Warnings)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseTUICPercentEncodedPasswordAndIPv6(t *testing.T) {
|
||||
p, err := ParseShareLink("tuic://33333333-3333-3333-3333-333333333333:p%40ss%3Aword@[2001:db8::2]:8443?allow_insecure=1")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.Server != "2001:db8::2" || p.Port != 8443 {
|
||||
t.Fatalf("IPv6 host/port mangled: %v/%v", p.Server, p.Port)
|
||||
}
|
||||
if p.Password != "p@ss:word" {
|
||||
t.Fatalf("password = %q", p.Password)
|
||||
}
|
||||
if !p.TLS.Insecure {
|
||||
t.Fatalf("allow_insecure not honoured: %+v", p.TLS)
|
||||
}
|
||||
if p.Name != "" {
|
||||
t.Fatalf("name from a fragment-less link = %q, want empty", p.Name)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseTUICWarnsOnUndeliverableKnobs(t *testing.T) {
|
||||
uri := "tuic://44444444-4444-4444-4444-444444444444:pw@h.example.com:443?congestion_control=bbr&udp_relay_mode=quic&disable_sni=1&zero_rtt_handshake=1&fp=chrome"
|
||||
p, err := ParseShareLink(uri)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, want := range []string{"congestion_control=bbr", "udp_relay_mode=quic", "disable_sni", "zero_rtt_handshake", "fp=chrome"} {
|
||||
if !warnedAbout(p, want) {
|
||||
t.Errorf("no warning mentioning %q: %v", want, p.Warnings)
|
||||
}
|
||||
}
|
||||
if p.TLS.Fingerprint != "" {
|
||||
t.Errorf("fingerprint kept for QUIC: %q", p.TLS.Fingerprint)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseTUICRejects(t *testing.T) {
|
||||
cases := map[string]string{
|
||||
// v4 token form: sing-box speaks v5 only.
|
||||
"tuic://sometoken@h.example.com:443": "uuid:password",
|
||||
"tuic://not-a-uuid:pw@h.example.com:443": "not a uuid",
|
||||
"tuic://44444444-4444-4444-4444-444444444444:pw@": "host",
|
||||
"tuic://44444444-4444-4444-4444-444444444444:pw@h.example.com": "port",
|
||||
"tuic://44444444-4444-4444-4444-444444444444:pw@h.example.com:notaport": "port",
|
||||
"tuic://44444444-4444-4444-4444-444444444444:pw@h.example.com:0": "port",
|
||||
"tuic://44444444-4444-4444-4444-444444444444:pw@h.example.com:99999": "port",
|
||||
"tuic://h.example.com:443": "uuid",
|
||||
}
|
||||
for uri, want := range cases {
|
||||
p, err := ParseShareLink(uri)
|
||||
if err == nil {
|
||||
t.Errorf("%s: accepted (%+v), want error", uri, p)
|
||||
continue
|
||||
}
|
||||
if !strings.Contains(err.Error(), want) {
|
||||
t.Errorf("%s: error %q does not mention %q", uri, err, want)
|
||||
}
|
||||
}
|
||||
// The bare 32-hex uuid spelling is legal (sing-box's uuid.FromString takes it).
|
||||
if _, err := ParseShareLink("tuic://44444444444444444444444444444444:pw@h.example.com:443"); err != nil {
|
||||
t.Errorf("bare 32-hex uuid refused: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// A subscription body is split by scheme prefix before anything is parsed, so a
|
||||
// missing prefix drops the node one layer EARLIER than ParseShareLink — silently.
|
||||
func TestSubscriptionBodyKeepsQUICLinks(t *testing.T) {
|
||||
body := "hysteria2://pw@a.example.com:443#hy2\n" +
|
||||
"hy2://pw@b.example.com:443#hy2-alias\n" +
|
||||
"tuic://55555555-5555-5555-5555-555555555555:pw@c.example.com:443#tuic\n"
|
||||
links := ParseSubscriptionBody([]byte(body))
|
||||
if len(links) != 3 {
|
||||
t.Fatalf("ParseSubscriptionBody kept %d of 3 QUIC links: %v", len(links), links)
|
||||
}
|
||||
ps, err := ParseSubBody([]byte(body))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(ps) != 3 {
|
||||
t.Fatalf("ParseSubBody produced %d of 3 proxies: %v", len(ps), names(ps))
|
||||
}
|
||||
}
|
||||
|
||||
// warnedAbout reports whether any Proxy warning mentions substr.
|
||||
func warnedAbout(p *Proxy, substr string) bool {
|
||||
for _, w := range p.Warnings {
|
||||
if strings.Contains(w, substr) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
@@ -43,6 +43,24 @@ type Proxy struct {
|
||||
// wireguard / AmneziaWG (nil unless Protocol == "wireguard")
|
||||
WG *WGConfig
|
||||
|
||||
// Warnings records, in human-readable form, everything the share-link asked
|
||||
// for that this Proxy CANNOT carry to the generate stage — hysteria2 port
|
||||
// hopping, tuic congestion control, a uTLS fingerprint on a QUIC link, and so
|
||||
// on. It is empty for a link that survived its translation intact.
|
||||
//
|
||||
// It exists because the alternative was to swallow those parameters: the node
|
||||
// would then connect, but not the way the link describes it, and nothing
|
||||
// anywhere would say so. A parameter whose loss makes the node UNUSABLE or
|
||||
// LESS SECURE is not put here — that is an error from ParseShareLink, so the
|
||||
// node is skipped with a reason instead. This field is only for losses the
|
||||
// node survives.
|
||||
//
|
||||
// Consumers: `shaterd nodes` prints it as `parse_warnings`. It is deliberately
|
||||
// NOT emitted per-apply by the generate stage: a 300-node subscription would
|
||||
// repeat the same lines on every apply and train the operator to ignore the
|
||||
// warning channel that also carries kill-switch and chain failures.
|
||||
Warnings []string
|
||||
|
||||
// Passthrough for the generate stage; the parser leaves these at zero. In the
|
||||
// v0.2 model, per-node multiplexing and sockopt/mark live on model.Node and are
|
||||
// combined with this Proxy at generate time — these fields exist so the contract
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user