The boot armor never armed on the router it shipped to. procd runs the
K-links on the way down with the action `shutdown`, and stop_service
classified actions with an OPEN default:
case $action in restart|reload) keep;; *) DISARM;; esac
`shutdown` matched nobody, fell into `*`, and deleted the arm token. The
mechanism erased itself at exactly the transition it exists for, so every
boot found nothing to load. Measured on the live router, one minute apart
across a reboot:
13:28 /etc/shater/boot.nft present
---- reboot
18s at_S22: NO_TABLE armor_file=NO_FILE
It did not fail every time, which is worse than failing always: on the way
down `rm` from this script raced a `SaveBootArmor` driven by the ifdown
hotplug storm, and whichever landed second won. Two reboots on the same box
an hour apart gave opposite outcomes.
Both lists are now positive and CLOSED. Only `stop` disarms; only
`restart`/`reload` hand off. An action nobody thought of changes nothing,
so the default now fails toward a boot that arms when it need not have --
recoverable in the second before the daemon applies, and still gated by
shater-armor's four state refusals. The old default failed toward the
plaintext window the feature was built to close.
Also closed, found while proving the above:
* Every restart left the LAN in the clear for 80-90ms. The exit path was
`Teardown(); armOnExit()`, and TeardownNft DELETES the table -- two nft
transactions with no `inet shater` between them, leaving fw4's
`lan -> wan ACCEPT` as the only policy. Every restart, every LuCI Save
& Apply. TeardownExiting arms first under the apply lock and skips the
delete iff a plane actually went in; RenderHoldNft is one `nft -f` that
REPLACES the table, so the kernel never observes its absence.
35k-sample instrument: 7 and 6 no-table hits before, 0 across three
runs after.
* SaveBootArmor fsynced the payload but not the directory, so a power cut
could lose the rename that publishes it -- a boot with no armor and no
error anywhere.
`stop` now also reads rc.d state, so a package transaction that stops the
service is not mistaken for a person switching it off. This one does not
reproduce on apk (it runs no pre-upgrade script and never calls prerm on an
upgrade; verified with apk adbdump and 245k samples across a real reinstall)
-- it is one returning opkg lane away from being live, and the removal case
is now stated rather than implicit.
Both new tests are mutation-checked: reverting the predicate fails naming
`shutdown`; reverting the teardown fails with `did [arm delete], want [arm]`.
initscript_test.go sources the SHIPPED shell and calls the real predicates
with every action procd uses -- a comment claiming `shutdown` was handled is
what shipped last time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
282 lines
12 KiB
Go
282 lines
12 KiB
Go
// Keeping the fail-closed plane alive across the moments the daemon is not.
|
|
//
|
|
// The daemon owns the `inet shater` table, which means the table exists exactly
|
|
// while the daemon does. Three of those moments are not covered by anything else,
|
|
// and all three are the same defect wearing different clothes: the protection is
|
|
// an in-process thing, and the process is not always there.
|
|
//
|
|
// BOOT /etc/init.d/shater is START=99. fw4 loaded `lan -> wan ACCEPT` at 19
|
|
// and netifd brought the LAN up at 20; the clients that reconnect in
|
|
// between are unprotected until the daemon has been decompressed off
|
|
// flash, waited out any predecessor, migrated UCI and applied.
|
|
// RESTART SIGTERM ran an unconditional Teardown — kill_switch was not so much
|
|
// as consulted — and the successor cannot apply until the init's
|
|
// shater_wait_stopped loop, `shaterd migrate` and engine start have all
|
|
// finished. `reload_service` is stop+start, and so is every package
|
|
// upgrade, so this ran on a routine `Save & Apply`.
|
|
// NO CONFIG model.ReadUCI failing left the arming call unreached: it sat in the
|
|
// else-branch of the successful read. Nothing recovered from it either
|
|
// — Reconcile returns before any plane work, and the cron watchdog sees
|
|
// a live pidof and its own `uci -q get` fails the same way.
|
|
//
|
|
// The answer to all three is one artifact: netplane's BOOT ARMOR, a persisted copy
|
|
// of the fail-closed holding plane (netplane/armor.go). This file is the daemon's
|
|
// half — it keeps that copy honest, and it reinstates it in the two cases the
|
|
// daemon is the only one who can.
|
|
//
|
|
// Everything here is deliberately conservative in ONE direction: it never installs
|
|
// a plane the operator did not ask for. `globals.enabled=0` or `kill_switch=open`
|
|
// removes the armor and installs nothing, and a deliberate `/etc/init.d/shater
|
|
// stop` is a handoff-free exit that leaves nothing behind. Fail-closed is a
|
|
// policy, and a kill switch that outlives its own off switch is not one.
|
|
package main
|
|
|
|
import (
|
|
"os"
|
|
|
|
"github.com/sagernet/sing-box/log"
|
|
"github.com/sagernet/sing-box/shater/model"
|
|
"github.com/sagernet/sing-box/shater/netplane"
|
|
)
|
|
|
|
// restartHandoffPath is raised by /etc/init.d/shater around a restart/reload and
|
|
// cleared by its start (and by a real stop). Its presence at SIGTERM means "this
|
|
// daemon is being REPLACED", as opposed to "this daemon is being switched off".
|
|
//
|
|
// tmpfs on purpose: a marker that survived a power cut would make the first boot
|
|
// after it look like a restart.
|
|
//
|
|
// A var, not a const, only so tests can point it at a temp dir.
|
|
var restartHandoffPath = "/var/run/shater.restarting"
|
|
|
|
// restartHandoffPending reports whether the init script announced a restart.
|
|
//
|
|
// Absent is read as "a real stop", which is the SAFE direction to be wrong in: it
|
|
// degrades to exactly the behaviour that shipped before this file existed (full
|
|
// teardown), whereas the other default would leave a stopped router blocked.
|
|
func restartHandoffPending() bool {
|
|
_, err := os.Stat(restartHandoffPath)
|
|
return err == nil
|
|
}
|
|
|
|
// armorPlan is what to do about the fail-closed plane at a decision point.
|
|
type armorPlan int
|
|
|
|
const (
|
|
// armorNothing: install nothing. Either the operator does not want a plane
|
|
// (disabled / kill_switch=open), or there is nothing to install from.
|
|
armorNothing armorPlan = iota
|
|
// armorRender: build the holding plane from the model we just read. Preferred
|
|
// whenever a model is readable — it reflects the CURRENT interface set, where a
|
|
// snapshot may predate an interface rename.
|
|
armorRender
|
|
// armorSnapshot: reinstate the persisted boot armor. The only option when the
|
|
// config cannot be read, which is precisely when it is needed.
|
|
armorSnapshot
|
|
)
|
|
|
|
// armorWanted reports whether m asks for a fail-closed plane at all: the stack is
|
|
// enabled AND the kill switch is closed. Both halves are the operator's explicit
|
|
// choice and neither may be second-guessed — `kill_switch=open` is a documented
|
|
// decision to let traffic through when the engine is down, not an oversight.
|
|
func armorWanted(m *model.Model) bool {
|
|
return m != nil && m.Globals.Enabled && netplane.KillSwitchClosed(m.Globals)
|
|
}
|
|
|
|
// planStartupArmor decides what to install when the daemon starts and could NOT
|
|
// read its config.
|
|
//
|
|
// It is only ever consulted on the read-failure path: with a readable model the
|
|
// applier's own ArmHold does this job (and does it better — it holds the apply
|
|
// lock while it works). With no model there is nothing to render from, so the
|
|
// persisted snapshot is the entire answer; with no snapshot either, nothing is
|
|
// installed, because "this router has never applied an enabled, fail-closed
|
|
// config" is then the most likely truth and blacking out a LAN on a guess is not
|
|
// a recovery.
|
|
func planStartupArmor(readErr error, snapshot bool) armorPlan {
|
|
if readErr == nil {
|
|
return armorNothing
|
|
}
|
|
if snapshot {
|
|
return armorSnapshot
|
|
}
|
|
return armorNothing
|
|
}
|
|
|
|
// planExitArmor decides what the daemon leaves behind when it is asked to exit.
|
|
//
|
|
// The handoff flag is the whole distinction the old code was missing. A restart,
|
|
// a reload and a package upgrade all reach this point, and in all three the
|
|
// operator has not asked for protection to end — only for this process to be
|
|
// replaced. A `stop` has asked for exactly that, and must be obeyed: it is the
|
|
// operator's escape hatch, and a kill switch that cannot be switched off is a
|
|
// brick.
|
|
func planExitArmor(handoff bool, m *model.Model, readErr error, snapshot bool) armorPlan {
|
|
if !handoff {
|
|
return armorNothing
|
|
}
|
|
if readErr != nil {
|
|
// Being replaced with an unreadable config: the snapshot is the last thing
|
|
// this router is known to have wanted, and it is still the honest answer.
|
|
if snapshot {
|
|
return armorSnapshot
|
|
}
|
|
return armorNothing
|
|
}
|
|
if !armorWanted(m) {
|
|
return armorNothing
|
|
}
|
|
return armorRender
|
|
}
|
|
|
|
// refreshBootArmor keeps the persisted holding plane in step with the desired
|
|
// state. Called after every successful UCI read, so the snapshot on flash always
|
|
// describes the config the router is actually running.
|
|
//
|
|
// Writing is content-gated inside netplane.SaveBootArmor (this runs once a minute
|
|
// under cron; rewriting an identical file that often is how flash dies), and the
|
|
// REMOVE half matters just as much as the write: turning the stack off, or opening
|
|
// the kill switch, has to disarm the next boot too, or the operator's change would
|
|
// silently come back after a power cut.
|
|
func refreshBootArmor(m *model.Model, logger log.ContextLogger) {
|
|
if !armorWanted(m) {
|
|
if netplane.BootArmorPresent() {
|
|
if err := netplane.RemoveBootArmor(); err != nil {
|
|
logger.Warn("boot armor: could not remove ", netplane.BootArmorPath, ": ", err)
|
|
} else {
|
|
logger.Info("boot armor removed (the stack is disabled or the kill switch is open): ",
|
|
"the LAN is no longer blocked at boot before the daemon starts")
|
|
}
|
|
}
|
|
return
|
|
}
|
|
ruleset, err := netplane.RenderHoldNft(m)
|
|
if err != nil {
|
|
// A transient render failure must not disarm: a stale fail-closed plane is
|
|
// recoverable (the daemon replaces it seconds into the next boot), an absent
|
|
// one is a leak.
|
|
logger.Warn("boot armor: could not render the fail-closed plane: ", err)
|
|
return
|
|
}
|
|
if ruleset == "" {
|
|
// No divert devices at all — this config intercepts nothing, so there is
|
|
// nothing for a boot-time plane to protect. Blocking the LAN at boot on
|
|
// behalf of a config that does not touch it would be a pure outage.
|
|
if netplane.BootArmorPresent() {
|
|
if rerr := netplane.RemoveBootArmor(); rerr != nil {
|
|
logger.Warn("boot armor: could not remove ", netplane.BootArmorPath, ": ", rerr)
|
|
}
|
|
}
|
|
return
|
|
}
|
|
changed, serr := netplane.SaveBootArmor(ruleset)
|
|
switch {
|
|
case serr != nil:
|
|
logger.Warn("boot armor: could not write ", netplane.BootArmorPath, ": ", serr,
|
|
" — the LAN will be unprotected between boot and this daemon's first apply")
|
|
case changed:
|
|
logger.Info("boot armor updated (", netplane.BootArmorPath,
|
|
"): the LAN is fail-closed from early boot until the engine is up")
|
|
}
|
|
}
|
|
|
|
// armFromSnapshot reinstates the persisted holding plane and reports whether a
|
|
// plane is now standing. why is a short phrase for the log ("config is
|
|
// unreadable", "restart handoff").
|
|
func armFromSnapshot(why string, logger log.ContextLogger) bool {
|
|
loaded, err := netplane.LoadBootArmor()
|
|
switch {
|
|
case err != nil:
|
|
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): the saved plane ",
|
|
netplane.BootArmorPath, " could not be loaded: ", err,
|
|
" — LAN traffic may be reaching the WAN unprotected")
|
|
return false
|
|
case loaded:
|
|
logger.Error("fail-closed plane reinstated from ", netplane.BootArmorPath,
|
|
" (", why, "): LAN->WAN forwarding is BLOCKED. ",
|
|
"SSH, LuCI and the admin panel remain reachable.")
|
|
return true
|
|
default:
|
|
logger.Warn("no saved fail-closed plane at ", netplane.BootArmorPath, " (", why,
|
|
"): nothing was installed")
|
|
return false
|
|
}
|
|
}
|
|
|
|
// armFromModel renders the holding plane for m, installs it, and reports whether
|
|
// a plane is now standing.
|
|
//
|
|
// The return value is load-bearing on the exit path: Applier.TeardownExiting keeps
|
|
// the nft table only when a plane really was installed, so a render failure or a
|
|
// fail-open config falls back to the old remove-everything behaviour instead of
|
|
// leaving whatever the engine happened to have in the kernel.
|
|
func armFromModel(m *model.Model, why string, logger log.ContextLogger) bool {
|
|
ruleset, err := netplane.RenderHoldNft(m)
|
|
if err != nil {
|
|
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): render failed: ", err)
|
|
return false
|
|
}
|
|
if ruleset == "" {
|
|
// No divert devices: there is nothing this plane would protect.
|
|
return false
|
|
}
|
|
// One `nft -f` that opens with `delete table` and closes with the new table:
|
|
// the swap is a single netlink transaction, so this REPLACES whatever plane is
|
|
// loaded without the table ever being absent. That property is why the exit
|
|
// path can arm before it tears down.
|
|
if err := netplane.ApplyNft(ruleset); err != nil {
|
|
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): ", err,
|
|
" — LAN traffic may be reaching the WAN unprotected")
|
|
return false
|
|
}
|
|
logger.Info("fail-closed plane left in place (", why,
|
|
"): LAN->WAN forwarding stays BLOCKED until the next daemon applies. ",
|
|
"SSH, LuCI and the admin panel remain reachable.")
|
|
return true
|
|
}
|
|
|
|
// armOnUnreadableConfig is the window-3 answer: the daemon is up but cannot read
|
|
// its own desired state, so it falls back to the last state it persisted.
|
|
//
|
|
// A table that is ALREADY loaded is left alone. This runs on every reconcile —
|
|
// cron fires one a minute — and a `nft -f` is a delete-and-recreate of the whole
|
|
// table plus a DNS conntrack flush, so re-installing an identical plane sixty
|
|
// times an hour would be pure churn, and each replacement is itself a brief hole.
|
|
// The question this path answers is "is there anything at all standing", and once
|
|
// the answer is yes it stays yes until an apply succeeds and replaces it properly.
|
|
func armOnUnreadableConfig(logger log.ContextLogger) {
|
|
if netplane.TableExists() {
|
|
return
|
|
}
|
|
if planStartupArmor(errUnreadableConfig, netplane.BootArmorPresent()) != armorSnapshot {
|
|
logger.Warn("the config could not be read and there is no saved fail-closed plane at ",
|
|
netplane.BootArmorPath, " — nothing is protecting the LAN; fix /etc/config/shater ",
|
|
"(a full /overlay is the usual cause) and reconcile")
|
|
return
|
|
}
|
|
armFromSnapshot("the config could not be read", logger)
|
|
}
|
|
|
|
// errUnreadableConfig is a stand-in for "the read failed" in the call above,
|
|
// where the concrete error has already been logged by the caller.
|
|
var errUnreadableConfig = os.ErrInvalid
|
|
|
|
// armOnExit is the window-2 answer: what this daemon leaves in the kernel when it
|
|
// is asked to go away, and whether anything is now standing there. See
|
|
// planExitArmor for the policy.
|
|
//
|
|
// It is called BY Applier.TeardownExiting, before the teardown and under the apply
|
|
// lock, so that the plane is swapped rather than removed-then-rebuilt. It must
|
|
// therefore never call back into the Applier — everything here goes straight to
|
|
// model.ReadUCI and netplane.
|
|
func armOnExit(handoff bool, logger log.ContextLogger) bool {
|
|
m, err := model.ReadUCI()
|
|
switch planExitArmor(handoff, m, err, netplane.BootArmorPresent()) {
|
|
case armorRender:
|
|
return armFromModel(m, "restart handoff", logger)
|
|
case armorSnapshot:
|
|
return armFromSnapshot("restart handoff, config unreadable", logger)
|
|
}
|
|
return false
|
|
}
|