Files
shater/shater/apply/warnings.go
T
omarandClaude Opus 5 88a82c7297 feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:

  * two condition-less rules retire each other, and the LAST one by Order
    wins, so an earlier "default -> direct" is dead while looking live;
  * a condition-less rule can NEVER retire a rule that HAS conditions —
    those are emitted ahead of Final whatever their Order.

A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.

model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.

Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.

Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.

Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.

GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.

Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00

383 lines
17 KiB
Go

package apply
// Surfacing apply-time warnings to the operator.
//
// Over this audit we deliberately made many failure modes FAIL OPEN with a
// warning instead of refusing — an unreachable blocklist must not abort engine
// start and take the LAN down with it. That was the right trade, but it left a
// growing set of states that the log knows about and the user does not: a
// blocklist configured but not applied, a rule skipped because nothing was left
// to match on, a DoH host left un-blocked because it is also the upstream
// resolver. The panel showed a healthy screen over a router that was not doing
// what its config said.
//
// So the warnings from the last successful apply are kept and published through
// Status. Three sources converge here — generate.GenerateWithWarnings,
// netplane.RenderNftWithWarnings and model.Validate — and they are normalised
// into one shape the panel can group and highlight.
import (
"crypto/sha256"
"encoding/hex"
"fmt"
"regexp"
"sort"
"strconv"
"strings"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// Warning severities. Deliberately three, not a taxonomy:
//
// critical — something the operator CONFIGURED as protection is not in effect
// (a blocklist that did not load, a DoH host left reachable, an
// interface the kill-switch does not cover). This is the class that
// makes a healthy-looking panel a lie, so the UI must show it loudly.
// warning — a configured thing was skipped or degraded, but no protection
// claim is broken (an unresolvable chain hop, a device with no lease).
// info — operational notes worth surfacing but not acting on (cache moves).
const (
SeverityInfo = "info"
SeverityWarning = "warning"
SeverityCritical = "critical"
)
// maxStatusWarnings caps what Status carries, so a pathological config (hundreds
// of broken rules) cannot bloat every status poll. Warnings are sorted
// critical-first BEFORE the cap, so truncation can only ever drop the least
// important ones.
const maxStatusWarnings = 50
// Warning is one operator-facing apply warning.
//
// It mirrors model.Warning's Section/Name/Message — that type is the right shape
// and is reused verbatim for rule validation — and adds the Severity the panel
// needs to decide what to highlight. It is not a third format: model.Warning
// converts into this one field-for-field.
type Warning struct {
// Severity is one of the Severity* constants above.
Severity string `json:"severity"`
// Section is the kind of entity the warning is about ("rule", "ruleset",
// "blocklist", "device", "resolver", "chain", "group", "interface"), or a
// source bucket ("generate", "netplane", "config") when it is not attributable
// to a single entity.
Section string `json:"section"`
// Name is the entity's configured name, or "" when the warning is global.
// Together with Section it is what lets the panel deep-link to the offending
// section instead of printing a wall of text.
Name string `json:"name"`
// Message is the operator-facing sentence.
Message string `json:"message"`
}
// criticalMarkers are substrings that identify a warning as "protection you
// configured is not in effect".
//
// This is a heuristic over free text, and it is one on purpose: generate emits
// plain strings today, and inventing a parallel structured warning API across a
// package boundary owned by another agent would be a far larger change than the
// problem warrants. The markers below are taken verbatim from the actual warning
// texts, so they are exact rather than speculative. If generate ever emits its
// own severity, this list becomes dead code and the conversion simplifies.
var criticalMarkers = []string{
"UNREACHABLE", // remote rule-set/blocklist not applied
"is NOT applied", // ''
"NOT emitted", // block_doh NXDOMAIN rules missing
"left un-blocked", // block_doh: upstream resolver excluded
"inert", // dns_filter / per-device DNS configured but not working
"has NO effect", // dns_mode=fakeip with no fakeip resolver
"not covered", // an interface outside the fail-closed guard
"fail-closed", // ''
"REJECTED", // an unusable interface name
// A condition-less rule retired by a later condition-less rule whose target is
// `direct` (generate/route.go warnUnreachableRules): the operator's default
// policy — a tunnel, or a block — is not the one the router uses, so everything
// unmatched leaves on the plain WAN. The marker is the full clause, not the
// shorter "leaves over the plain WAN" that ruleKillFallback's kill=open note
// also contains: that one is a DELIBERATE per-rule bypass the operator asked
// for, and grading it critical here would be a different decision made by
// accident.
"leaves over the plain WAN with your real IP address",
}
// protectionSections are entity kinds whose whole purpose is to block or divert
// something. ANY warning about one of them means that list is not fully in
// effect, which is by definition a protection gap — so severity is decided by the
// SECTION here rather than by matching the sentence. That is deliberately more
// robust than text markers: generate has half a dozen different ways to say a
// list was skipped ("unreadable", "empty url", "needs a category", "unknown
// source"), and a marker list would silently miss the next one someone adds.
var protectionSections = map[string]bool{
"ruleset": true,
"blocklist": true,
"allowlist": true,
}
// infoMarkers identify operational notes that are not protection gaps.
var infoMarkers = []string{"cache:"}
// entityRe matches the `kind "name": message` prefix that generate's warnings use
// consistently, so Section/Name can be recovered from a plain string.
var entityRe = regexp.MustCompile(`^([a-z_]+) "([^"]*)": (.*)$`)
// warningFromText normalises one free-text warning. defaultSection is used when
// the text carries no `kind "name":` prefix.
func warningFromText(text, defaultSection, severity string) Warning {
w := Warning{Severity: severity, Section: defaultSection, Message: strings.TrimSpace(text)}
if m := entityRe.FindStringSubmatch(w.Message); m != nil {
w.Section, w.Name, w.Message = m[1], m[2], m[3]
}
return w
}
// classify picks a severity from the parsed section plus the message text.
// Section wins where it is decisive (see protectionSections); the markers then
// catch the global warnings that carry no entity prefix at all.
func classify(section, text string) string {
for _, m := range infoMarkers {
if strings.Contains(text, m) {
return SeverityInfo
}
}
if protectionSections[section] {
return SeverityCritical
}
for _, m := range criticalMarkers {
if strings.Contains(text, m) {
return SeverityCritical
}
}
return SeverityWarning
}
// collectWarnings normalises and merges the three sources, sorts them
// critical-first and applies the cap.
//
// - generateWarnings come from generate.GenerateWithWarnings (free text).
// - netplaneWarnings come from netplane.RenderNftWithWarnings. Every one of
// them is a device name that could not be expressed, which means the
// fail-closed guard does not cover that interface — always critical.
// - configWarnings come from model.Validate (already structured).
func collectWarnings(g model.Globals, generateWarnings, netplaneWarnings []string, configWarnings []model.Warning, planWarnings ...string) []Warning {
out := make([]Warning, 0, len(generateWarnings)+len(netplaneWarnings)+len(configWarnings)+1)
// The untunnelable-protocol policy always reports what it costs the user; it is
// the only one of these that describes correct behaviour rather than a fault.
out = append(out, untunnelablePolicyWarnings(g, planWarnings)...)
for _, t := range generateWarnings {
// Parse the entity prefix FIRST, so severity can be decided from the section
// (a warning about a blocklist is a protection gap however it is worded).
w := warningFromText(t, "generate", SeverityWarning)
w.Severity = classify(w.Section, t)
out = append(out, w)
}
for _, t := range netplaneWarnings {
w := warningFromText(t, "interface", SeverityCritical)
out = append(out, w)
}
for _, cw := range configWarnings {
out = append(out, Warning{
Severity: SeverityWarning,
Section: cw.Section,
Name: cw.Name,
Message: cw.Message,
})
}
// Stable sort by descending severity so the cap can only drop the least
// important entries, and the panel gets the worst news first.
sort.SliceStable(out, func(i, j int) bool {
return severityRank(out[i].Severity) > severityRank(out[j].Severity)
})
if len(out) > maxStatusWarnings {
dropped := len(out) - maxStatusWarnings
out = out[:maxStatusWarnings]
out[maxStatusWarnings-1] = Warning{
Severity: SeverityInfo,
Section: "generate",
Message: fmt.Sprintf(
"%d further warning(s) suppressed; run `logread -e shater` for the full list", dropped+1),
}
}
return out
}
// untunnelablePolicyWarnings explains, in terms of what the user will actually
// experience, what the untunnelable-protocol policy costs them.
//
// This exists because the symptom is otherwise completely mute: the user pings
// 1.1.1.1, gets 100% loss, and there is nothing anywhere — no error, no log line
// they would think to look for — connecting that to a setting. The message is
// phrased as consequence ("ping and IPTV will not work") rather than mechanism
// ("non-TCP/UDP is dropped"), because the mechanism is only meaningful to someone
// who already knows why.
//
// Severity is INFO throughout: this is correct, intended behaviour, not a
// malfunction. Reporting it as a warning would train the operator to ignore the
// warning list, which is the one thing W7 must not do.
func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
const section = "untunnelable"
policy := netplane.EffectiveUntunnelable(g)
// Notes from the destination plan: what could not be resolved and is therefore
// staying blocked, and how large the resulting address sets are. All INFO —
// each describes the policy working as designed, including its deliberate
// conservative fallbacks, rather than a fault. They must NOT travel on the
// netplane warning channel, which is reserved for interfaces the fail-closed
// guard cannot cover and is critical by construction.
out := make([]Warning, 0, len(planNotes)+1)
for _, n := range planNotes {
out = append(out, Warning{Severity: SeverityInfo, Section: section, Message: n})
}
note := func(name, msg string) []Warning {
return append(out, Warning{Severity: SeverityInfo, Section: section, Name: name, Message: msg})
}
// With the kill switch open the forward chain has no drops at all, so nothing
// is restricted whatever the policy says. Saying that is more useful than
// repeating a promise which is not being kept.
if !killSwitchClosed(g) {
if policy == netplane.UntunnelableDirect {
return out
}
return note(policy,
"This setting has no effect while the kill switch is open: with the kill switch open "+
"nothing is blocked, so ping, IPTV and VPN passthrough all work — and all of them "+
"reach the internet with your real IP address.")
}
switch policy {
case netplane.UntunnelableDirect:
return note(policy,
"Ping, IPTV and VPN passthrough (IPsec/PPTP) work everywhere, but they go straight out "+
"with your real IP address instead of through the tunnel — they are the kinds of "+
"traffic a tunnel cannot carry.")
case netplane.UntunnelableICMP:
return note(policy,
"Ping and traceroute work everywhere, including addresses you send through the tunnel; "+
"the host you ping sees your real IP address. IPTV and VPN passthrough (IPsec/PPTP) "+
"work only toward addresses your rules route directly.")
default:
return note(netplane.UntunnelableBlock,
"Ping, traceroute, IPTV and VPN passthrough work only toward addresses your rules route "+
"directly — those already see your real IP address anyway. Toward addresses you send "+
"through the tunnel they will not work, because a tunnel cannot carry them and they "+
"would otherwise leak your real IP address.")
}
}
func severityRank(s string) int {
switch s {
case SeverityCritical:
return 2
case SeverityWarning:
return 1
default:
return 0
}
}
// volatileNumberRe matches runs of digits inside a warning message.
//
// It exists because the fingerprint below must answer "is this the SAME news as
// last time", and several warnings carry a number that legitimately moves between
// two runs describing an unchanged situation:
//
// generate/cache.go: "only %d MiB free on /overlay" — free space drifts
// generate/ruleset.go: "compiled blocklists total %d KiB" — ditto
// generate/ruleset.go: "compiled %d domains from %s" — list size drifts
// any warning ending %v: "dial tcp 1.2.3.4:443: i/o timeout" — address/port vary
//
// Left raw, those make the set "change" on most reconciles and the whole fix
// stops working — the noise it removes is exactly the noise it would reintroduce.
// So numbers are masked for COMPARISON ONLY; what gets logged and what Status
// publishes are always the untouched messages.
//
// The cost is deliberate and small: a set that differs from the previous one only
// in a number is treated as unchanged and not reprinted. That is the right trade —
// "13 MiB free" following "14 MiB free" is not news, and the exact current figure
// is one `GET /api/status` away regardless.
var volatileNumberRe = regexp.MustCompile(`\d+`)
// warningsFingerprint reduces a set to a value that changes when, and only when,
// the operator has something new to read: an entry appeared or disappeared, or an
// entry's severity, entity or wording changed.
//
// The set arrives already sorted and capped, so a plain ordered walk is enough —
// no re-sorting, no set arithmetic. A hash rather than the joined text so the
// retained state is 32 bytes instead of a copy of up to 50 sentences. Field
// separators are 0x00, which cannot occur in these messages, so no combination of
// section/name/message can forge a different set's fingerprint.
func warningsFingerprint(ws []Warning) string {
h := sha256.New()
for _, w := range ws {
fmt.Fprintf(h, "%s\x00%s\x00%s\x00%s\x00",
w.Severity, w.Section, w.Name, volatileNumberRe.ReplaceAllString(w.Message, "#"))
}
return hex.EncodeToString(h.Sum(nil))
}
// emptyWarningsFP is the fingerprint of a clean apply, kept so "did the previous
// set have anything in it" is a comparison rather than a rehash.
var emptyWarningsFP = warningsFingerprint(nil)
// logWarningsIfChanged writes the merged set to the daemon log, but ONLY when it
// differs from the set this process last logged. It reports whether it wrote.
//
// Why gate at all. Status is polled and reconcile runs constantly — cron every
// minute plus every hotplug event — and each run used to reprint the whole set.
// On a healthy router that is one identical line per minute forever: ~1500 a day
// of "ping and traceroute work", which is both how a genuine warning (a blocklist
// that did not load, an interface the kill-switch misses) goes unnoticed and how
// the ring buffer that `logread` reads — memory, not disk — loses the history an
// incident is diagnosed from. A log that repeats itself is a log nobody reads.
//
// The first call after daemon start always prints, even if the set is identical to
// what the previous process logged: the fingerprint is process-local by design (see
// Applier.loggedWarnOnce). An operator reading the log after a reboot must find the
// router's current state there, not an empty tail justified by history they cannot
// see.
//
// Caller holds a.mu.
func (a *Applier) logWarningsIfChanged(ws []Warning) bool {
fp := warningsFingerprint(ws)
if a.loggedWarnOnce && fp == a.loggedWarnFP {
return false
}
hadPrevious := a.loggedWarnOnce && a.loggedWarnFP != emptyWarningsFP
a.loggedWarnFP, a.loggedWarnOnce = fp, true
if len(ws) == 0 {
// Silence would be ambiguous here: it reads the same as "nothing has been
// applied". Say the gap closed — that is the one line an operator watching a
// fix land actually wants.
if hadPrevious {
a.log.Info("config warnings: all previously reported warnings are resolved")
}
return hadPrevious
}
for _, w := range ws {
prefix := w.Section
if w.Name != "" {
prefix += " " + strconv.Quote(w.Name)
}
line := []any{"config warning [", w.Severity, "] ", prefix, ": ", w.Message}
// The syslog level must match the severity the line itself declares. An
// `[info]` note emitted at WARN is a false positive for anyone filtering the
// log by level, and teaches them that WARN from this daemon means nothing.
switch w.Severity {
case SeverityCritical:
a.log.Error(line...)
case SeverityInfo:
a.log.Info(line...)
default:
a.log.Warn(line...)
}
}
return true
}