A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:
* two condition-less rules retire each other, and the LAST one by Order
wins, so an earlier "default -> direct" is dead while looking live;
* a condition-less rule can NEVER retire a rule that HAS conditions —
those are emitted ahead of Final whatever their Order.
A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.
model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.
Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.
Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.
Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.
GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.
Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
383 lines
17 KiB
Go
383 lines
17 KiB
Go
package apply
|
|
|
|
// Surfacing apply-time warnings to the operator.
|
|
//
|
|
// Over this audit we deliberately made many failure modes FAIL OPEN with a
|
|
// warning instead of refusing — an unreachable blocklist must not abort engine
|
|
// start and take the LAN down with it. That was the right trade, but it left a
|
|
// growing set of states that the log knows about and the user does not: a
|
|
// blocklist configured but not applied, a rule skipped because nothing was left
|
|
// to match on, a DoH host left un-blocked because it is also the upstream
|
|
// resolver. The panel showed a healthy screen over a router that was not doing
|
|
// what its config said.
|
|
//
|
|
// So the warnings from the last successful apply are kept and published through
|
|
// Status. Three sources converge here — generate.GenerateWithWarnings,
|
|
// netplane.RenderNftWithWarnings and model.Validate — and they are normalised
|
|
// into one shape the panel can group and highlight.
|
|
|
|
import (
|
|
"crypto/sha256"
|
|
"encoding/hex"
|
|
"fmt"
|
|
"regexp"
|
|
"sort"
|
|
"strconv"
|
|
"strings"
|
|
|
|
"github.com/sagernet/sing-box/shater/model"
|
|
"github.com/sagernet/sing-box/shater/netplane"
|
|
)
|
|
|
|
// Warning severities. Deliberately three, not a taxonomy:
|
|
//
|
|
// critical — something the operator CONFIGURED as protection is not in effect
|
|
// (a blocklist that did not load, a DoH host left reachable, an
|
|
// interface the kill-switch does not cover). This is the class that
|
|
// makes a healthy-looking panel a lie, so the UI must show it loudly.
|
|
// warning — a configured thing was skipped or degraded, but no protection
|
|
// claim is broken (an unresolvable chain hop, a device with no lease).
|
|
// info — operational notes worth surfacing but not acting on (cache moves).
|
|
const (
|
|
SeverityInfo = "info"
|
|
SeverityWarning = "warning"
|
|
SeverityCritical = "critical"
|
|
)
|
|
|
|
// maxStatusWarnings caps what Status carries, so a pathological config (hundreds
|
|
// of broken rules) cannot bloat every status poll. Warnings are sorted
|
|
// critical-first BEFORE the cap, so truncation can only ever drop the least
|
|
// important ones.
|
|
const maxStatusWarnings = 50
|
|
|
|
// Warning is one operator-facing apply warning.
|
|
//
|
|
// It mirrors model.Warning's Section/Name/Message — that type is the right shape
|
|
// and is reused verbatim for rule validation — and adds the Severity the panel
|
|
// needs to decide what to highlight. It is not a third format: model.Warning
|
|
// converts into this one field-for-field.
|
|
type Warning struct {
|
|
// Severity is one of the Severity* constants above.
|
|
Severity string `json:"severity"`
|
|
// Section is the kind of entity the warning is about ("rule", "ruleset",
|
|
// "blocklist", "device", "resolver", "chain", "group", "interface"), or a
|
|
// source bucket ("generate", "netplane", "config") when it is not attributable
|
|
// to a single entity.
|
|
Section string `json:"section"`
|
|
// Name is the entity's configured name, or "" when the warning is global.
|
|
// Together with Section it is what lets the panel deep-link to the offending
|
|
// section instead of printing a wall of text.
|
|
Name string `json:"name"`
|
|
// Message is the operator-facing sentence.
|
|
Message string `json:"message"`
|
|
}
|
|
|
|
// criticalMarkers are substrings that identify a warning as "protection you
|
|
// configured is not in effect".
|
|
//
|
|
// This is a heuristic over free text, and it is one on purpose: generate emits
|
|
// plain strings today, and inventing a parallel structured warning API across a
|
|
// package boundary owned by another agent would be a far larger change than the
|
|
// problem warrants. The markers below are taken verbatim from the actual warning
|
|
// texts, so they are exact rather than speculative. If generate ever emits its
|
|
// own severity, this list becomes dead code and the conversion simplifies.
|
|
var criticalMarkers = []string{
|
|
"UNREACHABLE", // remote rule-set/blocklist not applied
|
|
"is NOT applied", // ''
|
|
"NOT emitted", // block_doh NXDOMAIN rules missing
|
|
"left un-blocked", // block_doh: upstream resolver excluded
|
|
"inert", // dns_filter / per-device DNS configured but not working
|
|
"has NO effect", // dns_mode=fakeip with no fakeip resolver
|
|
"not covered", // an interface outside the fail-closed guard
|
|
"fail-closed", // ''
|
|
"REJECTED", // an unusable interface name
|
|
// A condition-less rule retired by a later condition-less rule whose target is
|
|
// `direct` (generate/route.go warnUnreachableRules): the operator's default
|
|
// policy — a tunnel, or a block — is not the one the router uses, so everything
|
|
// unmatched leaves on the plain WAN. The marker is the full clause, not the
|
|
// shorter "leaves over the plain WAN" that ruleKillFallback's kill=open note
|
|
// also contains: that one is a DELIBERATE per-rule bypass the operator asked
|
|
// for, and grading it critical here would be a different decision made by
|
|
// accident.
|
|
"leaves over the plain WAN with your real IP address",
|
|
}
|
|
|
|
// protectionSections are entity kinds whose whole purpose is to block or divert
|
|
// something. ANY warning about one of them means that list is not fully in
|
|
// effect, which is by definition a protection gap — so severity is decided by the
|
|
// SECTION here rather than by matching the sentence. That is deliberately more
|
|
// robust than text markers: generate has half a dozen different ways to say a
|
|
// list was skipped ("unreadable", "empty url", "needs a category", "unknown
|
|
// source"), and a marker list would silently miss the next one someone adds.
|
|
var protectionSections = map[string]bool{
|
|
"ruleset": true,
|
|
"blocklist": true,
|
|
"allowlist": true,
|
|
}
|
|
|
|
// infoMarkers identify operational notes that are not protection gaps.
|
|
var infoMarkers = []string{"cache:"}
|
|
|
|
// entityRe matches the `kind "name": message` prefix that generate's warnings use
|
|
// consistently, so Section/Name can be recovered from a plain string.
|
|
var entityRe = regexp.MustCompile(`^([a-z_]+) "([^"]*)": (.*)$`)
|
|
|
|
// warningFromText normalises one free-text warning. defaultSection is used when
|
|
// the text carries no `kind "name":` prefix.
|
|
func warningFromText(text, defaultSection, severity string) Warning {
|
|
w := Warning{Severity: severity, Section: defaultSection, Message: strings.TrimSpace(text)}
|
|
if m := entityRe.FindStringSubmatch(w.Message); m != nil {
|
|
w.Section, w.Name, w.Message = m[1], m[2], m[3]
|
|
}
|
|
return w
|
|
}
|
|
|
|
// classify picks a severity from the parsed section plus the message text.
|
|
// Section wins where it is decisive (see protectionSections); the markers then
|
|
// catch the global warnings that carry no entity prefix at all.
|
|
func classify(section, text string) string {
|
|
for _, m := range infoMarkers {
|
|
if strings.Contains(text, m) {
|
|
return SeverityInfo
|
|
}
|
|
}
|
|
if protectionSections[section] {
|
|
return SeverityCritical
|
|
}
|
|
for _, m := range criticalMarkers {
|
|
if strings.Contains(text, m) {
|
|
return SeverityCritical
|
|
}
|
|
}
|
|
return SeverityWarning
|
|
}
|
|
|
|
// collectWarnings normalises and merges the three sources, sorts them
|
|
// critical-first and applies the cap.
|
|
//
|
|
// - generateWarnings come from generate.GenerateWithWarnings (free text).
|
|
// - netplaneWarnings come from netplane.RenderNftWithWarnings. Every one of
|
|
// them is a device name that could not be expressed, which means the
|
|
// fail-closed guard does not cover that interface — always critical.
|
|
// - configWarnings come from model.Validate (already structured).
|
|
func collectWarnings(g model.Globals, generateWarnings, netplaneWarnings []string, configWarnings []model.Warning, planWarnings ...string) []Warning {
|
|
out := make([]Warning, 0, len(generateWarnings)+len(netplaneWarnings)+len(configWarnings)+1)
|
|
// The untunnelable-protocol policy always reports what it costs the user; it is
|
|
// the only one of these that describes correct behaviour rather than a fault.
|
|
out = append(out, untunnelablePolicyWarnings(g, planWarnings)...)
|
|
for _, t := range generateWarnings {
|
|
// Parse the entity prefix FIRST, so severity can be decided from the section
|
|
// (a warning about a blocklist is a protection gap however it is worded).
|
|
w := warningFromText(t, "generate", SeverityWarning)
|
|
w.Severity = classify(w.Section, t)
|
|
out = append(out, w)
|
|
}
|
|
for _, t := range netplaneWarnings {
|
|
w := warningFromText(t, "interface", SeverityCritical)
|
|
out = append(out, w)
|
|
}
|
|
for _, cw := range configWarnings {
|
|
out = append(out, Warning{
|
|
Severity: SeverityWarning,
|
|
Section: cw.Section,
|
|
Name: cw.Name,
|
|
Message: cw.Message,
|
|
})
|
|
}
|
|
|
|
// Stable sort by descending severity so the cap can only drop the least
|
|
// important entries, and the panel gets the worst news first.
|
|
sort.SliceStable(out, func(i, j int) bool {
|
|
return severityRank(out[i].Severity) > severityRank(out[j].Severity)
|
|
})
|
|
|
|
if len(out) > maxStatusWarnings {
|
|
dropped := len(out) - maxStatusWarnings
|
|
out = out[:maxStatusWarnings]
|
|
out[maxStatusWarnings-1] = Warning{
|
|
Severity: SeverityInfo,
|
|
Section: "generate",
|
|
Message: fmt.Sprintf(
|
|
"%d further warning(s) suppressed; run `logread -e shater` for the full list", dropped+1),
|
|
}
|
|
}
|
|
return out
|
|
}
|
|
|
|
// untunnelablePolicyWarnings explains, in terms of what the user will actually
|
|
// experience, what the untunnelable-protocol policy costs them.
|
|
//
|
|
// This exists because the symptom is otherwise completely mute: the user pings
|
|
// 1.1.1.1, gets 100% loss, and there is nothing anywhere — no error, no log line
|
|
// they would think to look for — connecting that to a setting. The message is
|
|
// phrased as consequence ("ping and IPTV will not work") rather than mechanism
|
|
// ("non-TCP/UDP is dropped"), because the mechanism is only meaningful to someone
|
|
// who already knows why.
|
|
//
|
|
// Severity is INFO throughout: this is correct, intended behaviour, not a
|
|
// malfunction. Reporting it as a warning would train the operator to ignore the
|
|
// warning list, which is the one thing W7 must not do.
|
|
func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
|
|
const section = "untunnelable"
|
|
policy := netplane.EffectiveUntunnelable(g)
|
|
|
|
// Notes from the destination plan: what could not be resolved and is therefore
|
|
// staying blocked, and how large the resulting address sets are. All INFO —
|
|
// each describes the policy working as designed, including its deliberate
|
|
// conservative fallbacks, rather than a fault. They must NOT travel on the
|
|
// netplane warning channel, which is reserved for interfaces the fail-closed
|
|
// guard cannot cover and is critical by construction.
|
|
out := make([]Warning, 0, len(planNotes)+1)
|
|
for _, n := range planNotes {
|
|
out = append(out, Warning{Severity: SeverityInfo, Section: section, Message: n})
|
|
}
|
|
|
|
note := func(name, msg string) []Warning {
|
|
return append(out, Warning{Severity: SeverityInfo, Section: section, Name: name, Message: msg})
|
|
}
|
|
|
|
// With the kill switch open the forward chain has no drops at all, so nothing
|
|
// is restricted whatever the policy says. Saying that is more useful than
|
|
// repeating a promise which is not being kept.
|
|
if !killSwitchClosed(g) {
|
|
if policy == netplane.UntunnelableDirect {
|
|
return out
|
|
}
|
|
return note(policy,
|
|
"This setting has no effect while the kill switch is open: with the kill switch open "+
|
|
"nothing is blocked, so ping, IPTV and VPN passthrough all work — and all of them "+
|
|
"reach the internet with your real IP address.")
|
|
}
|
|
|
|
switch policy {
|
|
case netplane.UntunnelableDirect:
|
|
return note(policy,
|
|
"Ping, IPTV and VPN passthrough (IPsec/PPTP) work everywhere, but they go straight out "+
|
|
"with your real IP address instead of through the tunnel — they are the kinds of "+
|
|
"traffic a tunnel cannot carry.")
|
|
case netplane.UntunnelableICMP:
|
|
return note(policy,
|
|
"Ping and traceroute work everywhere, including addresses you send through the tunnel; "+
|
|
"the host you ping sees your real IP address. IPTV and VPN passthrough (IPsec/PPTP) "+
|
|
"work only toward addresses your rules route directly.")
|
|
default:
|
|
return note(netplane.UntunnelableBlock,
|
|
"Ping, traceroute, IPTV and VPN passthrough work only toward addresses your rules route "+
|
|
"directly — those already see your real IP address anyway. Toward addresses you send "+
|
|
"through the tunnel they will not work, because a tunnel cannot carry them and they "+
|
|
"would otherwise leak your real IP address.")
|
|
}
|
|
}
|
|
|
|
func severityRank(s string) int {
|
|
switch s {
|
|
case SeverityCritical:
|
|
return 2
|
|
case SeverityWarning:
|
|
return 1
|
|
default:
|
|
return 0
|
|
}
|
|
}
|
|
|
|
// volatileNumberRe matches runs of digits inside a warning message.
|
|
//
|
|
// It exists because the fingerprint below must answer "is this the SAME news as
|
|
// last time", and several warnings carry a number that legitimately moves between
|
|
// two runs describing an unchanged situation:
|
|
//
|
|
// generate/cache.go: "only %d MiB free on /overlay" — free space drifts
|
|
// generate/ruleset.go: "compiled blocklists total %d KiB" — ditto
|
|
// generate/ruleset.go: "compiled %d domains from %s" — list size drifts
|
|
// any warning ending %v: "dial tcp 1.2.3.4:443: i/o timeout" — address/port vary
|
|
//
|
|
// Left raw, those make the set "change" on most reconciles and the whole fix
|
|
// stops working — the noise it removes is exactly the noise it would reintroduce.
|
|
// So numbers are masked for COMPARISON ONLY; what gets logged and what Status
|
|
// publishes are always the untouched messages.
|
|
//
|
|
// The cost is deliberate and small: a set that differs from the previous one only
|
|
// in a number is treated as unchanged and not reprinted. That is the right trade —
|
|
// "13 MiB free" following "14 MiB free" is not news, and the exact current figure
|
|
// is one `GET /api/status` away regardless.
|
|
var volatileNumberRe = regexp.MustCompile(`\d+`)
|
|
|
|
// warningsFingerprint reduces a set to a value that changes when, and only when,
|
|
// the operator has something new to read: an entry appeared or disappeared, or an
|
|
// entry's severity, entity or wording changed.
|
|
//
|
|
// The set arrives already sorted and capped, so a plain ordered walk is enough —
|
|
// no re-sorting, no set arithmetic. A hash rather than the joined text so the
|
|
// retained state is 32 bytes instead of a copy of up to 50 sentences. Field
|
|
// separators are 0x00, which cannot occur in these messages, so no combination of
|
|
// section/name/message can forge a different set's fingerprint.
|
|
func warningsFingerprint(ws []Warning) string {
|
|
h := sha256.New()
|
|
for _, w := range ws {
|
|
fmt.Fprintf(h, "%s\x00%s\x00%s\x00%s\x00",
|
|
w.Severity, w.Section, w.Name, volatileNumberRe.ReplaceAllString(w.Message, "#"))
|
|
}
|
|
return hex.EncodeToString(h.Sum(nil))
|
|
}
|
|
|
|
// emptyWarningsFP is the fingerprint of a clean apply, kept so "did the previous
|
|
// set have anything in it" is a comparison rather than a rehash.
|
|
var emptyWarningsFP = warningsFingerprint(nil)
|
|
|
|
// logWarningsIfChanged writes the merged set to the daemon log, but ONLY when it
|
|
// differs from the set this process last logged. It reports whether it wrote.
|
|
//
|
|
// Why gate at all. Status is polled and reconcile runs constantly — cron every
|
|
// minute plus every hotplug event — and each run used to reprint the whole set.
|
|
// On a healthy router that is one identical line per minute forever: ~1500 a day
|
|
// of "ping and traceroute work", which is both how a genuine warning (a blocklist
|
|
// that did not load, an interface the kill-switch misses) goes unnoticed and how
|
|
// the ring buffer that `logread` reads — memory, not disk — loses the history an
|
|
// incident is diagnosed from. A log that repeats itself is a log nobody reads.
|
|
//
|
|
// The first call after daemon start always prints, even if the set is identical to
|
|
// what the previous process logged: the fingerprint is process-local by design (see
|
|
// Applier.loggedWarnOnce). An operator reading the log after a reboot must find the
|
|
// router's current state there, not an empty tail justified by history they cannot
|
|
// see.
|
|
//
|
|
// Caller holds a.mu.
|
|
func (a *Applier) logWarningsIfChanged(ws []Warning) bool {
|
|
fp := warningsFingerprint(ws)
|
|
if a.loggedWarnOnce && fp == a.loggedWarnFP {
|
|
return false
|
|
}
|
|
hadPrevious := a.loggedWarnOnce && a.loggedWarnFP != emptyWarningsFP
|
|
a.loggedWarnFP, a.loggedWarnOnce = fp, true
|
|
|
|
if len(ws) == 0 {
|
|
// Silence would be ambiguous here: it reads the same as "nothing has been
|
|
// applied". Say the gap closed — that is the one line an operator watching a
|
|
// fix land actually wants.
|
|
if hadPrevious {
|
|
a.log.Info("config warnings: all previously reported warnings are resolved")
|
|
}
|
|
return hadPrevious
|
|
}
|
|
|
|
for _, w := range ws {
|
|
prefix := w.Section
|
|
if w.Name != "" {
|
|
prefix += " " + strconv.Quote(w.Name)
|
|
}
|
|
line := []any{"config warning [", w.Severity, "] ", prefix, ": ", w.Message}
|
|
// The syslog level must match the severity the line itself declares. An
|
|
// `[info]` note emitted at WARN is a false positive for anyone filtering the
|
|
// log by level, and teaches them that WARN from this daemon means nothing.
|
|
switch w.Severity {
|
|
case SeverityCritical:
|
|
a.log.Error(line...)
|
|
case SeverityInfo:
|
|
a.log.Info(line...)
|
|
default:
|
|
a.log.Warn(line...)
|
|
}
|
|
}
|
|
return true
|
|
}
|