Files
shater/shater/netplane/nft.go
T
omarandClaude Opus 5 cbda0fee0a fix(netplane): arm the fail-closed plane before the daemon can
The plane only ever existed while the daemon did. It starts at 99, after fw4 has
already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its
predecessor, migrating the schema, building the engine and reading UCI — with a
UPX-compressed binary decompressing off flash first. Every boot therefore had a
window with no protection at all, landing exactly when Wi-Fi comes up and every
client reconnects. A restart, a reload or a package upgrade opened the same
window on purpose: Teardown does not consult the kill switch, and the init script
guarantees the interval is non-empty.

The holding plane is now persisted to /etc/shater/boot.nft on every apply and
loaded by a small service at 21, right after fw4 and netifd. Its presence is the
arm token: it exists only while the last applied config was enabled AND
fail-closed, and goes away the moment either stops being true. Writes are
content-gated — the cron reconcile runs a minute — and atomic, because the one
boot that reads this file is the boot after a power cut.

The service refuses to arm four ways so it can never brick a box, and its
enabled-check reads /etc/rc.d directly rather than asking rc.common, which would
take a blocking flock in the middle of boot. On exit the daemon re-arms only for
restart and reload, read from a snapshot of rc.common's action; anything else,
including an unknown one, degrades to a real stop that also disarms.

An unreadable config used to leave the router bare forever: the arm call sat in
the branch that requires a successful read, and nothing downstream could recover
it. It now arms from the same path.

A network nobody named was neither diverted nor blocked — the divert set is built
from inbounds and rule sources, and the same set scopes the fail-closed drops. It
is now enumerated from the interfaces whose firewall zone the operator forwards
to a WAN zone — their own statement that those clients reach the internet through
this box — and reported critically, by name, with both resolutions. Deliberately
not closed automatically: this router cannot know a guest SSID was meant to be
off the tunnel, and guessing is an outage. A device name that resolved to nothing
is reported the same way, for the same reason: there is no fail-closed action
available for a device we cannot name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:16 +03:00

952 lines
40 KiB
Go

// Package netplane is the engine-agnostic data plane of the shater control-plane:
// the `inet shater` nftables ruleset (TPROXY divert of all LAN traffic incl. :53 +
// DoT/DoQ :853 reject + kill-switch +
// management-bypass), the policy-routing that delivers marked packets locally, and
// the sysctl knobs the plane depends on.
//
// It is deliberately independent of the proxy engine: everything here drives the
// external `nft` and `ip` (iproute2) binaries plus `sysctl`, never the engine's
// own config. It is ported near-verbatim from the v0.1 xrayctl nft/routing plane
// (egress.go, nftstats.go, and the RenderNft/applyNft/applyRouting parts of
// apply.go). Only the tproxy port and the loop-guard mark wiring are re-pointed
// at the sing-box engine; the ruleset skeleton is unchanged.
//
// It never touches fw4: the only table it owns is `inet shater`, the only marks
// it uses derive from Globals.FwmarkBase, and the only routing tables it uses
// derive from Globals.TableBase.
package netplane
import (
"fmt"
"net"
"regexp"
"strings"
"time"
"github.com/sagernet/sing-box/shater/model"
)
// Table/set identifiers and marks — the naming contract consumed by stats.go and
// (downstream) the LuCI dashboard.
const (
// Table is our nft table; never fw4.
Table = "inet shater"
// Client4 / Client6 are the per-LAN-source dynamic accounting sets.
Client4 = "clients"
Client6 = "clients6"
// loopMark is the loop-guard mark stamped on engine-originated egress (the
// engine sets RoutingMark = loopMark on its outbounds + tproxy inbound). The
// prerouting chain accepts it FIRST so re-entrant traffic is never re-diverted.
loopMark = 0xff
// egEgressMarkOffset / egEgressTableOffset keep per-egress marks/tables clear
// of the tproxy divert mark (FwmarkBase) and loop-guard mark (loopMark).
egEgressMarkOffset = 0x100
egEgressTableOffset = 0x10
// defaultFwmark / defaultTable are the model contract defaults when a raw
// Globals arrives with the fields unset (0).
defaultFwmark = 0x2000
defaultTable = 0x2000
)
// effFwmark / effTable normalise a possibly-zero Globals field to the contract
// default (0x2000), so RenderNft/ApplyRouting behave even on a bare model.
func effFwmark(g model.Globals) uint32 {
if g.FwmarkBase == 0 {
return defaultFwmark
}
return g.FwmarkBase
}
func effTable(g model.Globals) uint32 {
if g.TableBase == 0 {
return defaultTable
}
return g.TableBase
}
// EgressMark returns the deterministic SO_MARK for the idx-th egress (its
// position in Model.Egresses). The generator stamps the same mark on the egress
// outbound; ApplyRouting binds it to a routing table; RenderNft accepts it in the
// prerouting bypass so egress traffic is never re-tproxied.
func EgressMark(g model.Globals, idx int) uint32 {
return effFwmark(g) + egEgressMarkOffset + uint32(idx)
}
// EgressTable is the dedicated routing table for the idx-th egress, in the
// reserved block just above the main tproxy table (== table_base).
func EgressTable(g model.Globals, idx int) uint32 {
return effTable(g) + egEgressTableOffset + uint32(idx)
}
// EgressOutboundTag is the outbound tag the generator emits for an
// interface/tunnel egress (coordination contract with the generate stage).
func EgressOutboundTag(name string) string { return "egress-" + name }
// genGlobalClosed reports whether the effective global kill-switch is "closed"
// (fail-closed). Default is closed; only an explicit "open" fails open.
func genGlobalClosed(g model.Globals) bool {
return strings.ToLower(strings.TrimSpace(g.KillSwitch)) != "open"
}
// --- source classification (shared contract with the generate stage) ---
// SrcClasses is a rule's Src list split by kind. Only CIDRs belong in the
// engine's routing `source`; MAC/iface/zone are the netplane's nft matches.
type SrcClasses struct {
CIDRs []string // IP / CIDR / host — engine `source`
MACs []string // AA:BB:CC:DD:EE:FF — nft `ether saddr`
Ifaces []string // iface:<name> — nft `iifname`
Zones []string // zone:<name> — fw4 zone -> devices
}
var macRe = regexp.MustCompile(`^[0-9a-fA-F]{2}(:[0-9a-fA-F]{2}){5}$`)
// ClassifySrc classifies each rule.Src entry (ported from v0.1 genClassifySrc).
func ClassifySrc(src []string) SrcClasses {
var c SrcClasses
for _, raw := range src {
v := strings.TrimSpace(raw)
if v == "" {
continue
}
switch {
case strings.HasPrefix(v, "iface:"):
c.Ifaces = append(c.Ifaces, strings.TrimPrefix(v, "iface:"))
case strings.HasPrefix(v, "zone:"):
c.Zones = append(c.Zones, strings.TrimPrefix(v, "zone:"))
case macRe.MatchString(v):
c.MACs = append(c.MACs, v)
default:
c.CIDRs = append(c.CIDRs, v)
}
}
return c
}
// --- nft identifier / counter naming ---
// nftIdent maps an arbitrary name to a safe nft identifier ([A-Za-z0-9_]).
func nftIdent(s string) string {
if s == "" {
return "unnamed"
}
return strings.Map(func(r rune) rune {
if (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') || (r >= '0' && r <= '9') || r == '_' {
return r
}
return '_'
}, s)
}
// nftRuleKey is the stable key for a rule's counter (name, else order-N).
func nftRuleKey(r model.Rule) string {
if strings.TrimSpace(r.Name) != "" {
return r.Name
}
return fmt.Sprintf("order-%d", r.Order)
}
func nftRuleCounter(r model.Rule) string { return "c_rule_" + nftIdent(nftRuleKey(r)) }
func nftInCounter(in model.Inbound) string {
return "c_in_" + nftIdent(orDefault(in.Name, "lan"))
}
func orDefault(s, d string) string {
if s == "" {
return d
}
return s
}
// --- src fragments (first-match OR semantics) ---
// nftFrag is one nft match line derived from a single rule.Src entry. iif, when
// non-empty, overrides the inbound device set (iface:/zone: sources). match is an
// optional saddr expression ("ip saddr X" / "ip6 saddr X" / "ether saddr X").
// fam is 4, 6, or 0 (both).
type nftFrag struct {
iif []string
match string
fam int
}
// nftSrcFrags converts a rule's src list into independent match fragments,
// preserving first-match OR semantics (each src entry => its own line(s)).
// inboundDevs is the default iif for ip/mac sources.
func nftSrcFrags(r model.Rule, inboundDevs []string) []nftFrag {
var frags []nftFrag
for _, raw := range r.Src {
s := strings.TrimSpace(raw)
if s == "" {
continue
}
switch {
case strings.HasPrefix(s, "iface:"):
dev := IfaceDevice(strings.TrimPrefix(s, "iface:"))
frags = append(frags, nftFrag{iif: []string{dev}, fam: 0})
case strings.HasPrefix(s, "zone:"):
devs := nftZoneDevices(strings.TrimPrefix(s, "zone:"))
if len(devs) > 0 {
frags = append(frags, nftFrag{iif: devs, fam: 0})
}
case macRe.MatchString(s):
frags = append(frags, nftFrag{iif: inboundDevs, match: "ether saddr " + s, fam: 0})
default:
cidr, fam := nftNormalizeCIDR(s)
if cidr == "" {
continue
}
if fam == 6 {
frags = append(frags, nftFrag{iif: inboundDevs, match: "ip6 saddr " + cidr, fam: 6})
} else {
frags = append(frags, nftFrag{iif: inboundDevs, match: "ip saddr " + cidr, fam: 4})
}
}
}
return frags
}
// nftNormalizeCIDR returns a canonical CIDR string and family (4/6) for a CIDR
// or bare host; "" if unparseable.
func nftNormalizeCIDR(s string) (string, int) {
if strings.Contains(s, "/") {
if ip, _, err := net.ParseCIDR(s); err == nil {
if ip.To4() != nil {
return s, 4
}
return s, 6
}
return "", 0
}
ip := net.ParseIP(s)
if ip == nil {
return "", 0
}
if ip.To4() != nil {
return s + "/32", 4
}
return s + "/128", 6
}
// maxDevNameLen is IFNAMSIZ-1: the longest network-device name the kernel will
// accept, so anything longer cannot name a real interface.
const maxDevNameLen = 15
// nftValidDev reports whether dev is usable as an nft `iifname` operand, and
// returns it trimmed. It VALIDATES — it never rewrites.
//
// That distinction is the whole point. Device names reach us from UCI / `ubus
// network.interface status` and, when neither resolves, straight from an
// operator-supplied model field, so a name containing a quote, newline or `;`
// would break out of the `iifname "..."` token and inject nft statements into our
// table. The obvious defence — strip the offending characters — is a trap: it
// turns `br-lan;drop` into `br-landrop`, a syntactically valid name for an
// interface that does not exist. The resulting `iifname "br-landrop" ... drop`
// loads into the kernel without a murmur and matches NOTHING, so fail-closed
// silently stops covering that interface — reintroducing exactly the plaintext
// leak the forward-chain guard exists to prevent, but invisibly.
//
// So: a name is either exactly right or rejected outright, and rejection is
// surfaced to the operator (see RenderNftWithWarnings) rather than papered over.
//
// The accepted set is [A-Za-z0-9._-], 1..15 chars, excluding "." and "..". That
// covers every name OpenWrt produces (br-lan, eth0.100, pppoe-wan, 6in4-wan,
// wg0, wlan0-1) and every raw UCI section name IfaceDevice falls back to. The
// kernel's own dev_valid_name is marginally laxer (it only bars "", ".", "..",
// '/', ':' and whitespace), so an exotic-but-legal name like `br+lan` is rejected
// here — loudly, which is the safe direction to err.
func nftValidDev(dev string) (string, bool) {
d := strings.TrimSpace(dev)
if d == "" || len(d) > maxDevNameLen || d == "." || d == ".." {
return "", false
}
for _, r := range d {
switch {
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9':
case r == '.' || r == '_' || r == '-':
default:
return "", false
}
}
return d, true
}
// nftSplitDevs partitions devs into the valid names (deduped, order-stable) and
// the rejected ones (verbatim, so a warning can quote what the operator wrote).
func nftSplitDevs(devs []string) (valid, rejected []string) {
for _, d := range nftDedupStr(devs) {
if v, ok := nftValidDev(d); ok {
valid = append(valid, v)
} else {
rejected = append(rejected, d)
}
}
return nftDedupStr(valid), rejected
}
// nftIifExpr renders an iifname match for one or more devices. Invalid names are
// dropped here; callers that depend on complete coverage (the fail-closed guard)
// must have already refused the plan via nftSplitDevs — see RenderNftWithWarnings.
func nftIifExpr(devs []string) string {
valid, _ := nftSplitDevs(devs)
if len(valid) == 0 {
return ""
}
if len(valid) == 1 {
return fmt.Sprintf("iifname \"%s\"", valid[0])
}
q := make([]string, len(valid))
for i, d := range valid {
q[i] = "\"" + d + "\""
}
return "iifname { " + strings.Join(q, ", ") + " }"
}
func nftDedupStr(in []string) []string {
seen := map[string]bool{}
var out []string
for _, v := range in {
v = strings.TrimSpace(v)
if v == "" || seen[v] {
continue
}
seen[v] = true
out = append(out, v)
}
return out
}
// nftEnabledInboundDevs returns the L3 devices of all enabled tproxy inbounds.
func nftEnabledInboundDevs(m *model.Model) []string {
var out []string
for _, in := range m.Inbounds {
if model.IsTproxyInbound(in) {
out = append(out, IfaceDevice(in.Network))
}
}
return nftDedupStr(out)
}
// nftPlanRules returns the rules the data plane must honour: m.Rules with the
// ACTIVE PROFILE's EnableRules/DisableRules applied — via
// model.ResolveActiveProfile + model.ApplyProfileRuleOverrides, the exact
// resolution the generate stage uses for the engine plan — stably sorted by
// Order. m is never mutated (the model helper copies).
//
// Why not raw m.Rules: a rule DISABLED in UCI but force-enabled by the active
// profile gets its engine route rule from generate, so it MUST also get its
// per-rule tproxy divert, join the fail-closed drop coverage and receive
// accept_local — otherwise its traffic never enters the engine at all. The
// inverse holds too: a profile-DISABLED rule must not keep diverting a device
// nothing routes. Reading the raw Enabled flag here was exactly that
// divert/engine divergence.
//
// The resolution warnings are deliberately discarded: generate runs the same
// resolver on the same apply and is the single reporter, and netplane's own
// warning channel is reserved for fail-closed coverage gaps (graded critical
// wholesale).
//
// The clock parameter stays so the apply pipeline keeps threading the generate
// timestamp through the At-variants unchanged; profile selection itself is no
// longer time-dependent.
func nftPlanRules(m *model.Model, _ time.Time) []model.Rule {
prof, _ := model.ResolveActiveProfile(m)
rules, _ := model.ApplyProfileRuleOverrides(m.Rules, prof)
return sortedRules(rules)
}
// divertRef is one device this ruleset diverts, together with WHAT THE OPERATOR
// WROTE to bring it in — the inbound's `network`, an `iface:` source, a `zone:`
// source.
//
// The provenance is not decoration. IfaceDevice returns the UCI name unchanged
// when it cannot resolve it (see its doc comment), so a divert set is a list of
// strings in which "br-lan" and "lan" are indistinguishable — one is a device,
// the other is a name that will match nothing. Reporting that to an operator as
// `device "lan" does not exist` is useless; reporting it as `interface "lan"
// resolved to no device` is actionable, and only the reference knows which is
// which. coverage.go is the only consumer.
type divertRef struct {
// Src is the operator-facing name: the inbound's UCI network, the bare name
// from an `iface:` source, or "zone:<name>" for a device pulled out of a
// firewall zone.
Src string
// Dev is what IfaceDevice/nftZoneDevices resolved Src to — the string that
// actually appears in `iifname "..."`.
Dev string
}
// nftDivertRefs returns EVERY LAN-ingress device whose traffic this ruleset
// diverts into the engine, with its provenance: the enabled tproxy inbounds'
// devices PLUS the devices named by any enabled rule's `iface:`/`zone:` source.
// rules is the EFFECTIVE rule list (profile overrides applied — nftPlanRules),
// never raw m.Rules: the divert set must match what RenderNft actually emits
// per-rule.
//
// The distinction matters for fail-closed correctness. RenderNft emits per-rule
// tproxy lines for `iface:`/`zone:` sources, and those devices need NOT be tproxy
// inbounds. Scoping the forward-chain kill-switch to the inbound devices alone
// therefore left a hole: when the engine is down, `tproxy` returns NFT_BREAK, the
// packet falls out of prerouting UNMARKED, reaches the forward hook — and, not
// matching the inbound-only drop, was FORWARDED TO WAN in plaintext. Same story
// for ApplyIfaceSysctls: an ingress device without accept_local=1 cannot deliver
// the tproxied packet to the engine socket at all.
//
// So: everything we divert, we must also be able to fail closed on and set the
// ingress sysctls for. Result is deduped BY DEVICE and order-stable (inbounds
// first), so the first reference that produced a device is the one reported.
func nftDivertRefs(m *model.Model, rules []model.Rule) []divertRef {
var refs []divertRef
for _, in := range m.Inbounds {
if model.IsTproxyInbound(in) {
refs = append(refs, divertRef{
Src: orDefault(strings.TrimSpace(in.Network), "lan"),
Dev: IfaceDevice(in.Network),
})
}
}
// No tproxy inbound => RenderNft emits no per-rule divert lines either; and
// with no inbound device there is nothing for a rule source to add to.
if _, ok := nftPrimaryInbound(m); !ok || len(refs) == 0 {
return dedupRefs(refs)
}
for _, r := range rules {
if !r.Enabled {
continue
}
for _, raw := range r.Src {
s := strings.TrimSpace(raw)
switch {
case strings.HasPrefix(s, "iface:"):
name := strings.TrimPrefix(s, "iface:")
refs = append(refs, divertRef{Src: name, Dev: IfaceDevice(name)})
case strings.HasPrefix(s, "zone:"):
zone := strings.TrimPrefix(s, "zone:")
for _, dev := range nftZoneDevices(zone) {
refs = append(refs, divertRef{Src: "zone:" + zone, Dev: dev})
}
}
}
}
return dedupRefs(refs)
}
// dedupRefs drops references whose device was already seen (and blank devices),
// preserving order.
func dedupRefs(in []divertRef) []divertRef {
seen := map[string]bool{}
var out []divertRef
for _, r := range in {
dev := strings.TrimSpace(r.Dev)
if dev == "" || seen[dev] {
continue
}
seen[dev] = true
r.Dev = dev
out = append(out, r)
}
return out
}
// nftDivertDevs is nftDivertRefs projected onto the device names, which is what
// every renderer and the sysctl/fail-closed scoping consume.
func nftDivertDevs(m *model.Model, rules []model.Rule) []string {
refs := nftDivertRefs(m, rules)
devs := make([]string, 0, len(refs))
for _, r := range refs {
devs = append(devs, r.Dev)
}
return nftDedupStr(devs)
}
// nftPrimaryInbound returns the first enabled tproxy inbound (port/mark source).
func nftPrimaryInbound(m *model.Model) (model.Inbound, bool) {
for _, in := range m.Inbounds {
if model.IsTproxyInbound(in) {
return in, true
}
}
return model.Inbound{}, false
}
// sortedRules returns a copy of in stably sorted by ascending Order (ported).
func sortedRules(in []model.Rule) []model.Rule {
rules := append([]model.Rule(nil), in...)
for i := 1; i < len(rules); i++ {
for j := i; j > 0 && rules[j-1].Order > rules[j].Order; j-- {
rules[j-1], rules[j] = rules[j], rules[j-1]
}
}
return rules
}
// RenderNft builds the `inet shater` tproxy table text (ported verbatim from the
// v0.1 apply.RenderNft, re-typed onto *model.Model; the ruleset skeleton, marks,
// sets, counters and chain order are unchanged).
//
// mgmt-bypass (unconditional, in order): loop-guard mark, per-egress marks,
// router-local dsts (fib daddr type local — SSH/LuCI/DNS to any router IP),
// RFC1918 + loopback + link-local + multicast/broadcast, and the IPv6 mirror.
// Only LAN-ingress (iifname) traffic past those bypasses is diverted, so
// router-originated egress is structurally never touched.
//
// Chains: prerouting (mangle) holds the tproxy diverts and only accept verbs; a
// separate forward (filter) chain holds everything needing reject/drop — the
// DoT/DoQ :853 reject and the IPv6 fail-closed drop when ipv6 is disabled with a
// closed kill-switch (reject is illegal in prerouting-hook chains). LAN :53 is
// NOT special-cased: the tproxy catch-all diverts it into the engine, where the
// hijack-dns route action answers it (D14) — no nft :53 redirect / dnsnat chain.
//
// RenderHoldNft renders the FAIL-CLOSED HOLDING PLANE: the ruleset that protects
// the LAN while the engine is NOT running.
//
// Why this exists. The forward-chain kill-switch in the full ruleset only works
// if the ruleset is loaded. When the engine fails to START — an unreachable
// remote rule-set at boot, a bad node, a busy port, out of memory — the apply
// aborts before the netplane stage and NO table is installed at all. After a
// reboot (the previous daemon's teardown removed the old table) that is not a
// window, it is a permanent state: `kill_switch=closed` is configured, the daemon
// is alive, the panel answers, and the router has silently become a plain
// OpenWrt box forwarding every LAN packet straight to the WAN in the clear. The
// operator is looking at a working UI and believes they are protected. A leak
// caused by the absence of the chain is still a leak.
//
// The trigger is entirely mundane: power cut, router up in 30s, ISP up in 60s,
// geo list not downloaded yet.
//
// What it does. ONE forward chain, no prerouting/tproxy (there is no engine
// socket to divert to), scoped to exactly the devices the real plan would divert:
//
// - router-origin and egress-marked traffic is accepted (the daemon must stay
// able to reach the internet, or it could never fetch the rule-set it choked
// on and would never recover);
// - LAN-to-LAN, link-local and ICMPv6 ND stay up, so the local network keeps
// working;
// - everything else from those devices is DROPPED.
//
// Management access is unaffected by construction: SSH, LuCI and the panel are
// INPUT-hook traffic to the router itself, and this chain only hooks forward. The
// operator can always get in to fix the config.
//
// Honest limitation: LAN clients can still reach the router's own resolver and
// dnsmasq will forward those queries to the ISP over the router's egress, so DNS
// is unfiltered while holding. Blocking it would also cut the daemon's own name
// resolution and with it any chance of self-recovery. No client TRAFFIC reaches
// the WAN — that is the promise being kept here.
//
// Returns "" when there is nothing to protect (no divert devices), which the
// caller treats as "no hold plane needed".
func RenderHoldNft(m *model.Model) (string, error) {
return RenderHoldNftAt(m, time.Now())
}
// RenderHoldNftAt is RenderHoldNft with the evaluation clock injected: the
// divert-device set it protects honours the active profile's rule overrides at
// now (nftPlanRules), exactly like the full plan's — everything the full plan
// would divert, the holding plane fails closed on.
func RenderHoldNftAt(m *model.Model, now time.Time) (string, error) {
validDevs, _ := nftSplitDevs(nftDivertDevs(m, nftPlanRules(m, now)))
iif := nftIifExpr(validDevs)
if iif == "" {
return "", nil
}
var sb strings.Builder
sb.WriteString("#!/usr/sbin/nft -f\n")
sb.WriteString("table inet shater\n")
sb.WriteString("delete table inet shater\n")
sb.WriteString("table inet shater {\n")
sb.WriteString("\tchain forward {\n")
sb.WriteString("\t\ttype filter hook forward priority filter; policy accept;\n")
sb.WriteString("\t\t# HOLDING PLANE: the engine is down; LAN->WAN is blocked (fail-closed).\n")
// Router-origin / egress-marked traffic survives, so the daemon can still
// reach the network and recover on its own.
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", loopMark))
for i, eg := range m.Egresses {
if t := strings.ToLower(eg.Type); t == "interface" || t == "tunnel" {
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", EgressMark(m.Globals, i)))
}
}
// Keep the LAN itself usable.
sb.WriteString("\t\t" + iif + " ip daddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16 } accept\n")
sb.WriteString("\t\t" + iif + " icmpv6 type { nd-router-solicit, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept\n")
sb.WriteString("\t\t" + iif + " ip6 daddr { ::1, fc00::/7, fe80::/10, ff00::/8 } accept\n")
// The untunnelable-protocol policy applies here too. These protocols were never
// going through the tunnel, engine or no engine, so honouring the operator's
// choice is the consistent answer; silently flipping them to blocked only while
// holding would make ping stop working for a reason nothing on screen explains.
// nil plan: the engine is down, so there is no routing knowledge to consult and
// every destination is treated as tunnelled. `direct` still lets everything out
// (it needs no knowledge); `block` and `icmp` degrade to their conservative form.
sb.WriteString(untunnelableRules(m.Globals, iif, nil))
// Everything else from the protected devices is dropped, both families —
// IPv6 unconditionally, because with no engine there is no v6 divert either.
sb.WriteString("\t\t" + iif + " meta nfproto ipv4 drop\n")
sb.WriteString("\t\t" + iif + " meta nfproto ipv6 drop\n")
sb.WriteString("\t}\n")
sb.WriteString("}\n")
return sb.String(), nil
}
// RenderNft renders the ruleset, discarding the non-fatal warnings. See
// RenderNftWithWarnings for the full contract.
func RenderNft(m *model.Model) (string, error) {
rs, _, err := RenderNftWithWarnings(m)
return rs, err
}
// RenderNftPlan renders with a destination-aware untunnelable plan (see
// untunnelable.go). A nil plan is the conservative reading: no destination is
// known to route direct, so nothing untunnelable is let out.
func RenderNftPlan(m *model.Model, plan *UntunnelablePlan) (string, []string, error) {
return renderNft(m, plan, time.Now())
}
// RenderNftPlanAt is RenderNftPlan with the evaluation clock injected. The
// apply pipeline passes the SAME timestamp it gave generate, so the per-rule
// diverts and the fail-closed coverage resolve profile/rule schedules against
// the identical instant the engine plan did.
func RenderNftPlanAt(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, []string, error) {
return renderNft(m, plan, now)
}
// RenderNftWithWarnings renders the ruleset and reports device names it had to
// reject (see nftValidDev), mirroring generate.GenerateWithWarnings so netplane
// stays free of a logger dependency and the daemon does the logging.
//
// REFUSAL CONTRACT. A rejected device name is fatal when the kill-switch is
// closed, and the ruleset is NOT rendered. The reasoning:
//
// - A rejected name means some interface we intend to divert cannot be named in
// the ruleset. Rendering anyway would produce a plan that diverts that
// interface's traffic (or not) while the forward-chain guard provably does not
// cover it — an apply that reports SUCCESS with fail-closed quietly not
// working for that interface. That path must not exist.
// - The alternative, rendering a "safe" plan that drops everything, would black
// out the LAN because of a name-parsing failure. Too destructive, and it
// misdiagnoses a config error as a security event.
// - Refusing is safe precisely because of the caller's fail-closed policy: the
// apply aborts BEFORE ApplyNft, so the PREVIOUS ruleset stays loaded and keeps
// protecting the LAN, while the panel surfaces a real, actionable error.
//
// With the kill-switch OPEN there is no fail-closed guard to undermine, so a
// rejected name is only a warning: the device drops out of the divert set (its
// traffic is not proxied, which open mode already tolerates) and the plan applies.
func RenderNftWithWarnings(m *model.Model) (string, []string, error) {
return renderNft(m, nil, time.Now())
}
// RenderNftWithWarningsAt is RenderNftWithWarnings with the evaluation clock
// injected (see RenderNftPlanAt).
func RenderNftWithWarningsAt(m *model.Model, now time.Time) (string, []string, error) {
return renderNft(m, nil, now)
}
func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, []string, error) {
tproxyMark := effFwmark(m.Globals)
inboundDevs := nftEnabledInboundDevs(m)
primary, hasPrimary := nftPrimaryInbound(m)
// The EFFECTIVE rules (active-profile enable/disable applied at now) drive
// BOTH the per-rule divert emission and the divert-device set below, so the
// two can never disagree with each other — or with the engine plan generate
// built from the same resolution at the same instant.
planRules := nftPlanRules(m, now)
// Validate every device the plan wants to divert BEFORE rendering a single
// line, so a name we cannot express never reaches the kernel as a rule that
// matches nothing. See the refusal contract in the doc comment.
divertRefs := nftDivertRefs(m, planRules)
divertDevs := make([]string, 0, len(divertRefs))
for _, r := range divertRefs {
divertDevs = append(divertDevs, r.Dev)
}
validDevs, rejectedDevs := nftSplitDevs(divertDevs)
var warnings []string
for _, bad := range rejectedDevs {
warnings = append(warnings, fmt.Sprintf(
"interface name %q is not a usable device name and was REJECTED: it is not covered by the "+
"fail-closed guard and its traffic is not diverted", bad))
}
// What this plan does NOT cover: a device name that resolves to nothing the
// kernel has, and a client network of this router that the plan never mentions.
// Both are silent leaks of exactly the kind the rejected-name check above
// exists for, and both were previously visible only in the panel's interface
// picker (or nowhere at all). Collected BEFORE the refusal below, so a refused
// plan still tells the operator everything that is wrong with it.
warnings = append(warnings, coverageWarnings(m, divertRefs)...)
if len(rejectedDevs) > 0 && genGlobalClosed(m.Globals) {
return "", warnings, fmt.Errorf(
"refusing to apply: %d interface name(s) cannot be expressed in the ruleset (%s), so the "+
"fail-closed kill-switch could not cover them; fix the interface/network names or set "+
"kill_switch=open to accept the leak explicitly",
len(rejectedDevs), strings.Join(rejectedDevs, ", "))
}
// Build the classification/divert lines into a buffer, collecting referenced
// counters so we can declare exactly those.
var body strings.Builder
used := map[string]bool{}
var order []string
useCounter := func(name string) string {
if name != "" && !used[name] {
used[name] = true
order = append(order, name)
}
return name
}
emit := func(iif, l4, match string, fam int, port int, mark uint32, counter string) {
var b strings.Builder
b.WriteString("\t\t")
if iif != "" {
b.WriteString(iif + " ")
}
b.WriteString("meta l4proto " + l4 + " ")
if match != "" {
b.WriteString(match + " ")
}
if fam == 6 {
b.WriteString("update @" + Client6 + " { ip6 saddr } ")
} else {
b.WriteString("update @" + Client4 + " { ip saddr } ")
}
if counter != "" {
b.WriteString("counter name \"" + counter + "\" ")
}
if fam == 6 {
b.WriteString(fmt.Sprintf("tproxy ip6 to :%d ", port))
} else {
b.WriteString(fmt.Sprintf("tproxy ip to :%d ", port))
}
b.WriteString(fmt.Sprintf("meta mark set 0x%x accept\n", mark))
body.WriteString(b.String())
}
// emitFrag expands one src fragment into the l4 x family lines it needs.
emitFrag := func(f nftFrag, l4 string, port int, mark uint32, counter string) {
iif := nftIifExpr(f.iif)
// A fragment that NAMED devices but whose names were all rejected must be
// dropped, not emitted unscoped: emit() omits the iifname when the
// expression is empty, which would turn an interface-scoped divert into a
// catch-all matching every ingress including WAN. (Closed kill-switch has
// already refused such a plan; this keeps open mode honest too.)
if len(f.iif) > 0 && iif == "" {
return
}
switch f.fam {
case 4:
emit(iif, l4, f.match, 4, port, mark, counter)
case 6:
if m.Globals.IPv6 {
emit(iif, l4, f.match, 6, port, mark, counter)
}
default: // both
emit(iif, l4, f.match, 4, port, mark, counter)
if m.Globals.IPv6 {
emit(iif, l4, f.match, 6, port, mark, counter)
}
}
}
// 1) Per-rule src classification (first-match order). Only rules with a src
// selector produce nft lines; dst/target decisions stay inside the engine.
// planRules, not m.Rules: the Enabled flag here is the PROFILE-EFFECTIVE
// one, matching the engine plan and the divert-device set above.
if hasPrimary && len(inboundDevs) > 0 {
for _, r := range planRules {
if !r.Enabled || len(r.Src) == 0 {
continue
}
frags := nftSrcFrags(r, inboundDevs)
if len(frags) == 0 {
continue
}
cnt := useCounter(nftRuleCounter(r))
for _, f := range frags {
if primary.TCP {
emitFrag(f, "tcp", primary.TproxyPort, tproxyMark, cnt)
}
if primary.UDP {
emitFrag(f, "udp", primary.TproxyPort, tproxyMark, cnt)
}
}
}
}
// 2) Per-inbound catch-all divert (all remaining LAN-ingress traffic). Only
// tproxy inbounds divert; local socks/http/dokodemo listeners must not
// inject a bogus LAN tproxy rule pointing at their listener port.
for _, in := range m.Inbounds {
if !model.IsTproxyInbound(in) {
continue
}
dev := IfaceDevice(in.Network)
iif := nftIifExpr([]string{dev})
if iif == "" {
// Unusable device name: skip the inbound entirely rather than emit an
// unscoped catch-all divert (see emitFrag). Already fatal under a closed
// kill-switch; under an open one the warning has been recorded.
continue
}
cnt := useCounter(nftInCounter(in))
if in.TCP {
emit(iif, "tcp", "", 4, in.TproxyPort, tproxyMark, cnt)
if m.Globals.IPv6 {
emit(iif, "tcp", "", 6, in.TproxyPort, tproxyMark, cnt)
}
}
if in.UDP {
emit(iif, "udp", "", 4, in.TproxyPort, tproxyMark, cnt)
if m.Globals.IPv6 {
emit(iif, "udp", "", 6, in.TproxyPort, tproxyMark, cnt)
}
}
}
// Assemble the table: sets + counter declarations + chains.
var sb strings.Builder
sb.WriteString("#!/usr/sbin/nft -f\n")
sb.WriteString("table inet shater\n")
sb.WriteString("delete table inet shater\n")
sb.WriteString("table inet shater {\n")
sb.WriteString("\tset " + Client4 + " { type ipv4_addr; flags dynamic; counter; }\n")
if m.Globals.IPv6 {
sb.WriteString("\tset " + Client6 + " { type ipv6_addr; flags dynamic; counter; }\n")
}
for _, c := range order {
sb.WriteString(fmt.Sprintf("\tcounter %s { }\n", c))
}
// Destination sets for the untunnelable policy. Rendered INTO the ruleset text,
// which is what makes a refreshed address list take effect: the applier's
// idempotence check compares this text, so new addresses produce different text
// and a real reload instead of a stale plan left loaded (the D3 trap).
sb.WriteString(untunnelableSets(plan))
sb.WriteString("\tchain prerouting {\n")
sb.WriteString("\t\ttype filter hook prerouting priority mangle; policy accept;\n")
// Ingress-iface match expr shared by the :53/:853 DNS rules and — critically —
// by the forward-chain fail-closed guard below. It covers EVERY device we
// divert (inbounds PLUS rule-referenced iface:/zone: devices), not just the
// tproxy inbounds: anything diverted must also be droppable when the engine is
// down, or a dead engine leaks that device's traffic straight to WAN. Hoisted
// here so the DNS force-intercept rule can sit ABOVE the fib-local bypass.
//
// Built from the PRE-VALIDATED device set: with the kill-switch closed we have
// already refused the plan if any name was rejected, so this expression is
// either complete or empty. Empty means nothing is diverted at all, and every
// rule below that depends on it emits nothing — correct, because there is then
// also nothing to fail closed on.
dnsIif := nftIifExpr(validDevs)
// --- mgmt-bypass (must precede any divert) ---
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", loopMark))
// Egress-marked traffic (interface/tunnel egresses) must never be re-diverted.
for i, eg := range m.Egresses {
if t := strings.ToLower(eg.Type); t == "interface" || t == "tunnel" {
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", EgressMark(m.Globals, i)))
}
}
// --- DNS force-intercept (Globals.DNSIntercept) ---
// Divert ALL LAN plaintext :53 into the engine, INCLUDING queries addressed to
// the router itself, which the fib-local bypass below would otherwise hand to
// dnsmasq. Placed ABOVE that bypass but BELOW the loop-guard/egress marks so the
// engine's own DNS egress is never re-diverted. tproxy preserves the client
// source IP, so per-device DNS rules still match. :853 (DoT/DoQ) stays carved
// out for the forward-chain reject; :443 (DoH) cannot be port-intercepted.
if m.Globals.DNSIntercept && hasPrimary && dnsIif != "" {
sb.WriteString(fmt.Sprintf("\t\t%s meta l4proto tcp th dport 53 update @%s { ip saddr } tproxy ip to :%d meta mark set 0x%x accept\n", dnsIif, Client4, primary.TproxyPort, tproxyMark))
sb.WriteString(fmt.Sprintf("\t\t%s meta l4proto udp th dport 53 update @%s { ip saddr } tproxy ip to :%d meta mark set 0x%x accept\n", dnsIif, Client4, primary.TproxyPort, tproxyMark))
if m.Globals.IPv6 {
sb.WriteString(fmt.Sprintf("\t\t%s meta l4proto tcp th dport 53 update @%s { ip6 saddr } tproxy ip6 to :%d meta mark set 0x%x accept\n", dnsIif, Client6, primary.TproxyPort, tproxyMark))
sb.WriteString(fmt.Sprintf("\t\t%s meta l4proto udp th dport 53 update @%s { ip6 saddr } tproxy ip6 to :%d meta mark set 0x%x accept\n", dnsIif, Client6, primary.TproxyPort, tproxyMark))
}
}
sb.WriteString("\t\tfib daddr type local accept\n")
sb.WriteString("\t\tip daddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16, 224.0.0.0/4, 255.255.255.255 } accept\n")
sb.WriteString("\t\tip6 daddr { ::1, fc00::/7, fe80::/10, ff00::/8 } accept\n")
// --- DNS anti-leak: only :853 is carved out; plain :53 is DIVERTED (D14) ---
if dnsIif != "" {
// Keep encrypted-DNS bypass ports (DoT :853 tcp, DoQ :853 udp) OUT of
// tproxy here so those packets traverse the forward chain, where they are
// REJECTED. The reject cannot live in this chain: the kernel refuses
// `reject` outside input/forward/output hooks. DoH (:443) is not
// port-blockable without an IP list; it is still SNI-routed by the proxy.
sb.WriteString("\t\t" + dnsIif + " meta l4proto { tcp, udp } th dport 853 accept\n")
// Plain :53 is deliberately NOT carved out: it falls through to the tproxy
// catch-all below and is diverted into the engine, where the hijack-dns
// route action answers it via the resolver's detour (the anti-leak). This
// also fails closed when the engine is down — no leak to a WAN resolver.
}
sb.WriteString(body.String())
sb.WriteString("\t}\n")
// --- forward: policy enforcement that needs a reject/drop verb. ---
ipv6Off := !m.Globals.IPv6 && dnsIif != ""
if dnsIif != "" {
sb.WriteString("\tchain forward {\n")
sb.WriteString("\t\ttype filter hook forward priority filter; policy accept;\n")
// DoT/DoQ hard-block (accepted past tproxy in prerouting above): clients
// fall back to plain :53, which the tproxy catch-all diverts into the
// engine, where the hijack-dns route action answers it (D14).
sb.WriteString("\t\t" + dnsIif + " meta l4proto { tcp, udp } th dport 853 reject\n")
switch {
case genGlobalClosed(m.Globals):
// PRIMARY FAIL-CLOSED DROP.
//
// Diverted LAN traffic is delivered LOCALLY via the tproxy socket (input
// hook) and NEVER traverses the forward hook. Therefore ANY LAN-ingress
// packet that appears in this forward chain bound for a PUBLIC (non-private)
// destination has ESCAPED the tproxy divert -- the engine's tproxy socket
// is down, so nftables `tproxy` returned NFT_BREAK, aborting its own rule
// before the trailing `meta mark set ... accept` could run. The packet is
// left UNMARKED, falls through prerouting `policy accept`, and would be
// forwarded to WAN (fw4 accept_to_wan -> masquerade). In closed mode we
// must DROP it here so a dead engine cannot leak plaintext to the WAN.
//
// We deliberately keep alive, IN THIS ORDER, everything that legitimately
// belongs in forward: router/egress-marked traffic, and LAN-to-LAN /
// link-local. Only genuinely public-bound LAN-ingress traffic is dropped.
// 1) mgmt / interface+tunnel egress marks: router-origin and egress
// traffic must never be dropped (mirror of the prerouting bypass).
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", loopMark))
for i, eg := range m.Egresses {
if t := strings.ToLower(eg.Type); t == "interface" || t == "tunnel" {
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", EgressMark(m.Globals, i)))
}
}
// 2) LAN-to-LAN / link-local (v4): the LAN itself must keep working
// (same private-range bypass sets the prerouting chain uses).
sb.WriteString("\t\t" + dnsIif + " ip daddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16 } accept\n")
// 3) IPv6 plane. Public v6 escapes the divert exactly like v4, so it is
// dropped below in BOTH cases; only the surviving accepts differ:
if m.Globals.IPv6 {
// v6 is diverted upstream; keep LAN-to-LAN / link-local v6 working.
sb.WriteString("\t\t" + dnsIif + " ip6 daddr { ::1, fc00::/7, fe80::/10, ff00::/8 } accept\n")
} else {
// ipv6 disabled: no v6 divert is emitted, so dual-stack LAN clients
// would silently BYPASS the proxy over IPv6. Keep the link-local plane
// (ND/RA ICMPv6, fe80::/10, multicast) alive so the LAN still works.
sb.WriteString("\t\t" + dnsIif + " icmpv6 type { nd-router-solicit, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept\n")
sb.WriteString("\t\t" + dnsIif + " ip6 daddr fe80::/10 accept\n")
sb.WriteString("\t\t" + dnsIif + " ip6 daddr ff00::/8 accept\n")
}
// 4) Untunnelable protocols (never TCP/UDP, so never divertible): the
// configured policy decides whether they leave directly or are dropped
// with everything else. Must precede the drops to have any effect.
sb.WriteString(untunnelableRules(m.Globals, dnsIif, plan))
// 5) Fail-closed drops for public-bound LAN-ingress traffic.
sb.WriteString("\t\t" + dnsIif + " meta nfproto ipv4 drop\n")
sb.WriteString("\t\t" + dnsIif + " meta nfproto ipv6 drop\n")
case ipv6Off:
// kill_switch=open + ipv6 disabled: IPv6 from the LAN bypasses the proxy
// undiverted and undropped (documented fail-open behavior).
sb.WriteString("\t\t# ipv6=0, kill_switch=open: LAN IPv6 bypasses the proxy (not dropped)\n")
}
sb.WriteString("\t}\n")
}
sb.WriteString("}\n")
return sb.String(), warnings, nil
}