release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.
Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.
With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.
Fix (upstream file, lx:tproxy_writeback_connect):
* CONNECT the write-back socket to the one peer it ever talks to. The kernel's
compute_score() rejects a connected socket for any other peer, and a
connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
reuseport selection outright — so it can no longer be handed a datagram it
will not read. Nothing about the reply changes: same spoofed source, same
single peer, Write instead of WriteTo. The unconnected path is kept verbatim
for a destination that cannot be bound (domain socksaddr).
* A failed cached write now CLOSES the socket instead of only dropping the
reference (upstream left the fd to the GC finalizer).
* TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
but the cache evicts lazily, so after the inbound is gone nothing wakes the
live sessions and each strands its write-back socket. Invisible upstream
(one close at shutdown); on this fork the engine is rebuilt on every apply,
so it was one stranded generation per apply.
Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.
The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.
Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
245 lines
9.2 KiB
Go
245 lines
9.2 KiB
Go
package redirect
|
|
|
|
import (
|
|
"context"
|
|
"net"
|
|
"net/netip"
|
|
"time"
|
|
|
|
"github.com/sagernet/sing-box/adapter"
|
|
"github.com/sagernet/sing-box/adapter/inbound"
|
|
"github.com/sagernet/sing-box/common/listener"
|
|
"github.com/sagernet/sing-box/common/redir"
|
|
C "github.com/sagernet/sing-box/constant"
|
|
"github.com/sagernet/sing-box/log"
|
|
"github.com/sagernet/sing-box/option"
|
|
"github.com/sagernet/sing/common"
|
|
"github.com/sagernet/sing/common/buf"
|
|
"github.com/sagernet/sing/common/control"
|
|
E "github.com/sagernet/sing/common/exceptions" // lx:tproxy_writeback_connect
|
|
M "github.com/sagernet/sing/common/metadata"
|
|
N "github.com/sagernet/sing/common/network"
|
|
"github.com/sagernet/sing/common/udpnat2"
|
|
)
|
|
|
|
func RegisterTProxy(registry *inbound.Registry) {
|
|
inbound.Register[option.TProxyInboundOptions](registry, C.TypeTProxy, NewTProxy)
|
|
}
|
|
|
|
type TProxy struct {
|
|
inbound.Adapter
|
|
ctx context.Context
|
|
router adapter.Router
|
|
logger log.ContextLogger
|
|
listener *listener.Listener
|
|
udpNat *udpnat.Service
|
|
}
|
|
|
|
func NewTProxy(ctx context.Context, router adapter.Router, logger log.ContextLogger, tag string, options option.TProxyInboundOptions) (adapter.Inbound, error) {
|
|
tproxy := &TProxy{
|
|
Adapter: inbound.NewAdapter(C.TypeTProxy, tag),
|
|
ctx: ctx,
|
|
router: router,
|
|
logger: logger,
|
|
}
|
|
var udpTimeout time.Duration
|
|
if options.UDPTimeout != 0 {
|
|
udpTimeout = time.Duration(options.UDPTimeout)
|
|
} else {
|
|
udpTimeout = C.UDPTimeout
|
|
}
|
|
tproxy.udpNat = udpnat.New(tproxy, tproxy.preparePacketConnection, udpTimeout, false)
|
|
tproxy.listener = listener.New(listener.Options{
|
|
Context: ctx,
|
|
Logger: logger,
|
|
Network: options.Network.Build(),
|
|
Listen: options.ListenOptions,
|
|
ConnectionHandler: tproxy,
|
|
OOBPacketHandler: tproxy,
|
|
TProxy: true,
|
|
})
|
|
return tproxy, nil
|
|
}
|
|
|
|
func (t *TProxy) Start(stage adapter.StartStage) error {
|
|
if stage != adapter.StartStateStart {
|
|
return nil
|
|
}
|
|
return t.listener.Start()
|
|
}
|
|
|
|
func (t *TProxy) Close() error {
|
|
err := t.listener.Close()
|
|
// lx:begin tproxy_writeback_connect
|
|
// Closing the listener stops INGRESS but leaves every live UDP NAT session in
|
|
// the cache, and each session holds a write-back socket bound to its original
|
|
// destination. Nothing else ever wakes those sessions: the cache evicts
|
|
// lazily (on the next Get/Add), and after this inbound is gone there is no
|
|
// next Get. For a DNS session answered in-engine by hijack-dns there is not
|
|
// even an outbound connection whose failure could unwind it, so its socket
|
|
// survives until a GC finalizer happens to reach it.
|
|
//
|
|
// That is invisible upstream, where an inbound is closed once at shutdown. On
|
|
// this fork the engine is REBUILT on every config apply, so each apply would
|
|
// strand another generation of sockets on the router's own LAN :53. Purge
|
|
// evicts every session; the cache's OnEvict closes the conn, which unblocks
|
|
// the session's routing goroutine and runs its onClose — the one place the
|
|
// write-back socket is actually closed. Purge AFTER the listener so a packet
|
|
// arriving mid-teardown cannot re-create a session behind us.
|
|
t.udpNat.Purge()
|
|
// lx:end tproxy_writeback_connect
|
|
return err
|
|
}
|
|
|
|
func (t *TProxy) NewConnection(ctx context.Context, conn net.Conn, metadata adapter.InboundContext, onClose N.CloseHandlerFunc) {
|
|
metadata.Inbound = t.Tag()
|
|
metadata.InboundType = t.Type()
|
|
metadata.Destination = M.SocksaddrFromNet(conn.LocalAddr()).Unwrap()
|
|
t.logger.InfoContext(ctx, "inbound connection to ", metadata.Destination)
|
|
t.router.RouteConnectionEx(ctx, conn, metadata, onClose)
|
|
}
|
|
|
|
func (t *TProxy) NewPacketConnectionEx(ctx context.Context, conn N.PacketConn, source M.Socksaddr, destination M.Socksaddr, onClose N.CloseHandlerFunc) {
|
|
t.logger.InfoContext(ctx, "inbound packet connection from ", source)
|
|
t.logger.InfoContext(ctx, "inbound packet connection to ", destination)
|
|
var metadata adapter.InboundContext
|
|
metadata.Inbound = t.Tag()
|
|
metadata.InboundType = t.Type()
|
|
metadata.Source = source
|
|
metadata.Destination = destination
|
|
metadata.OriginDestination = t.listener.UDPAddr()
|
|
t.router.RoutePacketConnectionEx(ctx, conn, metadata, onClose)
|
|
}
|
|
|
|
func (t *TProxy) NewPacket(buffer *buf.Buffer, oob []byte, source M.Socksaddr) {
|
|
destination, err := redir.GetOriginalDestinationFromOOB(oob)
|
|
if err != nil {
|
|
t.logger.Warn("process packet from ", source, ": get tproxy destination: ", err)
|
|
return
|
|
}
|
|
t.udpNat.NewPacket([][]byte{buffer.Bytes()}, source, M.SocksaddrFromNetIP(destination), nil)
|
|
}
|
|
|
|
func (t *TProxy) preparePacketConnection(source M.Socksaddr, destination M.Socksaddr, userData any) (bool, context.Context, N.PacketWriter, N.CloseHandlerFunc) {
|
|
ctx := log.ContextWithNewID(t.ctx)
|
|
writer := &tproxyPacketWriter{
|
|
ctx: ctx,
|
|
listener: t.listener,
|
|
source: source.AddrPort(),
|
|
destination: destination,
|
|
}
|
|
return true, ctx, writer, func(it error) {
|
|
common.Close(common.PtrOrNil(writer.conn))
|
|
}
|
|
}
|
|
|
|
type tproxyPacketWriter struct {
|
|
ctx context.Context
|
|
listener *listener.Listener
|
|
source netip.AddrPort
|
|
destination M.Socksaddr
|
|
conn *net.UDPConn
|
|
}
|
|
|
|
// lx:begin tproxy_writeback_connect
|
|
//
|
|
// The TPROXY UDP write-back socket is bound to the ORIGINAL DESTINATION, so the
|
|
// client sees the reply coming from the address it addressed. Upstream leaves
|
|
// that socket UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort), and that is
|
|
// the bug this block exists for.
|
|
//
|
|
// An unconnected socket bound to <addr>:<port> is, as far as the kernel is
|
|
// concerned, a RECEIVER for that address:port — and because the socket also
|
|
// carries SO_REUSEADDR it silently joins the UDP demultiplex set of whatever
|
|
// else is bound there. Nothing ever reads from it: this writer only sends. So
|
|
// every datagram the kernel happens to hand it is lost.
|
|
//
|
|
// On a router that transparently intercepts LAN DNS towards its OWN address
|
|
// (`nft ... udp dport 53 tproxy ...`) the original destination IS the router's
|
|
// LAN address, so each intercepted DNS session parks another silent receiver on
|
|
// <router-lan-ip>:53 right next to dnsmasq's socket — observed on the stand:
|
|
// 33 such sockets against dnsmasq's one, several with a growing Recv-Q. The
|
|
// host's own queries to that address take the loopback path, are never diverted,
|
|
// and are therefore demultiplexed among all of them: they land in one of the
|
|
// silent sockets at random and time out, permanently and unpredictably, while
|
|
// every other DNS path on the box keeps working.
|
|
//
|
|
// Connecting the socket fixes it at the root. compute_score() in the kernel's
|
|
// UDP lookup REJECTS a connected socket for any peer other than the connected
|
|
// one, so a write-back socket can no longer be handed a datagram it will not
|
|
// read. Nothing about the reply changes — same spoofed source address, same
|
|
// single peer, one Write instead of one WriteTo.
|
|
//
|
|
// The unconnected path is kept verbatim for a destination that cannot be bound
|
|
// (a domain socksaddr), so no existing case regresses.
|
|
|
|
// newTProxyWriteBack creates the connected write-back socket: local address =
|
|
// the original destination (transparent bind), peer = the client. A package
|
|
// variable so the behaviour can be tested without CAP_NET_ADMIN.
|
|
var newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
|
|
var dialer net.Dialer
|
|
dialer.LocalAddr = destination.UDPAddr()
|
|
dialer.Control = control.Append(dialer.Control, control.ReuseAddr())
|
|
dialer.Control = control.Append(dialer.Control, redir.TProxyWriteBack())
|
|
conn, err := w.listener.DialContext(dialer, w.ctx, "udp", w.source.String())
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
udpConn, loaded := conn.(*net.UDPConn)
|
|
if !loaded {
|
|
conn.Close()
|
|
return nil, E.New("tproxy write back: unexpected connection type ", conn)
|
|
}
|
|
return udpConn, nil
|
|
}
|
|
|
|
// lx:end tproxy_writeback_connect
|
|
|
|
func (w *tproxyPacketWriter) WritePacket(buffer *buf.Buffer, destination M.Socksaddr) error {
|
|
defer buffer.Release()
|
|
if w.listener.ListenOptions().NetNs == "" {
|
|
conn := w.conn
|
|
if w.destination == destination && conn != nil {
|
|
// lx:begin tproxy_writeback_connect
|
|
// The cached socket is CONNECTED to the client, so this is a plain
|
|
// Write. A failed write also closes it: upstream only dropped the
|
|
// reference, leaving the fd to the GC finalizer.
|
|
_, err := conn.Write(buffer.Bytes())
|
|
if err != nil {
|
|
conn.Close()
|
|
w.conn = nil
|
|
}
|
|
return err
|
|
// lx:end tproxy_writeback_connect
|
|
}
|
|
}
|
|
// lx:begin tproxy_writeback_connect
|
|
if destination.IsIP() {
|
|
udpConn, err := newTProxyWriteBack(w, destination)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
if w.listener.ListenOptions().NetNs == "" && w.destination == destination {
|
|
w.conn = udpConn
|
|
} else {
|
|
defer udpConn.Close()
|
|
}
|
|
return common.Error(udpConn.Write(buffer.Bytes()))
|
|
}
|
|
// lx:end tproxy_writeback_connect
|
|
var listenConfig net.ListenConfig
|
|
listenConfig.Control = control.Append(listenConfig.Control, control.ReuseAddr())
|
|
listenConfig.Control = control.Append(listenConfig.Control, redir.TProxyWriteBack())
|
|
packetConn, err := w.listener.ListenPacket(listenConfig, w.ctx, "udp", destination.String())
|
|
if err != nil {
|
|
return err
|
|
}
|
|
udpConn := packetConn.(*net.UDPConn)
|
|
if w.listener.ListenOptions().NetNs == "" && w.destination == destination {
|
|
w.conn = udpConn
|
|
} else {
|
|
defer udpConn.Close()
|
|
}
|
|
return common.Error(udpConn.WriteToUDPAddrPort(buffer.Bytes(), w.source))
|
|
}
|