The daemon's own log (engine + control-plane) went only to os.Stderr →
procd → the logread RAM ring: no file, no wall-clock timestamps, no size
cap, and "LogLevel=none" silenced ONLY the engine while the control-plane
kept writing at trace. So "download last day/3d/all", "limit the size" and
"fully turn it off" were all unmet.
New shater/logsink: one long-lived, atomically-reconfigurable Sink that
receives BOTH halves' byte streams, stamps every complete line with a UTC
RFC3339 wall clock (what makes date ranges real), and fans each line to a
size-capped 2-segment rotated file (ToFile) and/or the real os.Stderr
(ToSyslog). Both off = the line is dropped — the only true full silence.
Persistent path sits behind a stats-style disk-free guard (suspend+warn
once, auto-resume); tmpfs path is bounded by the cap itself. ANSI stripped
from the file copy only.
Wiring: control-plane via log.SetStdLogger over the sink; engine via a new
box.Options.DefaultLogWriter threaded into all three box.New sites
(apply/close-then-start/restore) by engine.SetDefaultLogWriter; live
reconfigure on every apply.Reconcile (SIGHUP / control socket / panel
apply) so panel changes take effect without a daemon restart.
controlLogLevel now makes the control-plane respect Globals.LogLevel
(silent vocab → panic-only; unknown → warn, mirroring generate).
Globals: LogToSyslog/LogToFile (default true), LogPersist (default false =
/var/log tmpfs; true = /etc/shater flash), LogMaxKB (default 2048, clamped
[128,8192]; 0 = default, not off — LogToFile is the off switch). UCI
parse/render/aliases + ValidateGlobals clamp-warn.
Endpoint GET /api/log?range=1d|3d|all (session-gated): streams the log line
by line, oldest segment first, filtered by the timestamp prefix; UTC
attachment filename. Honest fallbacks — file off + syslog on → a
"# note: … syslog ring only, ranges approximate" comment then a
`logread -e shater` scrape; both off → "# logging disabled". Unknown range
→ 400.
init.d: shater/shater-cron gate their `logger -t` status lines on
log_syslog so "logread off" is honest at the shell layer too.
Tests: logsink rotation-cap/timestamp/toggle-gating/engine→sink,
model round-trip + validate, endpoint session-gate/range/fallbacks.
VM-verified on QEMU (x86_64, OpenWrt 24.10): download+ranges, size-cap
rotation, file-off/full-off, persistent path, live reconfigure — all green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DNS events carry no client IP, so per-device DOMAIN stats need connection
events. box.go now builds+registers the trafficcontrol.Manager + AppendTracker
UNCONDITIONALLY (moved out of the needObservable gate) — an in-process
connection observable with NO clash/api port opened (Emit is non-blocking, so
an unsubscribed tracker never stalls the hot path). Engine.ConnManager()
exposes it (box-owned; pointer changes each Apply swap). stats connLoop
subscribes (pointer-identity resubscribe like dnsLoop), folding
{Source.Addr, Domain||Destination.Fqdn} into deviceDomains (bounded 512
clients / 200 domains-each). Snapshot gains device_domains
[{ip,name,domains:[{domain,count}]}]; Insights shows a per-device domain view.
Configurable retention: Globals StatsRingSize/StatsTimelineMinutes/
StatsMaxDomains/StatsRetentionDisabled (0=built-in defaults 200/60/5000);
stats.New resolves them, RetentionDisabled skips all pruning (RAM-bounded);
Settings gains a Statistics-retention section with a disable-trim toggle.
Verified: root+shater build (router tags)/vet 0, go test ok (conn-event fold,
retention, TestEngineConnManagerWired drives a real proxied conn + asserts a
live ConnectionEventNew + pointer-change-on-swap), VM box.New still Applies
with the tracker wired. panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Two nits surfaced by the lx-vs-upstream cleanliness audit (no runtime impact):
- box.go: the dnstrack registration comment said "service.FromContext" — the
§180 dead-stream signature. The actual readers use PtrFromContext (pairs with
MustRegisterPtr). Fixed the comment + noted why FromContext[*T] returns nil,
so a future debugger doesn't "fix" the readers back into §180.
- common/dnstrack/manager.go: removed the unused SourceRejected constant —
rejected resolutions are folded into SourceFailed at the emit site, so
"rejected" never reaches the wire. Replaced with a comment to prevent re-adding
an unreachable client case.
Audit verdict: code clean — no concurrency/wire/behaviour issues; dns/client.go
byte-identical to upstream, emits additive and subscriber-gated.
Hijacked DNS (the norm on an Android VPN) is answered before a connection becomes
a traffic tracker, so DNS queries never reach the connections stream — the only
egress was the text log, which carries no app attribution. Add common/dnstrack
(a Subscriber[QueryEvent] mirror of trafficcontrol) emitting one event per
resolution from dns/client.go, attributed via adapter.ContextFrom(ctx).ProcessInfo
(same ctx on cache-hit and miss, so cached queries are attributed too).
Failures are first-class: timeout/loopback/rejected-cached/SERVFAIL-reject emit
failed=true + error + rcode=-1 (no response) — without this the stream is blind to
DNS failures, the primary throttling signal. CNAME chains preserved: with
includeAnswers, each event carries the full response.Answer in wire order (CNAME
hops + final A/AAAA, not filtered to IPs).
Wire: rpc SubscribeDNSQueries(SubscribeDNSQueriesRequest) returns (stream
DnsQueryEvent) + DnsAnswer; event-driven server stream (no ticker); libbox
SubscribeDNSQueries(includeAnswers, handler). Tag-less core -> Unimplemented.
Detour/Chain and other streams unchanged.
Docs: SPECS/018, lx-changelog rc.7.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
Upstream box.go forced needClashAPI whenever PlatformLogWriter is set (always
on Android/libbox), because the Clash server was historically the only log/
traffic observer. With with_clash_api dropped (rc.1), that made every Android
start fatal: 'clash api is not included in this build' — even with no clash_api
in the config.
Split the concern behind a // lx: seam: PlatformLogWriter now requests
observability (Observable log factory + connection/traffic tracker), served by
the native CommandClient (SubscribeLog/SubscribeConnections), NOT the Clash
server. Only an explicit experimental.clash_api block still creates the Clash
server (and still fails fast without the tag). daemon is already nil-safe to a
missing clashServer, so Clash-mode degrades gracefully. Desktop unaffected.
Verified: core starts with no clash_api config; still fail-fast with one.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
We mistakenly believed that `libresolv`'s `search` function worked correctly in NetworkExtension, but it seems only `getaddrinfo` does.
This commit changes the behavior of the `local` DNS server in NetworkExtension to prefer DHCP, falling back to `getaddrinfo` if DHCP servers are unavailable.
It's worth noting that `prefer_go` does not disable DHCP since it respects Dial Fields, but `getaddrinfo` does the opposite. The new behavior only applies to NetworkExtension, not to all scenarios (primarily command-line binaries) as it did previously.
In addition, this commit also improves the DHCP DNS server to use the same robust query logic as `local`.