Run owner jobs in parallel with atomic reservations and harden recovery
Worker - Per-owner max_active_attempts (default 1, <= concurrent) replaces the one-active-attempt-per-owner claim rule; claims fill free slots each tick in deterministic created_at/profile_name/id order and record claimed_at. - Disk, memory and artifact reservations are admitted and inserted in one owner-locked transaction; clones reserve measured source disk bytes and source RAM plus overhead; pending clone storage and aggregate artifact headroom are accounted. - nextid lock is released right after the allocation reservation; source lock waits follow the caller context (20 min ceiling) with specific text. - Shared settleJobSQL for worker and watchdog; cancellation is reported as cancelled instead of failed/error. - Safety deadline is refreshed when collection starts so the watchdog cannot take over a live collection. - Evidence release continues past failures, keeps the original error and stops retrying until an admin re-requests. - Job ISO media is reconciled after all attempts finish (fixes leaks and the parallel-finish race); readiness errors carry the last guest/transport cause and distinguish cancellation/lease loss. - Watchdog prunes expired sessions, login limits and old orphan events; worker health events are throttled. PVE client - Bounded retries for idempotent GETs on transport errors/5xx only; task polling starts at 250 ms; task exitstatus is logged for operators; QGA file reads use 1 MiB chunks. API/store - Indexes for hot job/claim/quota queries; claimed_at column. - Native UUID path comparisons with early 404; deterministic run ordering; queue_ahead; admin held-attempts listing; stale revision publish guard; retry refused while any attempt is unfinished. - Uploads stream outside the DB transaction with a short locked quota recheck; explicit pool sizes; `otche healthcheck` CLI mode. Recorder - Network reads and FFmpeg writes happen outside the framebuffer lock; wall-clock frame pacing with bounded catch-up and explicit gaps; fsync at most once per second after the first durable fragment. Windows runner - UTC-safe timestamp parsing everywhere; readiness checks desktop first and computes the baseline once; per-nonce probe results; collector publishes initial telemetry immediately and reads event logs incrementally. - Baseline records UAC, Smart App Control, Secure Boot and Device Guard registry state; installer stores the LSA secret last and gains read-only -Verify and signed -UpdateRunner maintenance modes.
This commit is contained in:
@@ -15,9 +15,9 @@ docker build --target api -t otche-api:local .
|
||||
docker build --target worker -t otche-worker:local .
|
||||
```
|
||||
|
||||
Use Go 1.26 (the module minimum is in `go.mod`). Docker builds use Go 1.26/Bookworm and Debian Trixie runtime; worker adds FFmpeg and xorriso. `internal/store/schema.sql` is embedded in the binary. Runtime user is UID/GID 10001. No private key or operator credential is included.
|
||||
The module minimum is Go 1.24 (`go 1.24.0` in `go.mod`); Docker builds use Go 1.26/Bookworm and Debian Trixie runtime. Worker adds FFmpeg and xorriso. `internal/store/schema.sql` is embedded in the binary. Runtime user is UID/GID 10001. No private key or operator credential is included.
|
||||
|
||||
The command modes are `api`, `worker`, `watchdog`, `migrate`, `user-create`, `bindings-sync`, `source-maintenance` and `source-validate`. See [configuration](docs/CONFIG.md) and [API contract](docs/API.md). `DATABASE_URL_FILE` is required for all database commands. API additionally uses `PUBLIC_ORIGIN`, `ARTIFACT_ROOT`, optional `LISTEN_ADDR` (default :8080), and explicitly opt-in `ALLOW_INSECURE_HTTP=true` only for loopback development. Run `migrate` before the API. Use the sibling deploy repository for a complete database-backed startup and secure first-user creation.
|
||||
The command modes are `api`, `worker`, `watchdog`, `healthcheck`, `migrate`, `user-create`, `bindings-sync`, `source-maintenance` and `source-validate`. See [configuration](docs/CONFIG.md) and [API contract](docs/API.md). `DATABASE_URL_FILE` is required for all database commands. `otche healthcheck` needs no database credentials: it requests `http://127.0.0.1:<LISTEN_ADDR port>/api/v1/health` (default 8080) with a two-second timeout and exits successfully only for HTTP 200. API additionally uses `PUBLIC_ORIGIN`, `ARTIFACT_ROOT`, optional `LISTEN_ADDR` (default :8080), and explicitly opt-in `ALLOW_INSECURE_HTTP=true` for trusted development only. Startup warns when `PUBLIC_ORIGIN` is unset or insecure HTTP is enabled on a non-loopback listener; neither warning refuses container-network development. Run `migrate` before the API. Use the sibling deploy repository for a complete database-backed startup and secure first-user creation.
|
||||
|
||||
## Verification
|
||||
|
||||
|
||||
+32
-2
@@ -7,6 +7,7 @@ import (
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net"
|
||||
"net/http"
|
||||
"os"
|
||||
"os/signal"
|
||||
@@ -34,7 +35,7 @@ func env(name, fallback string) string {
|
||||
}
|
||||
func run() error {
|
||||
if len(os.Args) < 2 {
|
||||
return errors.New("usage: otche api|worker|watchdog|migrate|user-create|source-maintenance|source-validate|bindings-sync")
|
||||
return errors.New("usage: otche api|worker|watchdog|healthcheck|migrate|user-create|source-maintenance|source-validate|bindings-sync")
|
||||
}
|
||||
mode := os.Args[1]
|
||||
flags := flag.NewFlagSet(mode, flag.ContinueOnError)
|
||||
@@ -49,11 +50,40 @@ func run() error {
|
||||
}
|
||||
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
|
||||
defer stop()
|
||||
if mode == "healthcheck" {
|
||||
_, port, err := net.SplitHostPort(env("LISTEN_ADDR", ":8080"))
|
||||
if err != nil || port == "" {
|
||||
return errors.New("healthcheck requires LISTEN_ADDR with a port")
|
||||
}
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "http://127.0.0.1:"+port+"/api/v1/health", nil)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
client := &http.Client{Timeout: 2 * time.Second, CheckRedirect: func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse }}
|
||||
resp, err := client.Do(req)
|
||||
if err != nil {
|
||||
return fmt.Errorf("healthcheck failed: %w", err)
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return fmt.Errorf("healthcheck returned HTTP %d", resp.StatusCode)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
url, err := store.Secret("DATABASE_URL")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
db, err := store.Open(ctx, url)
|
||||
maxConns := int32(4)
|
||||
switch mode {
|
||||
case "api":
|
||||
maxConns = 16
|
||||
case "worker":
|
||||
maxConns = 32
|
||||
case "watchdog":
|
||||
maxConns = 6
|
||||
}
|
||||
db, err := store.Open(ctx, url, maxConns)
|
||||
if err != nil {
|
||||
return errors.New("database connection failed; check DATABASE_URL secret and server")
|
||||
}
|
||||
|
||||
+13
-6
@@ -2,6 +2,8 @@
|
||||
|
||||
Base `/api/v1`; JSON snake_case; RFC3339 UTC dates; UUID IDs; errors `{ "error": { "code": "invalid_request", "message": "..." } }`. Lists `{ "items": [...] }`. Same origin only; browser fetch `credentials: include`. Session cookie HttpOnly SameSite=Strict Secure (local development flag permits HTTP). `GET /auth/session` returns `{user:{id,username,role},csrf_token}` or 401; `POST /auth/login` `{username,password}` returns same shape and cookie. All authenticated writes require `X-CSRF-Token` plus matching Origin. `POST /auth/logout` returns 204. Roles `admin|operator`. Admin reads all jobs; operators only own jobs. No browser PVE secrets.
|
||||
|
||||
Malformed `{id}` path parameters return HTTP 404 `not_found` before database lookup.
|
||||
|
||||
## Job submission and settings
|
||||
|
||||
`POST /uploads` body raw bytes (`application/octet-stream`), filename in `X-Filename` (encodeURIComponent). Response 201 `{id,filename,size,sha256,created_at}`. Original bytes immutable. No multipart. `GET /uploads/{id}/content` authorized attachment.
|
||||
@@ -14,15 +16,19 @@ Each Attempt with Grub produces one private `kind: "grub_archive"`, `filename: "
|
||||
|
||||
Readable live logs are captured as bounded **best-effort open-length snapshots**, not atomic filesystem snapshots. Writers that allow reads are supported; growth is not chased. Detected length/content/write-time changes yield `changed`, preserve the captured bytes when possible and make the archive `partial`. A sharing-denied file may have no snapshot. Limits: per file `min(8 MiB,max_artifact_bytes/2)`, total captured bytes `min(32 MiB,max_artifact_bytes/2)`; existing collection deadline and owner quota apply. No guest-supplied ZIP is trusted or imported.
|
||||
|
||||
`GET /jobs?page=1&page_size=25&q=&status=` -> `{items:Job[],total,page,page_size,next_page:null|number}`; `GET /jobs/{id}` Job with runs; `POST /jobs/{id}/cancel` -> Job; `POST /jobs/{id}/retry` -> same Job with a new Attempt for each terminal Run, preserving every prior Attempt and original immutable settings/revisions, admission rechecked. Active Jobs cannot retry. Job `{id,owner_id,upload_id,filename,execution_filename,sha256,size,settings,status,created_at,updated_at,cancel_requested,runs:Run[]}`. Status `queued|running|completed|cancelled|failed`.
|
||||
`GET /jobs?page=1&page_size=25&q=&status=` -> `{items:Job[],total,page,page_size,next_page:null|number}`; `GET /jobs/{id}` Job with runs; `POST /jobs/{id}/cancel` -> Job; `POST /jobs/{id}/retry` -> same Job with a new Attempt for every Run (including previously successful Runs), preserving every prior Attempt and original immutable settings/revisions, admission rechecked. Retry returns HTTP 409 `job_active` if the Job is queued/running **or any of its Attempts is unfinished**, even if the stored Job status is terminal. Job `{id,owner_id,upload_id,filename,execution_filename,sha256,size,settings,status,created_at,updated_at,cancel_requested,runs:Run[]}`. Status `queued|running|completed|cancelled|failed`.
|
||||
|
||||
Run `{id,job_id,profile_id,profile_name,revision_id,antivirus:"defender",status,attempts:Attempt[]}`; status `queued|running|completed|cancelled|failed`. `profile_name` is the fixed full name captured when that Run was created, including any manually entered build text; retries preserve it.
|
||||
Attempt `{id,run_id,command_id,phase,outcome,findings,telemetry,cleanup,error,created_at,started_at,finished_at,deadline_at,allocation,report,artifacts:Artifact[]}`.
|
||||
Job detail additionally returns `queue_ahead: number|null`: the number of queued, never-claimed Attempts from **other Jobs** created before this Job's earliest queued, never-claimed Attempt. The count includes all owners for both operators and admins; it exposes no other Job/owner IDs. It is null when this Job has no queued, never-claimed Attempt. It counts Attempts, not Jobs, and is not a completion-time estimate.
|
||||
|
||||
Run `{id,job_id,profile_id,profile_name,revision_id,antivirus:"defender",status,attempts:Attempt[]}`; status `queued|running|completed|cancelled|failed`. `profile_name` is the fixed full name captured when that Run was created, including any manually entered build text; retries preserve it. Runs are ordered by `profile_name`, then `id` in detail and list/dashboard summaries; detail Attempts are ordered by `created_at` ascending.
|
||||
Attempt `{id,run_id,command_id,phase,outcome,findings,telemetry,cleanup,error,created_at,claimed_at,started_at,finished_at,deadline_at,allocation,report,video,artifacts:Artifact[]}`.
|
||||
|
||||
- phase: `queued|provisioning|booting|recording|delivering|preparing|running|collecting|stopping|cleanup|finished`.
|
||||
- outcome: `pending|executed|blocked_before_execution|incompatible|policy_blocked|delivery_error|interrupted|cancelled|error`.
|
||||
- findings: `unknown|detected|not_observed`; telemetry: `pending|complete|partial|unavailable`; cleanup: `pending|complete|evidence_held|failed`.
|
||||
- Timestamps are null when not applicable. Early quarantine has no actual start/PID/duration; unknown session is null, **not Session 0**.
|
||||
- `claimed_at` is null until the first worker claim, then remains the first claim timestamp; it is distinct from actual sample `started_at` and includes provisioning/readiness overhead.
|
||||
- `video` is the same object returned by `GET /attempts/{id}/video`: `{state:"pending"|"recording"|"complete"|"partial"|"unavailable",segments:Artifact[],gaps:[{at,reason}],started_at,finished_at}`. Clients polling Job detail need not separately poll every Attempt's video endpoint.
|
||||
- allocation is null for operators; admins receive `{id,node,vmid,state}` diagnostics only, never PVE URLs or credentials.
|
||||
- report is null before collection. Partially recovered data may have null environment/Defender state; UI must show N/A, not infer successful execution.
|
||||
|
||||
@@ -94,7 +100,7 @@ type Detection = {
|
||||
|
||||
Missing or partial telemetry is not an empty clean report. Fixed collector fields, not arbitrary raw guest objects, populate the DTO.
|
||||
|
||||
`GET /jobs/{id}/events?after=N` -> `{items:[{id,job_id,attempt_id,kind,message,created_at}]}`; polling supported, no PVE WebSocket exposed. `GET /jobs/{id}/artifacts` -> list Artifact `{id,job_id,attempt_id,kind,filename,content_type,size,sha256,created_at,url}`. `GET /artifacts/{id}/content` owner-authorized, supports video Range, forces attachment except safe video. HTML always attachment + sandbox. `GET /attempts/{id}/video` -> `{state:"pending"|"recording"|"complete"|"partial"|"unavailable",segments:Artifact[],gaps:[{at,reason}],started_at,finished_at}`.
|
||||
`GET /jobs/{id}/events?after=N` -> `{items:[{id,job_id,attempt_id,kind,message,created_at}]}` ordered by ascending `id`, capped at **500 rows per response**. Continue with `after` set to the last returned ID until fewer than 500 rows are returned; polling supported, no PVE WebSocket exposed. `GET /jobs/{id}/artifacts` -> list Artifact `{id,job_id,attempt_id,kind,filename,content_type,size,sha256,created_at,url}`. `GET /artifacts/{id}/content` owner-authorized, supports video Range, forces attachment except safe video. HTML always attachment + sandbox. `GET /attempts/{id}/video` -> `{state:"pending"|"recording"|"complete"|"partial"|"unavailable",segments:Artifact[],gaps:[{at,reason}],started_at,finished_at}`.
|
||||
|
||||
## Profiles and administration
|
||||
|
||||
@@ -103,10 +109,11 @@ Missing or partial telemetry is not an empty clean report. Fixed collector field
|
||||
Windows names are entered manually, for example `Windows 11 Pro — build 26200.6584`. Neither qualification nor an OS update changes that label; only an explicit profile-name edit does. Interfaces render `Profile.name` and the immutable `Run.profile_name` verbatim. Actual `report.environment.os_build` remains technical telemetry for reports/fingerprints/drift, never a naming source. Migration freezes legacy Runs at their currently known profile name before subsequent edits; it cannot reconstruct names that were never stored and does not invent historical builds. Revisions retain separate identifiers and technical metadata, not generated build-qualified names.
|
||||
|
||||
Admin: `GET /admin/users`; `POST /admin/users` `{username,password,role}`; `PATCH /admin/users/{id}` `{role?,disabled?,password?}`. User `{id,username,role,disabled,created_at}`.
|
||||
`GET /admin/profiles`; `POST /admin/profiles` `{name,os,architecture}` -> Profile (maintenance, unqualified). `PATCH /admin/profiles/{id}` `{name?,enabled?}`. `POST /admin/profiles/{id}/maintenance` closes admission, returns profile plus drain status; never powers off live master. `POST /admin/profiles/{id}/publish` `{revision_id}` selects a previously worker-validated stopped source revision; never accepts VMID from browser. `POST /admin/profiles/{id}/qualify` queues qualification, returns `{id,status}`. Qualified revision configuration and credentials are worker-only CLI/file-backed, not browser secrets.
|
||||
`GET /admin/profiles`; `POST /admin/profiles` `{name,os,architecture}` -> Profile (maintenance, unqualified). `PATCH /admin/profiles/{id}` `{name?,enabled?}`. `POST /admin/profiles/{id}/maintenance` closes admission, returns profile plus drain status; never powers off live master. `POST /admin/profiles/{id}/publish` `{revision_id}` selects a previously worker-validated stopped source revision; never accepts VMID from browser. Returns HTTP 409 `unqualified_revision` without a passed qualification, `stale_revision` if a newer revision exists for the same source reference (including another profile), or `source_not_drained` while old-revision Attempts await cloning. `POST /admin/profiles/{id}/qualify` queues qualification, returns `{id,status}`. Qualified revision configuration and credentials are worker-only CLI/file-backed, not browser secrets.
|
||||
`GET /admin/bindings` -> list `{owner_id,pool,iso_storage,disk_storage,node,configured,reason,isolation_expires_at}` (metadata only). `isolation_expires_at` is nullable RFC3339: null means not validated; effective `configured` requires successful worker validation **and** a future expiry at request time. Expired/missing expiry closes dashboard readiness and `POST /jobs` admission even if a stale stored flag was true. Bindings are provisioned by operator CLI with file-backed credentials and genuine probe evidence, no secret edits/browser storage; UI shows actual expiry/prerequisites rather than a fake credential form.
|
||||
`GET /admin/health` -> `{database,worker,last_worker_seen,integration_ready,blockers:string[]}`.
|
||||
`POST /admin/attempts/{id}/release-evidence` `{confirm:true}` explicit irreversible authorized cleanup of owned retained disposable VM/disk only, never master/control.
|
||||
`GET /admin/attempts/held` -> `{items:[{attempt_id,job_id,run_id,profile_name,filename,outcome,cleanup,error,finished_at,release_requested,allocation:{id,node,vmid,state}|null}]}`. Admin-only; finished Attempts with `cleanup:"evidence_held"|"failed"`, newest `finished_at` first (nulls last), capped at 200. Includes entries without an allocation; no PVE URLs, tokens or credentials.
|
||||
`POST /admin/attempts/{id}/release-evidence` `{confirm:true}` -> HTTP 202 `{release_requested:true}` for retained Attempts with cleanup `evidence_held` or `failed` and an allocation. Explicit irreversible authorized cleanup of owned retained disposable VM/disk only, never master/control.
|
||||
`GET /dashboard` owner-scoped `{metrics:{jobs,queued,running,completed,failed,cancelled,detected},recent_jobs:Job[],queue:{queued,running},integration:{ready:boolean,blockers:string[]}}`.
|
||||
`GET /admin/profiles/{id}/revisions` -> `{items:[{id,profile_id,fingerprint,config_digest,qualification,created_at}]}`. Revision `qualification` is null or an object with `worker_validated:boolean`, `status:"passed"|"failed"`, `state:"qualified"|"unqualified"`, `baseline_fingerprint:string` and observed control details; fields absent before that stage must remain unknown. It is not the profile's string qualification enum.
|
||||
`GET /admin/profiles/{id}/qualifications` -> `{items:[{id,profile_id,status,result,created_at}]}`; `GET /admin/qualifications/{id}` same qualification object. Status `queued|running|passed|failed`; result actual controls/attempt IDs/errors only, null until available. Qualification bypasses *qualified-profile* admission only; still requires explicit configured stopped source and isolated disposable allocation, never puts controls into master or a real Run.
|
||||
|
||||
@@ -43,6 +43,20 @@ Binding metadata exposes `isolation_expires_at` (RFC3339 or null). API admission
|
||||
|
||||
Readiness now uses the signed fixed `Otche-DesktopReady` Limited/Interactive task, with a fresh nonce, bounded response and deadline, matching boot/account/console session, unlocked Default input desktop, real Explorer shell/taskbar and no visible setup/sign-in host. Resident UserOOBEBroker or generic WWAHost processes alone are not blockers. During a sample, the worker uses `-BootOnly`; post-run it uses `-BaselineOnly`, so sample-owned GUI is not reclassified as setup and no repeated interactive probe is launched. Capture and review your own protected source/fresh-clone evidence. Observed Defender and Windows versions are recorded, never inferred from marketing release names.
|
||||
|
||||
## Parallel execution
|
||||
|
||||
Top-level `concurrent` is the global live-attempt cap (1–16, default 1). Each owner binding has `max_active_attempts` (1–16, default 1), which must not exceed `concurrent`. A binding omitted from `owners` cannot claim work. Set both limits to 9 for nine simultaneous runs of one Job; leaving the owner limit at 1 deliberately preserves serial execution. Claims fill available slots each tick and order equal-age runs by their immutable `profile_name`, then run ID. A recovered lease preserves the original `claimed_at` timestamp.
|
||||
|
||||
Reservation admission and allocation insertion run in one transaction under the owner's lock. The VMID lock ends as soon as the reservation commits; only the source-specific lock spans the full clone and both source digest checks. Lock waits respect caller cancellation and have a 20-minute ceiling. Different sources can clone concurrently without weakening protected-VMID, ownership, isolation or qualification checks.
|
||||
|
||||
- **Disk:** `max_disk_bytes` is the source validation ceiling, not a charge per clone. Each Windows allocation reserves the measured sum of source virtual disks (including EFI/TPM). `max_owned_disk_bytes` includes every non-deleted allocation and extractor, including retained evidence. Storage admission additionally reserves all not-yet-materialized clones (`reserved`, `clone_intent`, `cloning`) against available bytes above `min_storage_free_bytes`; it does not subtract full logical disk capacity again for already-materialized thin clones. Thin provisioning is not a guarantee against later disk growth: size the underlying storage for the workload and retain free-space monitoring.
|
||||
- **Memory:** a clone reserves source `memory` MiB plus 1 GiB for QEMU overhead. Active owner reservations must fit `node total - min_memory_free_bytes - host baseline`; the baseline is node used memory minus measured `mem` of running VMs recorded in that owner's pool. Stopped/held evidence keeps its disk reservation but no running-memory reservation. Existing live free-memory checks remain mandatory.
|
||||
- **Artifacts:** owner quota includes inputs, retained artifacts, job media and non-deleted allocation reservations. Each new allocation reserves `max_video_bytes + 3 * max_artifact_bytes`, plus one artifact bucket when Grub is requested. The private artifact filesystem must keep `min_artifact_free_bytes` after **all** live artifact reservations, not just one encoder. The configured video ceiling must realistically cover the longest observation plus collection tail; measure actual recordings before reducing it.
|
||||
|
||||
For a 32-thread / approximately 125 GiB node, nine 8 GiB / 4-vCPU clones reserve approximately 81 GiB of memory (9 × (8 + 1)), leaving room for a measured host baseline and, for example, 12 GiB minimum free memory. Their 36 virtual CPUs are overcommitted against 32 hardware threads: benchmark payload and encoder contention rather than treating the cap as guaranteed throughput. Set `max_owned_disk_bytes >= 9 × measured source bytes + retained evidence headroom` (and extractor capacity when enabled). Nine measured 64 GiB sources need at least 576 GiB before EFI/TPM, extractors and retained evidence; a 160 GiB owner quota cannot admit them. Available clone storage must separately satisfy the pending-clone check.
|
||||
|
||||
For an illustrative `max_video_bytes = 512 MiB`, `max_artifact_bytes = 64 MiB` and no Grub, nine runs reserve 6.19 GiB of artifacts before existing uploads/media/evidence and the filesystem free-space floor. This video limit is an example, not a measured bound: long or high-motion captures may require more. Give the worker container enough CPU for nine ffmpeg processes; a 2-CPU container cap can bottleneck recording even if the Windows guests have ample node CPU. Adjust the deployment's `WORKER_CPUS`, memory and PID limits after measuring, and restart the worker/watchdog when changing runtime configuration. The shipped example remains conservative at one active attempt.
|
||||
|
||||
## Optional online policy
|
||||
|
||||
Under the owner binding set:
|
||||
|
||||
+25
-10
@@ -25,13 +25,13 @@ func (s *Server) adminRoutes() {
|
||||
m.HandleFunc("POST /api/v1/admin/profiles/{id}/publish", s.protected(true, s.publish))
|
||||
m.HandleFunc("POST /api/v1/admin/profiles/{id}/qualify", s.protected(true, s.qualify))
|
||||
m.HandleFunc("GET /api/v1/admin/profiles/{id}/revisions", s.protected(true, func(w http.ResponseWriter, r *http.Request) {
|
||||
s.list(w, r, `SELECT to_jsonb(v)-'source_ref' FROM revisions v WHERE profile_id::text=$1 ORDER BY created_at DESC`, r.PathValue("id"))
|
||||
s.list(w, r, `SELECT to_jsonb(v)-'source_ref' FROM revisions v WHERE profile_id=$1::uuid ORDER BY created_at DESC`, r.PathValue("id"))
|
||||
}))
|
||||
m.HandleFunc("GET /api/v1/admin/profiles/{id}/qualifications", s.protected(true, func(w http.ResponseWriter, r *http.Request) {
|
||||
s.list(w, r, `SELECT to_jsonb(q) FROM qualifications q WHERE profile_id::text=$1 ORDER BY created_at DESC`, r.PathValue("id"))
|
||||
s.list(w, r, `SELECT to_jsonb(q) FROM qualifications q WHERE profile_id=$1::uuid ORDER BY created_at DESC`, r.PathValue("id"))
|
||||
}))
|
||||
m.HandleFunc("GET /api/v1/admin/qualifications/{id}", s.protected(true, func(w http.ResponseWriter, r *http.Request) {
|
||||
v, e := s.queryObject(r.Context(), `SELECT to_jsonb(q) FROM qualifications q WHERE id::text=$1`, r.PathValue("id"))
|
||||
v, e := s.queryObject(r.Context(), `SELECT to_jsonb(q) FROM qualifications q WHERE id=$1::uuid`, r.PathValue("id"))
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -42,6 +42,7 @@ func (s *Server) adminRoutes() {
|
||||
s.list(w, r, `SELECT to_jsonb(b)||jsonb_build_object('configured',configured AND COALESCE(isolation_expires_at>now(),false),'reason',CASE WHEN configured AND NOT COALESCE(isolation_expires_at>now(),false) THEN 'Isolation proof expired or requires revalidation' ELSE reason END) FROM bindings b ORDER BY owner_id`)
|
||||
}))
|
||||
m.HandleFunc("GET /api/v1/admin/health", s.protected(true, s.health))
|
||||
m.HandleFunc("GET /api/v1/admin/attempts/held", s.protected(true, s.heldAttempts))
|
||||
m.HandleFunc("POST /api/v1/admin/attempts/{id}/release-evidence", s.protected(true, s.releaseEvidence))
|
||||
}
|
||||
func (s *Server) createUser(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -104,7 +105,7 @@ func (s *Server) updateUser(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
var id, role string
|
||||
var disabled bool
|
||||
e = tx.QueryRow(r.Context(), `SELECT id::text,role,disabled FROM users WHERE id::text=$1 FOR UPDATE`, r.PathValue("id")).Scan(&id, &role, &disabled)
|
||||
e = tx.QueryRow(r.Context(), `SELECT id::text,role,disabled FROM users WHERE id=$1::uuid FOR UPDATE`, r.PathValue("id")).Scan(&id, &role, &disabled)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -172,7 +173,7 @@ func (s *Server) updateProfile(w http.ResponseWriter, r *http.Request) {
|
||||
fail(w, 400, "invalid_profile", "Invalid name")
|
||||
return
|
||||
}
|
||||
v, e := s.queryObject(r.Context(), `UPDATE profiles SET name=COALESCE($2,name),enabled=COALESCE($3,enabled) WHERE id::text=$1 RETURNING to_jsonb(profiles)`, r.PathValue("id"), in.Name, in.Enabled)
|
||||
v, e := s.queryObject(r.Context(), `UPDATE profiles SET name=COALESCE($2,name),enabled=COALESCE($3,enabled) WHERE id=$1::uuid RETURNING to_jsonb(profiles)`, r.PathValue("id"), in.Name, in.Enabled)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -180,13 +181,13 @@ func (s *Server) updateProfile(w http.ResponseWriter, r *http.Request) {
|
||||
jsonResponse(w, 200, v)
|
||||
}
|
||||
func (s *Server) maintenance(w http.ResponseWriter, r *http.Request) {
|
||||
v, e := s.queryObject(r.Context(), `UPDATE profiles SET state='maintenance',enabled=false,reason='Maintenance admission closed; drain queued clones before changing source disks' WHERE id::text=$1 RETURNING to_jsonb(profiles)`, r.PathValue("id"))
|
||||
v, e := s.queryObject(r.Context(), `UPDATE profiles SET state='maintenance',enabled=false,reason='Maintenance admission closed; drain queued clones before changing source disks' WHERE id=$1::uuid RETURNING to_jsonb(profiles)`, r.PathValue("id"))
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
var pending int
|
||||
e = s.db.QueryRow(r.Context(), `SELECT count(*) FROM runs r JOIN attempts a ON a.run_id=r.id WHERE r.profile_id::text=$1 AND a.phase IN('queued','provisioning')`, r.PathValue("id")).Scan(&pending)
|
||||
e = s.db.QueryRow(r.Context(), `SELECT count(*) FROM runs r JOIN attempts a ON a.run_id=r.id WHERE r.profile_id=$1::uuid AND a.phase IN('queued','provisioning')`, r.PathValue("id")).Scan(&pending)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -212,7 +213,7 @@ func (s *Server) publish(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
defer tx.Rollback(r.Context())
|
||||
var profile string
|
||||
e = tx.QueryRow(r.Context(), `SELECT id::text FROM profiles WHERE id::text=$1 FOR UPDATE`, r.PathValue("id")).Scan(&profile)
|
||||
e = tx.QueryRow(r.Context(), `SELECT id::text FROM profiles WHERE id=$1::uuid FOR UPDATE`, r.PathValue("id")).Scan(&profile)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -223,6 +224,16 @@ func (s *Server) publish(w http.ResponseWriter, r *http.Request) {
|
||||
fail(w, 409, "unqualified_revision", "Worker-validated source and passed qualification required")
|
||||
return
|
||||
}
|
||||
var stale bool
|
||||
e = tx.QueryRow(r.Context(), `SELECT EXISTS(SELECT 1 FROM revisions newer WHERE newer.source_ref=v.source_ref AND newer.created_at>v.created_at) FROM revisions v WHERE v.id=$1`, in.RevisionID).Scan(&stale)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
if stale {
|
||||
fail(w, 409, "stale_revision", "A newer revision exists for this source")
|
||||
return
|
||||
}
|
||||
var pending int
|
||||
e = tx.QueryRow(r.Context(), `SELECT count(*) FROM runs r JOIN attempts a ON a.run_id=r.id WHERE r.profile_id=$1 AND r.revision_id<>$2 AND a.phase IN('queued','provisioning')`, profile, in.RevisionID).Scan(&pending)
|
||||
if e != nil {
|
||||
@@ -257,7 +268,7 @@ func (s *Server) qualify(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
defer tx.Rollback(r.Context())
|
||||
var id string
|
||||
e = tx.QueryRow(r.Context(), `SELECT id::text FROM profiles WHERE id::text=$1 FOR UPDATE`, r.PathValue("id")).Scan(&id)
|
||||
e = tx.QueryRow(r.Context(), `SELECT id::text FROM profiles WHERE id=$1::uuid FOR UPDATE`, r.PathValue("id")).Scan(&id)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -324,7 +335,7 @@ func (s *Server) releaseEvidence(w http.ResponseWriter, r *http.Request) {
|
||||
fail(w, 400, "confirmation_required", "Explicit confirm true required")
|
||||
return
|
||||
}
|
||||
tag, e := s.db.Exec(r.Context(), `UPDATE attempts SET release_requested=true WHERE id::text=$1 AND phase='finished' AND cleanup IN ('evidence_held','failed') AND EXISTS(SELECT 1 FROM allocations WHERE attempt_id=attempts.id)`, r.PathValue("id"))
|
||||
tag, e := s.db.Exec(r.Context(), `UPDATE attempts SET release_requested=true WHERE id=$1::uuid AND phase='finished' AND cleanup IN ('evidence_held','failed') AND EXISTS(SELECT 1 FROM allocations WHERE attempt_id=attempts.id)`, r.PathValue("id"))
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -335,3 +346,7 @@ func (s *Server) releaseEvidence(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
jsonResponse(w, 202, map[string]any{"release_requested": true})
|
||||
}
|
||||
|
||||
func (s *Server) heldAttempts(w http.ResponseWriter, r *http.Request) {
|
||||
s.list(w, r, `SELECT jsonb_build_object('attempt_id',a.id,'job_id',j.id,'run_id',r.id,'profile_name',r.profile_name,'filename',u.filename,'outcome',a.outcome,'cleanup',a.cleanup,'error',a.error,'finished_at',a.finished_at,'release_requested',a.release_requested,'allocation',CASE WHEN v.id IS NULL THEN NULL ELSE jsonb_build_object('id',v.id,'node',v.node,'vmid',v.vmid,'state',v.state) END) FROM attempts a JOIN runs r ON r.id=a.run_id JOIN jobs j ON j.id=r.job_id JOIN uploads u ON u.id=j.upload_id LEFT JOIN allocations v ON v.attempt_id=a.id WHERE a.phase='finished' AND a.cleanup IN ('evidence_held','failed') ORDER BY a.finished_at DESC NULLS LAST,a.created_at DESC,a.id LIMIT 200`)
|
||||
}
|
||||
|
||||
@@ -120,3 +120,40 @@ func TestAccountLoginLimitAndSessionRevocation(t *testing.T) {
|
||||
}
|
||||
fmt.Fprintln(os.Stdout, "17 valid proxied logins; account failure throttling; unrelated user unaffected; revoked session denied; DB failure remains 500")
|
||||
}
|
||||
|
||||
func TestInvalidUUIDPathsReturnNotFound(t *testing.T) {
|
||||
s := &Server{mux: http.NewServeMux()}
|
||||
s.routes()
|
||||
for _, route := range []struct {
|
||||
method string
|
||||
path string
|
||||
}{
|
||||
{"GET", "/api/v1/jobs/not-a-uuid"},
|
||||
{"POST", "/api/v1/jobs/not-a-uuid/retry"},
|
||||
{"GET", "/api/v1/uploads/not-a-uuid/content"},
|
||||
{"GET", "/api/v1/artifacts/not-a-uuid/content"},
|
||||
{"GET", "/api/v1/attempts/not-a-uuid/video"},
|
||||
{"PATCH", "/api/v1/admin/users/not-a-uuid"},
|
||||
{"POST", "/api/v1/admin/profiles/not-a-uuid/publish"},
|
||||
{"GET", "/api/v1/admin/profiles/not-a-uuid/revisions"},
|
||||
{"GET", "/api/v1/admin/qualifications/not-a-uuid"},
|
||||
{"POST", "/api/v1/admin/attempts/not-a-uuid/release-evidence"},
|
||||
} {
|
||||
t.Run(route.method+" "+route.path, func(t *testing.T) {
|
||||
r := httptest.NewRequest(route.method, route.path, strings.NewReader(`{}`))
|
||||
w := httptest.NewRecorder()
|
||||
s.ServeHTTP(w, r)
|
||||
var body struct {
|
||||
Error struct {
|
||||
Code string `json:"code"`
|
||||
} `json:"error"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &body); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if w.Code != 404 || body.Error.Code != "not_found" {
|
||||
t.Fatalf("invalid UUID: %d %s", w.Code, w.Body)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,8 +3,10 @@ package api
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"net/http/httptest"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
@@ -239,4 +241,171 @@ func TestJobSummariesIncludeAssignedProfile(t *testing.T) {
|
||||
if get("/api/v1/jobs/" + createdID)["runs"].([]any)[0].(map[string]any)["profile_name"] != manualName {
|
||||
t.Fatal("queued run name followed later profile edit")
|
||||
}
|
||||
|
||||
requestError := func(method, path, body string, status int, code string) {
|
||||
t.Helper()
|
||||
req := httptest.NewRequest(method, path, strings.NewReader(body))
|
||||
req.AddCookie(cookie)
|
||||
req.Header.Set("Origin", "http://localhost")
|
||||
req.Header.Set("X-CSRF-Token", session.CSRF)
|
||||
res := httptest.NewRecorder()
|
||||
s.ServeHTTP(res, req)
|
||||
var value struct {
|
||||
Error struct {
|
||||
Code string `json:"code"`
|
||||
} `json:"error"`
|
||||
}
|
||||
if err := json.Unmarshal(res.Body.Bytes(), &value); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if res.Code != status || value.Error.Code != code {
|
||||
t.Fatalf("%s %s: %d %s", method, path, res.Code, res.Body)
|
||||
}
|
||||
}
|
||||
// A stale terminal job status must not allow duplicate unfinished attempts.
|
||||
{
|
||||
if _, err := db.Exec(ctx, `UPDATE jobs SET status='failed' WHERE id=$1`, createdID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
requestError("POST", "/api/v1/jobs/"+createdID+"/retry", `{}`, 409, "job_active")
|
||||
body := get("/api/v1/jobs/" + createdID)
|
||||
attempts := body["runs"].([]any)[0].(map[string]any)["attempts"].([]any)
|
||||
if body["status"] != "failed" || len(attempts) != 1 || attempts[0].(map[string]any)["phase"] != "queued" {
|
||||
t.Fatalf("rejected retry changed the job or duplicated its unfinished attempt: %v", body)
|
||||
}
|
||||
if _, err := db.Exec(ctx, `UPDATE attempts SET phase='finished',outcome='error' WHERE run_id IN (SELECT id FROM runs WHERE job_id=$1)`, createdID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
body = writeRequest("POST", "/api/v1/jobs/"+createdID+"/retry", map[string]any{})
|
||||
attempts = body["runs"].([]any)[0].(map[string]any)["attempts"].([]any)
|
||||
if body["status"] != "queued" || len(attempts) != 2 || attempts[0].(map[string]any)["phase"] != "finished" || attempts[1].(map[string]any)["phase"] != "queued" {
|
||||
t.Fatalf("terminal retry did not preserve history and queue one new attempt: %v", body)
|
||||
}
|
||||
body = writeRequest("POST", "/api/v1/jobs/"+job+"/retry", map[string]any{})
|
||||
attempts = body["runs"].([]any)[0].(map[string]any)["attempts"].([]any)
|
||||
if len(attempts) != 2 || attempts[0].(map[string]any)["outcome"] != "executed" || attempts[1].(map[string]any)["phase"] != "queued" {
|
||||
t.Fatalf("whole-job retry omitted a previously successful run: %v", body)
|
||||
}
|
||||
}
|
||||
// Source seals are shared even when revisions belong to different profiles.
|
||||
{
|
||||
if _, err := db.Exec(ctx, `UPDATE revisions SET qualification='{"status":"passed"}' WHERE id=$1`, revision); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
otherProfile, newer := store.NewID(), store.NewID()
|
||||
if _, err := db.Exec(ctx, `INSERT INTO profiles(id,name,os,architecture) VALUES($1,'Other manual label','Windows','x64')`, otherProfile); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := db.Exec(ctx, `INSERT INTO revisions(id,profile_id,source_ref,fingerprint,config_digest,created_at) SELECT $1,$2,source_ref,'new-fingerprint','new-config',created_at+interval '1 second' FROM revisions WHERE id=$3`, newer, otherProfile, revision); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
writeRequest("POST", "/api/v1/admin/profiles/"+profile+"/maintenance", map[string]any{})
|
||||
requestError("POST", "/api/v1/admin/profiles/"+profile+"/publish", `{"revision_id":"`+revision+`"}`, 409, "stale_revision")
|
||||
for _, item := range get("/api/v1/admin/profiles")["items"].([]any) {
|
||||
p := item.(map[string]any)
|
||||
if p["id"] == profile && (p["state"] != "maintenance" || p["enabled"] != false) {
|
||||
t.Fatalf("rejected stale publish reopened admission: %v", p)
|
||||
}
|
||||
}
|
||||
}
|
||||
if _, err := db.Exec(ctx, `UPDATE attempts SET created_at='2026-01-01T00:00:00Z' WHERE run_id=$1 AND phase='queued'`, run); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := db.Exec(ctx, `UPDATE attempts SET created_at='2026-01-02T00:00:00Z' WHERE run_id IN (SELECT id FROM runs WHERE job_id=$1) AND phase='queued'`, createdID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if ahead := get("/api/v1/jobs/" + createdID)["queue_ahead"]; ahead != float64(1) {
|
||||
t.Fatalf("earlier queued attempt was not counted: %v", ahead)
|
||||
}
|
||||
if _, err := db.Exec(ctx, `UPDATE attempts SET claimed_at=now() WHERE run_id=$1 AND phase='queued'`, run); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if ahead := get("/api/v1/jobs/" + createdID)["queue_ahead"]; ahead != float64(0) {
|
||||
t.Fatalf("already claimed attempt was counted ahead: %v", ahead)
|
||||
}
|
||||
claimedJob := get("/api/v1/jobs/" + job)
|
||||
if claimedJob["queue_ahead"] != nil {
|
||||
t.Fatalf("job without unclaimed queued attempts has a queue position: %v", claimedJob["queue_ahead"])
|
||||
}
|
||||
for _, value := range claimedJob["runs"].([]any)[0].(map[string]any)["attempts"].([]any) {
|
||||
a := value.(map[string]any)
|
||||
if a["phase"] == "queued" && a["claimed_at"] == nil {
|
||||
t.Fatal("claimed_at missing from job detail")
|
||||
}
|
||||
}
|
||||
if _, err := db.Exec(ctx, `UPDATE attempts SET cleanup='failed',finished_at=now() WHERE id=$1`, attempt); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
held := get("/api/v1/admin/attempts/held")["items"].([]any)
|
||||
if len(held) != 1 || held[0].(map[string]any)["attempt_id"] != attempt || held[0].(map[string]any)["allocation"] != nil {
|
||||
t.Fatalf("failed cleanup without allocation missing from held evidence: %v", held)
|
||||
}
|
||||
allocation := store.NewID()
|
||||
if _, err := db.Exec(ctx, `INSERT INTO allocations(id,attempt_id,owner_id,node,vmid,pool,state) VALUES($1,$2,$3,'test-node',7100,'test-pool','stopped')`, allocation, attempt, owner); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
held = get("/api/v1/admin/attempts/held")["items"].([]any)
|
||||
if held[0].(map[string]any)["allocation"].(map[string]any)["id"] != allocation {
|
||||
t.Fatalf("held evidence lost its diagnostic allocation: %v", held)
|
||||
}
|
||||
// A competing commit while bytes arrive must not block on the owner or exceed quota.
|
||||
s.cfg.MaxOwnerBytes = 1024
|
||||
s.cfg.MaxUploadBytes = 512
|
||||
streamed := false
|
||||
payload := strings.NewReader("test")
|
||||
probe := uploadReaderFunc(func(p []byte) (int, error) {
|
||||
if !streamed {
|
||||
streamed = true
|
||||
tx, err := db.Begin(ctx)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
if _, err = tx.Exec(ctx, `SELECT id FROM users WHERE id=$1 FOR UPDATE NOWAIT`, owner); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
var used int64
|
||||
if err = tx.QueryRow(ctx, ownerStorageQuery, owner).Scan(&used); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
competitor := store.NewID()
|
||||
if _, err = tx.Exec(ctx, `INSERT INTO uploads(id,owner_id,filename,size,sha256,storage_key) VALUES($1,$2,'concurrent.exe',$3,$4,$5)`, competitor, owner, s.cfg.MaxOwnerBytes-used-1, strings.Repeat("c", 64), competitor); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if err = tx.Commit(ctx); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
return payload.Read(p)
|
||||
})
|
||||
uploadRequest := httptest.NewRequest("POST", "/api/v1/uploads", io.NopCloser(probe))
|
||||
uploadRequest.AddCookie(cookie)
|
||||
uploadRequest.Header.Set("Origin", "http://localhost")
|
||||
uploadRequest.Header.Set("X-CSRF-Token", session.CSRF)
|
||||
uploadRequest.Header.Set("X-Filename", "race.exe")
|
||||
uploadResponse := httptest.NewRecorder()
|
||||
s.ServeHTTP(uploadResponse, uploadRequest)
|
||||
var uploadError struct {
|
||||
Error struct {
|
||||
Code string `json:"code"`
|
||||
} `json:"error"`
|
||||
}
|
||||
if err := json.Unmarshal(uploadResponse.Body.Bytes(), &uploadError); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if uploadResponse.Code != 429 || uploadError.Error.Code != "quota_exceeded" {
|
||||
t.Fatalf("upload held a streaming lock or skipped final quota check: %d %s", uploadResponse.Code, uploadResponse.Body)
|
||||
}
|
||||
files, err := os.ReadDir(filepath.Join(s.cfg.ArtifactRoot, "uploads", owner))
|
||||
if err != nil || len(files) != 0 {
|
||||
t.Fatalf("rejected upload retained private file: %v %v", files, err)
|
||||
}
|
||||
var uploaded int
|
||||
if err := db.QueryRow(ctx, `SELECT count(*) FROM uploads WHERE filename='race.exe'`).Scan(&uploaded); err != nil || uploaded != 0 {
|
||||
t.Fatalf("rejected upload was committed: %d %v", uploaded, err)
|
||||
}
|
||||
}
|
||||
|
||||
type uploadReaderFunc func([]byte) (int, error)
|
||||
|
||||
func (f uploadReaderFunc) Read(p []byte) (int, error) { return f(p) }
|
||||
|
||||
+40
-19
@@ -144,6 +144,8 @@ func (s *Server) jobRoutes() {
|
||||
}))
|
||||
m.HandleFunc("GET /api/v1/dashboard", s.protected(false, s.dashboard))
|
||||
}
|
||||
const ownerStorageQuery = `SELECT COALESCE((SELECT sum(size) FROM uploads WHERE owner_id=$1),0)+COALESCE((SELECT sum(a.size) FROM artifacts a JOIN jobs j ON j.id=a.job_id WHERE j.owner_id=$1),0)`
|
||||
|
||||
func (s *Server) upload(w http.ResponseWriter, r *http.Request) {
|
||||
a := user(r)
|
||||
name, e := url.PathUnescape(r.Header.Get("X-Filename"))
|
||||
@@ -151,19 +153,8 @@ func (s *Server) upload(w http.ResponseWriter, r *http.Request) {
|
||||
fail(w, 400, "invalid_filename", "A safe filename with one of the ten supported extensions is required")
|
||||
return
|
||||
}
|
||||
tx, e := s.db.Begin(r.Context())
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
defer tx.Rollback(r.Context())
|
||||
var id string
|
||||
if e = tx.QueryRow(r.Context(), `SELECT id::text FROM users WHERE id=$1 FOR UPDATE`, a.ID).Scan(&id); e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
var used int64
|
||||
e = tx.QueryRow(r.Context(), `SELECT COALESCE((SELECT sum(size) FROM uploads WHERE owner_id=$1),0)+COALESCE((SELECT sum(a.size) FROM artifacts a JOIN jobs j ON j.id=a.job_id WHERE j.owner_id=$1),0)`, a.ID).Scan(&used)
|
||||
e = s.db.QueryRow(r.Context(), ownerStorageQuery, a.ID).Scan(&used)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -218,6 +209,25 @@ func (s *Server) upload(w http.ResponseWriter, r *http.Request) {
|
||||
fail(w, 500, "storage_error", "Could not finalize upload")
|
||||
return
|
||||
}
|
||||
tx, e := s.db.Begin(r.Context())
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
defer tx.Rollback(r.Context())
|
||||
var id string
|
||||
if e = tx.QueryRow(r.Context(), `SELECT id::text FROM users WHERE id=$1 FOR UPDATE`, a.ID).Scan(&id); e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
if e = tx.QueryRow(r.Context(), ownerStorageQuery, a.ID).Scan(&used); e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
if n > s.cfg.MaxOwnerBytes-used {
|
||||
fail(w, 429, "quota_exceeded", "Owner storage quota reached")
|
||||
return
|
||||
}
|
||||
sum := hex.EncodeToString(h.Sum(nil))
|
||||
var created any
|
||||
e = tx.QueryRow(r.Context(), `INSERT INTO uploads(id,owner_id,filename,size,sha256,storage_key) VALUES($1,$2,$3,$4,$5,$6) RETURNING created_at`, uploadID, a.ID, name, n, sum, key).Scan(&created)
|
||||
@@ -264,7 +274,7 @@ func (s *Server) servePrivate(w http.ResponseWriter, r *http.Request, key, name,
|
||||
func (s *Server) uploadContent(w http.ResponseWriter, r *http.Request) {
|
||||
a := user(r)
|
||||
var key, name string
|
||||
e := s.db.QueryRow(r.Context(), `SELECT storage_key,filename FROM uploads WHERE id::text=$1 AND(owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role).Scan(&key, &name)
|
||||
e := s.db.QueryRow(r.Context(), `SELECT storage_key,filename FROM uploads WHERE id=$1::uuid AND(owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role).Scan(&key, &name)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -274,7 +284,7 @@ func (s *Server) uploadContent(w http.ResponseWriter, r *http.Request) {
|
||||
func (s *Server) artifactContent(w http.ResponseWriter, r *http.Request) {
|
||||
a := user(r)
|
||||
var key, name, ct string
|
||||
e := s.db.QueryRow(r.Context(), `SELECT a.storage_key,a.filename,a.content_type FROM artifacts a JOIN jobs j ON j.id=a.job_id WHERE a.id::text=$1 AND(j.owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role).Scan(&key, &name, &ct)
|
||||
e := s.db.QueryRow(r.Context(), `SELECT a.storage_key,a.filename,a.content_type FROM artifacts a JOIN jobs j ON j.id=a.job_id WHERE a.id=$1::uuid AND(j.owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role).Scan(&key, &name, &ct)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -392,9 +402,11 @@ func (s *Server) admit(r *http.Request, tx pgx.Tx, owner string) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
const jobRunsSummary = `COALESCE((SELECT jsonb_agg(to_jsonb(r) ORDER BY r.id) FROM runs r WHERE r.job_id=j.id),'[]'::jsonb)`
|
||||
const jobRunsSummary = `COALESCE((SELECT jsonb_agg(to_jsonb(r) ORDER BY r.profile_name,r.id) FROM runs r WHERE r.job_id=j.id),'[]'::jsonb)`
|
||||
|
||||
const jobQuery = `SELECT to_jsonb(j)||jsonb_build_object('filename',u.filename,'sha256',u.sha256,'size',u.size,'runs',COALESCE((SELECT jsonb_agg(to_jsonb(r)||jsonb_build_object('attempts',COALESCE((SELECT jsonb_agg((to_jsonb(a)-'lease_owner'-'lease_until'-'release_requested')||jsonb_build_object('allocation',CASE WHEN $3='admin' THEN(SELECT jsonb_build_object('id',v.id,'node',v.node,'vmid',v.vmid,'state',v.state) FROM allocations v WHERE v.attempt_id=a.id) ELSE NULL END,'artifacts',COALESCE((SELECT jsonb_agg(to_jsonb(ar)-'storage_key'||jsonb_build_object('url','/api/v1/artifacts/'||ar.id||'/content')) FROM artifacts ar WHERE ar.attempt_id=a.id),'[]'::jsonb)) ORDER BY a.created_at) FROM attempts a WHERE a.run_id=r.id),'[]'::jsonb))) FROM runs r WHERE r.job_id=j.id),'[]'::jsonb)) FROM jobs j JOIN uploads u ON u.id=j.upload_id WHERE j.id::text=$1 AND(j.owner_id=$2 OR $3='admin')`
|
||||
const jobQuery = `SELECT to_jsonb(j)||jsonb_build_object('filename',u.filename,'sha256',u.sha256,'size',u.size,
|
||||
'queue_ahead',(SELECT (SELECT count(*) FROM attempts ahead JOIN runs ar ON ar.id=ahead.run_id WHERE ar.job_id<>j.id AND ahead.phase='queued' AND ahead.claimed_at IS NULL AND ahead.created_at<first_queued.created_at) FROM (SELECT min(qa.created_at) AS created_at FROM attempts qa JOIN runs qr ON qr.id=qa.run_id WHERE qr.job_id=j.id AND qa.phase='queued' AND qa.claimed_at IS NULL) first_queued WHERE first_queued.created_at IS NOT NULL),
|
||||
'runs',COALESCE((SELECT jsonb_agg(to_jsonb(r)||jsonb_build_object('attempts',COALESCE((SELECT jsonb_agg((to_jsonb(a)-'lease_owner'-'lease_until'-'release_requested')||jsonb_build_object('allocation',CASE WHEN $3='admin' THEN(SELECT jsonb_build_object('id',v.id,'node',v.node,'vmid',v.vmid,'state',v.state) FROM allocations v WHERE v.attempt_id=a.id) ELSE NULL END,'artifacts',COALESCE((SELECT jsonb_agg(to_jsonb(ar)-'storage_key'||jsonb_build_object('url','/api/v1/artifacts/'||ar.id||'/content')) FROM artifacts ar WHERE ar.attempt_id=a.id),'[]'::jsonb)) ORDER BY a.created_at) FROM attempts a WHERE a.run_id=r.id),'[]'::jsonb)) ORDER BY r.profile_name,r.id) FROM runs r WHERE r.job_id=j.id),'[]'::jsonb)) FROM jobs j JOIN uploads u ON u.id=j.upload_id WHERE j.id=$1::uuid AND(j.owner_id=$2 OR $3='admin')`
|
||||
|
||||
func (s *Server) respondJob(w http.ResponseWriter, r *http.Request, id string, status int) {
|
||||
a := user(r)
|
||||
@@ -411,7 +423,7 @@ func (s *Server) getJob(w http.ResponseWriter, r *http.Request) {
|
||||
func (s *Server) ownJob(r *http.Request) (string, error) {
|
||||
a := user(r)
|
||||
var id string
|
||||
e := s.db.QueryRow(r.Context(), `SELECT id::text FROM jobs WHERE id::text=$1 AND(owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role).Scan(&id)
|
||||
e := s.db.QueryRow(r.Context(), `SELECT id::text FROM jobs WHERE id=$1::uuid AND(owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role).Scan(&id)
|
||||
return id, e
|
||||
}
|
||||
func (s *Server) jobs(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -486,7 +498,13 @@ func (s *Server) retryJob(w http.ResponseWriter, r *http.Request) {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
if status == "queued" || status == "running" {
|
||||
var unfinished bool
|
||||
e = tx.QueryRow(r.Context(), `SELECT EXISTS(SELECT 1 FROM attempts a JOIN runs r ON r.id=a.run_id WHERE r.job_id=$1 AND a.phase<>'finished')`, id).Scan(&unfinished)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
}
|
||||
if unfinished || status == "queued" || status == "running" {
|
||||
fail(w, 409, "job_active", "An active job cannot be retried")
|
||||
return
|
||||
}
|
||||
@@ -511,6 +529,9 @@ func (s *Server) retryJob(w http.ResponseWriter, r *http.Request) {
|
||||
ids = append(ids, run)
|
||||
}
|
||||
rows.Close()
|
||||
if e == nil {
|
||||
e = rows.Err()
|
||||
}
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
@@ -561,7 +582,7 @@ func (s *Server) artifacts(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
func (s *Server) video(w http.ResponseWriter, r *http.Request) {
|
||||
a := user(r)
|
||||
v, e := s.queryObject(r.Context(), `SELECT a.video FROM attempts a JOIN runs r ON r.id=a.run_id JOIN jobs j ON j.id=r.job_id WHERE a.id::text=$1 AND(j.owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role)
|
||||
v, e := s.queryObject(r.Context(), `SELECT a.video FROM attempts a JOIN runs r ON r.id=a.run_id JOIN jobs j ON j.id=r.job_id WHERE a.id=$1::uuid AND(j.owner_id=$2 OR $3='admin')`, r.PathValue("id"), a.ID, a.Role)
|
||||
if e != nil {
|
||||
s.dbError(w, e)
|
||||
return
|
||||
|
||||
@@ -10,6 +10,7 @@ import (
|
||||
"errors"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net"
|
||||
"net/http"
|
||||
"os"
|
||||
"regexp"
|
||||
@@ -41,6 +42,19 @@ var uuidRE = regexp.MustCompile(`^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4
|
||||
var usernameRE = regexp.MustCompile(`^[A-Za-z0-9_.-]{3,64}$`)
|
||||
|
||||
func New(db *pgxpool.Pool, c Config) (*Server, error) {
|
||||
if os.Getenv("PUBLIC_ORIGIN") == "" {
|
||||
slog.Warn("PUBLIC_ORIGIN is unset; configure the external origin before exposing the API")
|
||||
}
|
||||
if os.Getenv("ALLOW_INSECURE_HTTP") == "true" {
|
||||
addr := os.Getenv("LISTEN_ADDR")
|
||||
if addr == "" {
|
||||
addr = ":8080"
|
||||
}
|
||||
host, _, err := net.SplitHostPort(addr)
|
||||
if err != nil || !strings.EqualFold(host, "localhost") && !net.ParseIP(host).IsLoopback() {
|
||||
slog.Warn("ALLOW_INSECURE_HTTP is enabled on a non-loopback listener; session cookies are not secure and HTTP must not be exposed to untrusted networks", "listen_addr", addr)
|
||||
}
|
||||
}
|
||||
if c.Origin == "" || c.ArtifactRoot == "" {
|
||||
return nil, errors.New("origin and artifact root required")
|
||||
}
|
||||
@@ -117,6 +131,10 @@ func (s *Server) origin(w http.ResponseWriter, r *http.Request) bool {
|
||||
}
|
||||
func (s *Server) protected(admin bool, h http.HandlerFunc) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
if id := r.PathValue("id"); id != "" && !uuidRE.MatchString(id) {
|
||||
fail(w, 404, "not_found", "Resource not found")
|
||||
return
|
||||
}
|
||||
c, err := r.Cookie("otche_session")
|
||||
if err != nil || len(c.Value) != 64 {
|
||||
fail(w, 401, "unauthenticated", "Sign in required")
|
||||
|
||||
+52
-7
@@ -9,6 +9,7 @@ import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"mime/multipart"
|
||||
"net/http"
|
||||
"net/url"
|
||||
@@ -78,24 +79,49 @@ func (c *Client) endpoint(path string, q url.Values) string {
|
||||
u.RawQuery = q.Encode()
|
||||
return u.String()
|
||||
}
|
||||
const responseLimit = 24 << 20
|
||||
|
||||
var errTransport = errors.New("PVE transport failed")
|
||||
var errResponseRead = errors.New("PVE response read failed")
|
||||
|
||||
func (c *Client) decode(req *http.Request, out any) error {
|
||||
for attempt := 0; ; attempt++ {
|
||||
err := c.decodeOnce(req, out)
|
||||
var status *HTTPError
|
||||
retry := errors.Is(err, errTransport) || errors.Is(err, errResponseRead) || (errors.As(err, &status) && status.Status >= 500 && status.Status < 600)
|
||||
if req.Method != http.MethodGet || !retry || attempt == 2 {
|
||||
return err
|
||||
}
|
||||
timer := time.NewTimer((500 * time.Millisecond) << attempt)
|
||||
select {
|
||||
case <-req.Context().Done():
|
||||
timer.Stop()
|
||||
return req.Context().Err()
|
||||
case <-timer.C:
|
||||
}
|
||||
}
|
||||
}
|
||||
func (c *Client) decodeOnce(req *http.Request, out any) error {
|
||||
req.Header.Set("Authorization", c.auth)
|
||||
resp, err := c.http.Do(req)
|
||||
if err != nil {
|
||||
if req.Context().Err() != nil {
|
||||
return req.Context().Err()
|
||||
}
|
||||
return errors.New("PVE transport failed")
|
||||
return errTransport
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
|
||||
return &HTTPError{resp.StatusCode}
|
||||
}
|
||||
b, err := io.ReadAll(io.LimitReader(resp.Body, 24<<20+1))
|
||||
b, err := io.ReadAll(io.LimitReader(resp.Body, responseLimit+1))
|
||||
if err != nil {
|
||||
return errors.New("PVE response read failed")
|
||||
if req.Context().Err() != nil {
|
||||
return req.Context().Err()
|
||||
}
|
||||
return errResponseRead
|
||||
}
|
||||
if len(b) > 24<<20 {
|
||||
if len(b) > responseLimit {
|
||||
return errors.New("PVE response exceeded bound")
|
||||
}
|
||||
var envelope struct {
|
||||
@@ -143,8 +169,7 @@ func (c *Client) Task(ctx context.Context, node, upid string) error {
|
||||
if err := ValidateUPID(node, upid); err != nil {
|
||||
return err
|
||||
}
|
||||
ticker := time.NewTicker(time.Second)
|
||||
defer ticker.Stop()
|
||||
interval := 250 * time.Millisecond
|
||||
for {
|
||||
var s struct {
|
||||
Status string `json:"status"`
|
||||
@@ -156,6 +181,23 @@ func (c *Client) Task(ctx context.Context, node, upid string) error {
|
||||
switch s.Status {
|
||||
case "stopped":
|
||||
if s.ExitStatus != "OK" {
|
||||
// Detailed task failures belong only in operator logs, never owner errors.
|
||||
detail := s.ExitStatus
|
||||
for _, credential := range strings.Split(strings.TrimPrefix(c.auth, "PVEAPIToken="), "=") {
|
||||
if credential != "" {
|
||||
detail = strings.ReplaceAll(detail, credential, "[redacted]")
|
||||
}
|
||||
}
|
||||
detail = strings.Map(func(r rune) rune {
|
||||
if r < 32 || r > 126 {
|
||||
return ' '
|
||||
}
|
||||
return r
|
||||
}, detail)
|
||||
if len(detail) > 120 {
|
||||
detail = detail[:120]
|
||||
}
|
||||
slog.Error("PVE task failed", "exitstatus", detail)
|
||||
return errors.New("PVE task failed")
|
||||
}
|
||||
return nil
|
||||
@@ -163,11 +205,14 @@ func (c *Client) Task(ctx context.Context, node, upid string) error {
|
||||
default:
|
||||
return errors.New("invalid PVE task status")
|
||||
}
|
||||
timer := time.NewTimer(interval)
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
timer.Stop()
|
||||
return ctx.Err()
|
||||
case <-ticker.C:
|
||||
case <-timer.C:
|
||||
}
|
||||
interval = time.Second
|
||||
}
|
||||
}
|
||||
func VMPath(node string, vmid int) string { return "/nodes/" + node + "/qemu/" + strconv.Itoa(vmid) }
|
||||
|
||||
+106
-3
@@ -6,6 +6,7 @@ import (
|
||||
"encoding/base64"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"net/url"
|
||||
@@ -93,11 +94,30 @@ func TestMalformedUPIDNeverPolled(t *testing.T) {
|
||||
}
|
||||
}
|
||||
func TestTaskRequiresTerminalOK(t *testing.T) {
|
||||
var logs bytes.Buffer
|
||||
previous := slog.Default()
|
||||
slog.SetDefault(slog.New(slog.NewJSONHandler(&logs, nil)))
|
||||
t.Cleanup(func() { slog.SetDefault(previous) })
|
||||
c := fixture(t, func(w http.ResponseWriter, r *http.Request) {
|
||||
data(w, map[string]any{"status": "stopped", "exitstatus": "ERROR: cloning failed"})
|
||||
data(w, map[string]any{"status": "stopped", "exitstatus": "ERROR: private\n\x1b cloning failed " + strings.Repeat("x", 140)})
|
||||
})
|
||||
if c.Task(context.Background(), "node", "UPID:node:00000001:00000002:00000003:qmclone:9001:user@pve:") == nil {
|
||||
t.Fatal("task failure treated as success")
|
||||
err := c.Task(context.Background(), "node", "UPID:node:00000001:00000002:00000003:qmclone:9001:user@pve:")
|
||||
if err == nil || err.Error() != "PVE task failed" {
|
||||
t.Fatalf("task failure lost or exposed internal status: %v", err)
|
||||
}
|
||||
var entry struct {
|
||||
ExitStatus string `json:"exitstatus"`
|
||||
}
|
||||
if err := json.Unmarshal(logs.Bytes(), &entry); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(entry.ExitStatus, "cloning failed") || strings.Contains(entry.ExitStatus, "private") || len(entry.ExitStatus) > 120 {
|
||||
t.Fatalf("operator log lost detail, leaked credentials or exceeded bound: %q", entry.ExitStatus)
|
||||
}
|
||||
for _, r := range entry.ExitStatus {
|
||||
if r < 32 || r > 126 {
|
||||
t.Fatalf("unprintable task detail: %q", entry.ExitStatus)
|
||||
}
|
||||
}
|
||||
}
|
||||
func TestQGAReadOffsetsUseDecodedBytesAndEnforceLimit(t *testing.T) {
|
||||
@@ -271,3 +291,86 @@ func TestUploadStreamsKnownLengthMultipartFromCurrentOffset(t *testing.T) {
|
||||
t.Fatalf("upload result=%q error=%v", got, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOnlyGETRetriesTransientFailures(t *testing.T) {
|
||||
for _, method := range []string{http.MethodGet, http.MethodPost, http.MethodPut, http.MethodDelete} {
|
||||
t.Run(method, func(t *testing.T) {
|
||||
calls := 0
|
||||
c := fixture(t, func(w http.ResponseWriter, r *http.Request) {
|
||||
calls++
|
||||
if calls == 1 {
|
||||
w.WriteHeader(http.StatusBadGateway)
|
||||
return
|
||||
}
|
||||
data(w, "recovered")
|
||||
})
|
||||
var result string
|
||||
err := c.Do(context.Background(), method, "/fixture", nil, &result)
|
||||
if method == http.MethodGet {
|
||||
if err != nil || result != "recovered" || calls != 2 {
|
||||
t.Fatalf("GET failed to recover: result=%q calls=%d error=%v", result, calls, err)
|
||||
}
|
||||
} else if err == nil || calls != 1 {
|
||||
t.Fatalf("mutation retried or accepted: calls=%d error=%v", calls, err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestGETRetryBudgetAndCancellation(t *testing.T) {
|
||||
t.Run("exhausted", func(t *testing.T) {
|
||||
calls := 0
|
||||
c := fixture(t, func(w http.ResponseWriter, r *http.Request) {
|
||||
calls++
|
||||
w.WriteHeader(http.StatusServiceUnavailable)
|
||||
})
|
||||
if err := c.Do(context.Background(), http.MethodGet, "/fixture", nil, nil); err == nil || calls != 3 {
|
||||
t.Fatalf("GET retry budget changed: calls=%d error=%v", calls, err)
|
||||
}
|
||||
})
|
||||
t.Run("cancelled", func(t *testing.T) {
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
defer cancel()
|
||||
calls := 0
|
||||
c := fixture(t, func(w http.ResponseWriter, r *http.Request) {
|
||||
calls++
|
||||
w.WriteHeader(http.StatusBadGateway)
|
||||
cancel()
|
||||
})
|
||||
if err := c.Do(ctx, http.MethodGet, "/fixture", nil, nil); err != context.Canceled || calls != 1 {
|
||||
t.Fatalf("GET ignored cancellation: calls=%d error=%v", calls, err)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
func TestQGAOneMiBChunkFitsResponseBound(t *testing.T) {
|
||||
payload := bytes.Repeat([]byte{0, 255, 17, 32}, (1<<20)/4)
|
||||
c := fixture(t, func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Query().Get("count") != "1048576" {
|
||||
t.Errorf("unexpected chunk count: %s", r.URL.Query().Get("count"))
|
||||
}
|
||||
data(w, map[string]any{"content": base64.StdEncoding.EncodeToString(payload), "bytes-read": len(payload), "truncated": false})
|
||||
})
|
||||
var out bytes.Buffer
|
||||
n, err := c.ReadFile(context.Background(), "node", 9001, "file", &out, 1<<20)
|
||||
if err != nil || n != int64(len(payload)) || !bytes.Equal(out.Bytes(), payload) {
|
||||
t.Fatalf("one MiB binary chunk failed: n=%d error=%v", n, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGETRetriesInterruptedResponse(t *testing.T) {
|
||||
calls := 0
|
||||
c := fixture(t, func(w http.ResponseWriter, r *http.Request) {
|
||||
calls++
|
||||
if calls == 1 {
|
||||
w.Header().Set("Content-Length", "1000")
|
||||
_, _ = io.WriteString(w, `{"data":`)
|
||||
return
|
||||
}
|
||||
data(w, "recovered")
|
||||
})
|
||||
var result string
|
||||
if err := c.Do(context.Background(), http.MethodGet, "/fixture", nil, &result); err != nil || result != "recovered" || calls != 2 {
|
||||
t.Fatalf("interrupted response not retried: result=%q calls=%d error=%v", result, calls, err)
|
||||
}
|
||||
}
|
||||
|
||||
+2
-1
@@ -11,7 +11,8 @@ import (
|
||||
"time"
|
||||
)
|
||||
|
||||
const FileChunk = 256 << 10
|
||||
// A base64-encoded chunk plus its JSON envelope fits well within responseLimit.
|
||||
const FileChunk = 1 << 20
|
||||
const ControlLimit = 40 << 10
|
||||
|
||||
type Bool bool
|
||||
|
||||
@@ -28,6 +28,7 @@ type fragmentSink struct {
|
||||
kind string
|
||||
sawType, sawMovie, sawFragment bool
|
||||
fragments int
|
||||
lastSync time.Time
|
||||
ready chan struct{}
|
||||
err error
|
||||
}
|
||||
@@ -113,8 +114,11 @@ func (s *fragmentSink) endBox() error {
|
||||
if !s.sawType || !s.sawMovie || !s.sawFragment {
|
||||
return errors.New("encoded MP4 fragment has no initialization")
|
||||
}
|
||||
if err := s.file.Sync(); err != nil {
|
||||
return errors.New("recording sink sync failed")
|
||||
if s.fragments == 0 || time.Since(s.lastSync) >= time.Second {
|
||||
if err := s.file.Sync(); err != nil {
|
||||
return errors.New("recording sink sync failed")
|
||||
}
|
||||
s.lastSync = time.Now()
|
||||
}
|
||||
s.fragments++
|
||||
s.sawFragment = false
|
||||
|
||||
@@ -288,6 +288,14 @@ func syncDirectory(root *os.Root) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// framesDue includes the frame at elapsed zero. Dropped frames advance the
|
||||
// accounting clock only after the caller records their absence in the manifest.
|
||||
func framesDue(elapsed time.Duration, fps int, accounted int64) (frames, dropped int64) {
|
||||
due := max(int64(0), int64(elapsed)*int64(fps)/int64(time.Second)+1-accounted)
|
||||
frames = min(due, int64(2*fps+1))
|
||||
return frames, due - frames
|
||||
}
|
||||
|
||||
func (s *Session) run() {
|
||||
readerDone := make(chan struct{})
|
||||
go func() {
|
||||
@@ -328,14 +336,44 @@ func (s *Session) run() {
|
||||
}()
|
||||
var current *encoder
|
||||
var generation uint64
|
||||
var pixels []byte
|
||||
var segmentStart time.Time
|
||||
var droppedFrames int64
|
||||
catchUp := func(at time.Time, reserve int64) (int64, error) {
|
||||
frames, dropped := framesDue(at.Sub(segmentStart), s.fps, current.frames+droppedFrames)
|
||||
if dropped > 0 {
|
||||
gapAt := segmentStart.Add(time.Duration(current.frames+droppedFrames) * time.Second / time.Duration(s.fps))
|
||||
s.manifest.Gaps = append(s.manifest.Gaps, Gap{At: gapAt.UTC(), Reason: fmt.Sprintf("encoder fell behind wall clock; skipped %d frames at %d FPS", dropped, s.fps)})
|
||||
droppedFrames += dropped
|
||||
if err := s.saveManifest(); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
reserved := min(frames, reserve)
|
||||
for range frames - reserved {
|
||||
if err := current.writeFrame(pixels); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
return reserved, nil
|
||||
}
|
||||
announced := false
|
||||
finishSegment := func() {
|
||||
if current == nil {
|
||||
return
|
||||
}
|
||||
err := current.finish()
|
||||
// End capture at this instant, before draining the encoder. Include any
|
||||
// delayed final tick, but do not count FFmpeg finalization as capture time.
|
||||
end := time.Now()
|
||||
var err error
|
||||
if s.Err() == nil {
|
||||
_, err = catchUp(end, 0)
|
||||
}
|
||||
if finishErr := current.finish(); err == nil {
|
||||
err = finishErr
|
||||
}
|
||||
segment := &s.manifest.Segments[len(s.manifest.Segments)-1]
|
||||
now := time.Now().UTC()
|
||||
now := end.UTC()
|
||||
segment.FinishedAt = &now
|
||||
segment.Frames, segment.Size = current.frames, current.sink.bytes
|
||||
segment.State = "complete"
|
||||
@@ -391,7 +429,8 @@ func (s *Session) run() {
|
||||
select {
|
||||
case <-s.ctx.Done():
|
||||
return
|
||||
case tick := <-ticker.C:
|
||||
case <-ticker.C:
|
||||
tick := time.Now()
|
||||
if current != nil {
|
||||
select {
|
||||
case <-current.done:
|
||||
@@ -415,10 +454,11 @@ func (s *Session) run() {
|
||||
}
|
||||
}
|
||||
s.frame.mu.Lock()
|
||||
resized := current != nil && generation != s.frame.generation
|
||||
rotate := current != nil && tick.Sub(s.manifest.Segments[len(s.manifest.Segments)-1].StartedAt) >= segmentDuration
|
||||
width, height, frameGeneration, complete := s.frame.width, s.frame.height, s.frame.generation, s.frame.complete
|
||||
s.frame.mu.Unlock()
|
||||
resized := current != nil && generation != frameGeneration
|
||||
rotate := current != nil && tick.Sub(segmentStart) >= segmentDuration
|
||||
if resized || rotate {
|
||||
s.frame.mu.Unlock()
|
||||
finishSegment()
|
||||
if resized {
|
||||
s.manifest.Gaps = append(s.manifest.Gaps, Gap{At: tick.UTC(), Reason: "framebuffer resized; awaiting complete new frame"})
|
||||
@@ -430,33 +470,53 @@ func (s *Session) run() {
|
||||
if s.Err() != nil {
|
||||
return
|
||||
}
|
||||
s.frame.mu.Lock()
|
||||
}
|
||||
if !s.frame.complete {
|
||||
if !complete {
|
||||
continue
|
||||
}
|
||||
if current != nil {
|
||||
frames, err := catchUp(time.Now(), 1)
|
||||
if err != nil {
|
||||
s.fail(err)
|
||||
return
|
||||
}
|
||||
if frames == 0 {
|
||||
continue
|
||||
}
|
||||
}
|
||||
// Reuse one private snapshot. Network reads and encoder writes never
|
||||
// own the framebuffer lock; only the bounded pixel copy does.
|
||||
size := width * height * 3
|
||||
if cap(pixels) < size {
|
||||
pixels = make([]byte, size)
|
||||
}
|
||||
pixels = pixels[:size]
|
||||
s.frame.mu.Lock()
|
||||
if s.frame.generation != frameGeneration || !s.frame.complete {
|
||||
s.frame.mu.Unlock()
|
||||
continue
|
||||
}
|
||||
newSegment := current == nil
|
||||
if newSegment {
|
||||
copy(pixels, s.frame.pixels)
|
||||
s.frame.mu.Unlock()
|
||||
if current == nil {
|
||||
if len(s.manifest.Segments) >= maxSegments {
|
||||
s.frame.mu.Unlock()
|
||||
s.fail(errors.New("recording segment limit exceeded"))
|
||||
return
|
||||
}
|
||||
name := fmt.Sprintf("segment-%06d.mp4", len(s.manifest.Segments)+1)
|
||||
segmentStart = time.Now()
|
||||
var err error
|
||||
current, err = startEncoder(s.ffmpeg, s.root, name, s.frame.width, s.frame.height, s.fps, &s.remaining)
|
||||
current, err = startEncoder(s.ffmpeg, s.root, name, width, height, s.fps, &s.remaining)
|
||||
if err != nil {
|
||||
s.frame.mu.Unlock()
|
||||
s.fail(err)
|
||||
return
|
||||
}
|
||||
generation = s.frame.generation
|
||||
generation = frameGeneration
|
||||
droppedFrames = 0
|
||||
announced = false
|
||||
s.manifest.Segments = append(s.manifest.Segments, Segment{Filename: name, StartedAt: tick.UTC(), Width: s.frame.width, Height: s.frame.height, FPS: s.fps, State: "recording"})
|
||||
s.manifest.Segments = append(s.manifest.Segments, Segment{Filename: name, StartedAt: segmentStart.UTC(), Width: width, Height: height, FPS: s.fps, State: "recording"})
|
||||
}
|
||||
err := current.writeFrame(s.frame.pixels)
|
||||
s.frame.mu.Unlock()
|
||||
err := current.writeFrame(pixels)
|
||||
if err != nil {
|
||||
s.fail(err)
|
||||
return
|
||||
|
||||
@@ -563,3 +563,66 @@ func TestFFmpegCancellationCannotBecomeComplete(t *testing.T) {
|
||||
t.Fatalf("cancelled recording marked %q", manifest.State)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFFmpegPacingDelayedTicks(t *testing.T) {
|
||||
requireFFmpeg(t)
|
||||
directory := t.TempDir()
|
||||
root, err := os.OpenRoot(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer root.Close()
|
||||
remaining := int64(4 << 20)
|
||||
const fps = 5
|
||||
encoder, err := startEncoder("ffmpeg", root, "pacing.mp4", 2, 2, fps, &remaining)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
finished := false
|
||||
defer func() {
|
||||
if !finished {
|
||||
encoder.finish()
|
||||
}
|
||||
}()
|
||||
var dropped int64
|
||||
pixels := bytes.Repeat([]byte{255, 0, 0}, 4)
|
||||
// Virtual elapsed time makes missed ticks and a delay beyond the catch-up
|
||||
// limit reproducible, independent of the host's encoder/scheduler speed.
|
||||
for _, elapsed := range []time.Duration{0, 200 * time.Millisecond, 1100 * time.Millisecond, 1300 * time.Millisecond, 6 * time.Second, 6200 * time.Millisecond} {
|
||||
frames, skipped := framesDue(elapsed, fps, encoder.frames+dropped)
|
||||
if frames > 2*fps+1 {
|
||||
t.Fatalf("unbounded catch-up: %d frames", frames)
|
||||
}
|
||||
if elapsed < 6*time.Second && skipped != 0 {
|
||||
t.Fatalf("recoverable delay lost %d frames", skipped)
|
||||
}
|
||||
dropped += skipped
|
||||
for range frames {
|
||||
if err := encoder.writeFrame(pixels); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
accountedDuration := time.Duration(encoder.frames+dropped) * time.Second / fps
|
||||
if accountedDuration <= elapsed || accountedDuration > elapsed+time.Second/fps {
|
||||
t.Fatalf("pacing at %s: encoded=%d dropped=%d duration=%s", elapsed, encoder.frames, dropped, accountedDuration)
|
||||
}
|
||||
if pending, skipped := framesDue(elapsed, fps, encoder.frames+dropped); pending != 0 || skipped != 0 {
|
||||
t.Fatal("same elapsed time encoded twice")
|
||||
}
|
||||
}
|
||||
if dropped != 13 {
|
||||
t.Fatalf("long delay must report precisely 13 missing frames, got %d", dropped)
|
||||
}
|
||||
err = encoder.finish()
|
||||
finished = true
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
decoded, err := exec.Command("ffmpeg", "-v", "error", "-i", filepath.Join(directory, "pacing.mp4"), "-f", "rawvideo", "-pix_fmt", "rgb24", "pipe:1").Output()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if int64(len(decoded)) != encoder.frames*int64(len(pixels)) {
|
||||
t.Fatalf("encoded duration differs from scheduled frames: %d decoded bytes for %d frames", len(decoded), encoder.frames)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -71,6 +71,7 @@ type framebuffer struct {
|
||||
mu sync.Mutex
|
||||
width, height int
|
||||
pixels []byte
|
||||
row [maxWidth * 4]byte
|
||||
coverage []uint64
|
||||
covered int
|
||||
complete bool
|
||||
@@ -275,9 +276,6 @@ func (f *framebuffer) readUpdate() error {
|
||||
if count > 4096 {
|
||||
return errors.New("RFB rectangle count exceeds limit")
|
||||
}
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
var row [maxWidth * 4]byte
|
||||
var rectangle [12]byte
|
||||
var total int64
|
||||
for i := range count {
|
||||
@@ -291,7 +289,10 @@ func (f *framebuffer) readUpdate() error {
|
||||
if x != 0 || y != 0 || i != count-1 {
|
||||
return errors.New("invalid RFB DesktopSize rectangle")
|
||||
}
|
||||
if err := f.resize(w, h); err != nil {
|
||||
f.mu.Lock()
|
||||
err := f.resize(w, h)
|
||||
f.mu.Unlock()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
continue
|
||||
@@ -307,14 +308,15 @@ func (f *framebuffer) readUpdate() error {
|
||||
return errors.New("RFB update exceeds limit")
|
||||
}
|
||||
for dy := range h {
|
||||
if _, err := io.ReadFull(f.stream, row[:w*4]); err != nil {
|
||||
if _, err := io.ReadFull(f.stream, f.row[:w*4]); err != nil {
|
||||
return err
|
||||
}
|
||||
f.mu.Lock()
|
||||
start := (y+dy)*f.width + x
|
||||
for dx := range w {
|
||||
pixel := start + dx
|
||||
offset := pixel * 3
|
||||
f.pixels[offset], f.pixels[offset+1], f.pixels[offset+2] = row[dx*4+2], row[dx*4+1], row[dx*4]
|
||||
f.pixels[offset], f.pixels[offset+1], f.pixels[offset+2] = f.row[dx*4+2], f.row[dx*4+1], f.row[dx*4]
|
||||
if f.awaitingFull {
|
||||
mask := uint64(1) << (pixel & 63)
|
||||
if f.coverage[pixel/64]&mask == 0 {
|
||||
@@ -323,13 +325,16 @@ func (f *framebuffer) readUpdate() error {
|
||||
}
|
||||
}
|
||||
}
|
||||
f.mu.Unlock()
|
||||
}
|
||||
}
|
||||
f.mu.Lock()
|
||||
if f.awaitingFull && f.covered == f.width*f.height {
|
||||
f.complete = true
|
||||
f.awaitingFull = false
|
||||
f.lastFull.Store(time.Now().UnixNano())
|
||||
}
|
||||
f.mu.Unlock()
|
||||
f.updates.Add(1)
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -28,3 +28,19 @@ ALTER TABLE bindings ADD COLUMN IF NOT EXISTS isolation_expires_at timestamptz;
|
||||
ALTER TABLE runs ADD COLUMN IF NOT EXISTS profile_name text;
|
||||
UPDATE runs r SET profile_name=p.name FROM profiles p WHERE r.profile_id=p.id AND r.profile_name IS NULL;
|
||||
ALTER TABLE runs ALTER COLUMN profile_name SET NOT NULL;
|
||||
ALTER TABLE attempts ADD COLUMN IF NOT EXISTS claimed_at timestamptz;
|
||||
CREATE INDEX IF NOT EXISTS runs_job ON runs(job_id);
|
||||
CREATE INDEX IF NOT EXISTS attempts_run ON attempts(run_id,created_at);
|
||||
CREATE INDEX IF NOT EXISTS attempts_unfinished ON attempts(created_at) WHERE phase<>'finished';
|
||||
CREATE INDEX IF NOT EXISTS artifacts_attempt_kind ON artifacts(attempt_id,kind);
|
||||
CREATE INDEX IF NOT EXISTS uploads_owner ON uploads(owner_id);
|
||||
CREATE INDEX IF NOT EXISTS allocations_owner_active ON allocations(owner_id) WHERE state<>'deleted';
|
||||
CREATE INDEX IF NOT EXISTS media_owner ON media(owner_id);
|
||||
CREATE INDEX IF NOT EXISTS sessions_user ON sessions(user_id);
|
||||
CREATE INDEX IF NOT EXISTS sessions_expires ON sessions(expires_at);
|
||||
CREATE INDEX IF NOT EXISTS login_limits_reset ON login_limits(reset_at);
|
||||
CREATE INDEX IF NOT EXISTS revisions_profile ON revisions(profile_id,created_at DESC);
|
||||
CREATE INDEX IF NOT EXISTS revisions_source ON revisions(source_ref,created_at DESC);
|
||||
CREATE INDEX IF NOT EXISTS qualifications_profile ON qualifications(profile_id,created_at DESC);
|
||||
CREATE INDEX IF NOT EXISTS qualifications_active ON qualifications(created_at) WHERE status IN ('queued','running');
|
||||
CREATE INDEX IF NOT EXISTS events_orphan ON events(created_at) WHERE job_id IS NULL;
|
||||
|
||||
@@ -15,8 +15,13 @@ import (
|
||||
//go:embed schema.sql
|
||||
var schema embed.FS
|
||||
|
||||
func Open(ctx context.Context, url string) (*pgxpool.Pool, error) {
|
||||
p, e := pgxpool.New(ctx, url)
|
||||
func Open(ctx context.Context, url string, maxConns int32) (*pgxpool.Pool, error) {
|
||||
cfg, e := pgxpool.ParseConfig(url)
|
||||
if e != nil {
|
||||
return nil, e
|
||||
}
|
||||
cfg.MaxConns = maxConns
|
||||
p, e := pgxpool.NewWithConfig(ctx, cfg)
|
||||
if e != nil {
|
||||
return nil, e
|
||||
}
|
||||
|
||||
@@ -0,0 +1,232 @@
|
||||
package worker
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"encoding/pem"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/jackc/pgx/v5/pgxpool"
|
||||
"otche/internal/pve"
|
||||
)
|
||||
|
||||
func queuedClaimFixture(t *testing.T, ctx context.Context, db *pgxpool.Pool, names ...string) (string, []string) {
|
||||
t.Helper()
|
||||
owner, profile, revision, upload, job := newID(), newID(), newID(), newID(), newID()
|
||||
tx, err := db.Begin(ctx)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
for _, statement := range []struct {
|
||||
sql string
|
||||
args []any
|
||||
}{
|
||||
{"INSERT INTO users(id,username,password_hash,role)VALUES($1::uuid,$1::text,'unused','operator')", []any{owner}},
|
||||
{"INSERT INTO profiles(id,name,os,architecture)VALUES($1,'claim-profile','Windows','x64')", []any{profile}},
|
||||
{"INSERT INTO revisions(id,profile_id,source_ref,fingerprint,config_digest)VALUES($1,$2,'source','fp','cfg')", []any{revision, profile}},
|
||||
{"INSERT INTO uploads(id,owner_id,filename,size,sha256,storage_key)VALUES($1::uuid,$2,'safe.exe',1,'hash',$1::text)", []any{upload, owner}},
|
||||
{"INSERT INTO jobs(id,owner_id,upload_id,execution_filename,settings)VALUES($1,$2,$3,'safe.exe','{}')", []any{job, owner, upload}},
|
||||
} {
|
||||
if _, err = tx.Exec(ctx, statement.sql, statement.args...); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
ids := make([]string, len(names))
|
||||
for i, name := range names {
|
||||
run := newID()
|
||||
ids[i] = newID()
|
||||
if _, err = tx.Exec(ctx, "INSERT INTO runs(id,job_id,profile_id,revision_id,profile_name)VALUES($1,$2,$3,$4,$5)", run, job, profile, revision, name); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err = tx.Exec(ctx, "INSERT INTO attempts(id,run_id,command_id)VALUES($1,$2,$3)", ids[i], run, newID()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err = tx.Commit(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return owner, ids
|
||||
}
|
||||
|
||||
func TestClaimOwnerLimitsAndStableProfileOrder(t *testing.T) {
|
||||
ctx, db := workerTestDB(t)
|
||||
_, _ = queuedClaimFixture(t, ctx, db, "Unconfigured owner")
|
||||
ownerA, ids := queuedClaimFixture(t, ctx, db, "Zulu", "Alpha", "Beta")
|
||||
e := engine{db: db, cfg: Config{Concurrent: 3, Owners: map[string]Binding{ownerA: {MaxActiveAttempts: 2}}}}
|
||||
for _, want := range []string{ids[1], ids[2]} {
|
||||
a, err := e.claim(ctx)
|
||||
if err != nil || a.ID != want {
|
||||
t.Fatalf("claim = %s, %v; want %s", a.ID, err, want)
|
||||
}
|
||||
}
|
||||
if _, err := e.claim(ctx); !isNoRows(err) {
|
||||
t.Fatalf("owner A exceeded two live attempts: %v", err)
|
||||
}
|
||||
ownerB, idsB := queuedClaimFixture(t, ctx, db, "Other owner")
|
||||
e.cfg.Owners[ownerB] = Binding{MaxActiveAttempts: 1}
|
||||
if a, err := e.claim(ctx); err != nil || a.ID != idsB[0] {
|
||||
t.Fatalf("owner B blocked by saturated owner A: %s, %v", a.ID, err)
|
||||
}
|
||||
if _, err := e.claim(ctx); !isNoRows(err) {
|
||||
t.Fatalf("global cap exceeded: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestConcurrentClaimsRespectOwnerLimit(t *testing.T) {
|
||||
ctx, db := workerTestDB(t)
|
||||
owner, _ := queuedClaimFixture(t, ctx, db, "A", "B", "C", "D")
|
||||
e := engine{db: db, cfg: Config{Concurrent: 4, Owners: map[string]Binding{owner: {MaxActiveAttempts: 2}}}}
|
||||
results := make(chan error, 4)
|
||||
start := make(chan struct{})
|
||||
for range 4 {
|
||||
go func() {
|
||||
<-start
|
||||
_, err := e.claim(ctx)
|
||||
results <- err
|
||||
}()
|
||||
}
|
||||
close(start)
|
||||
claimed := 0
|
||||
for range 4 {
|
||||
err := <-results
|
||||
if err == nil {
|
||||
claimed++
|
||||
} else if !isNoRows(err) {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if claimed != 2 {
|
||||
t.Fatalf("concurrent claimants acquired %d leases, want 2", claimed)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReservationAdmissionIsAtomic(t *testing.T) {
|
||||
ctx, db := workerTestDB(t)
|
||||
owner, ids := queuedClaimFixture(t, ctx, db, "A", "B")
|
||||
root := t.TempDir()
|
||||
server := httptest.NewTLSServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
var data any
|
||||
switch {
|
||||
case strings.Contains(r.URL.Path, "/storage/"):
|
||||
data = map[string]any{"avail": 100 << 30, "active": 1}
|
||||
case strings.HasSuffix(r.URL.Path, "/qemu"):
|
||||
data = []any{}
|
||||
default:
|
||||
data = map[string]any{"memory": map[string]int64{"total": 128 << 30, "used": 16 << 30, "free": 112 << 30}}
|
||||
}
|
||||
_ = json.NewEncoder(w).Encode(map[string]any{"data": data})
|
||||
}))
|
||||
defer server.Close()
|
||||
ca, token := filepath.Join(root, "ca.pem"), filepath.Join(root, "token.json")
|
||||
if err := os.WriteFile(ca, pem.EncodeToMemory(&pem.Block{Type: "CERTIFICATE", Bytes: server.Certificate().Raw}), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(token, []byte(`{"token_id":"test@pve!test","secret":"test"}`), 0600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
client, err := pve.New(server.URL, ca, token)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer client.Close()
|
||||
binding := Binding{Endpoint: server.URL, Node: "test", Pool: "owner", DiskStorage: "disk", ISOStorage: "iso", MaxOwnedDiskBytes: 3 << 30, MaxOwnedArtifactBytes: 1 << 30, MinStorageFreeBytes: 1 << 30, MinMemoryFreeBytes: 12 << 30}
|
||||
e := engine{db: db, root: root, id: "reservation-test", cfg: Config{Owners: map[string]Binding{owner: binding}}}
|
||||
if _, err = db.Exec(ctx, "UPDATE attempts SET lease_owner=$1,lease_until=now()+interval '1 minute'", e.id); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
start := make(chan struct{})
|
||||
results := make(chan error, 2)
|
||||
var wg sync.WaitGroup
|
||||
for i, id := range ids {
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
<-start
|
||||
tx, err := db.Begin(ctx)
|
||||
if err != nil {
|
||||
results <- err
|
||||
return
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
err = e.reserveResources(ctx, tx, attempt{ID: id, OwnerID: owner}, binding, client, 2<<30, 9<<30, 1<<20)
|
||||
if err == nil {
|
||||
_, err = tx.Exec(ctx, `INSERT INTO allocations(id,attempt_id,owner_id,node,vmid,pool,state,metadata)VALUES($1,$2,$3,'test',$4,'owner','reserved','{"reserved_disk_bytes":2147483648,"reserved_memory_bytes":9663676416,"artifact_reservation":1048576}')`, newID(), id, owner, 8000+i)
|
||||
}
|
||||
if err == nil {
|
||||
err = tx.Commit(ctx)
|
||||
}
|
||||
results <- err
|
||||
}()
|
||||
}
|
||||
close(start)
|
||||
wg.Wait()
|
||||
close(results)
|
||||
accepted, refused := 0, 0
|
||||
for err := range results {
|
||||
if err == nil {
|
||||
accepted++
|
||||
} else if strings.HasPrefix(err.Error(), "owner disk reservation quota exceeded") {
|
||||
refused++
|
||||
} else {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if accepted != 1 || refused != 1 {
|
||||
t.Fatalf("competing reservations: accepted=%d refused=%d", accepted, refused)
|
||||
}
|
||||
// Pending clones must consume capacity before their disks/RAM materialize.
|
||||
binding.MaxOwnedDiskBytes = 1 << 40
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
disk, memory, artifact int64
|
||||
want string
|
||||
}{
|
||||
{"memory-boundary", 0, 91 << 30, 0, ""},
|
||||
{"memory-exhausted", 0, 92 << 30, 0, "node memory reservation insufficient for another clone"},
|
||||
{"pending-storage", 98 << 30, 0, 0, "storage headroom insufficient"},
|
||||
{"artifact-quota", 0, 0, 1 << 30, "owner artifact reservation quota exceeded"},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
tx, err := db.Begin(ctx)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
err = e.reserveResources(ctx, tx, attempt{ID: ids[0], OwnerID: owner}, binding, client, tc.disk, tc.memory, tc.artifact)
|
||||
if tc.want == "" {
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
} else if err == nil || !strings.HasPrefix(err.Error(), tc.want) {
|
||||
t.Fatalf("reservation error = %v; want %s", err, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestSourceLockHonorsCallerDeadline(t *testing.T) {
|
||||
ctx, db := workerTestDB(t)
|
||||
ref := "test-lock-" + newID()
|
||||
conn, err := sourceLock(ctx, db, ref)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer unlockSource(conn, ref)
|
||||
bounded, cancel := context.WithTimeout(ctx, 50*time.Millisecond)
|
||||
defer cancel()
|
||||
if extra, err := sourceLock(bounded, db, ref); err == nil {
|
||||
unlockSource(extra, ref)
|
||||
t.Fatal("contended source lock ignored caller deadline")
|
||||
} else if err.Error() != fmt.Sprintf("waiting for lock %s timed out", ref) {
|
||||
t.Fatalf("unexpected lock timeout: %v", err)
|
||||
}
|
||||
}
|
||||
@@ -58,6 +58,7 @@ type Binding struct {
|
||||
IsolationProofFile string `json:"isolation_proof_file"`
|
||||
Online *OnlinePolicy `json:"online"`
|
||||
MaxOwnedArtifactBytes int64 `json:"max_owned_artifact_bytes"`
|
||||
MaxActiveAttempts int `json:"max_active_attempts"`
|
||||
}
|
||||
|
||||
// IsolationProof is generated by operator-run negative ACL/network probes, not by
|
||||
@@ -141,6 +142,13 @@ func loadConfig(path string) (Config, error) {
|
||||
pools := map[string]bool{}
|
||||
storage := map[string]bool{}
|
||||
for owner, b := range c.Owners {
|
||||
if b.MaxActiveAttempts == 0 {
|
||||
b.MaxActiveAttempts = 1
|
||||
}
|
||||
if b.MaxActiveAttempts < 1 || b.MaxActiveAttempts > 16 || b.MaxActiveAttempts > c.Concurrent {
|
||||
return c, errors.New("max_active_attempts must be 1..16 and not exceed concurrent")
|
||||
}
|
||||
c.Owners[owner] = b
|
||||
if !uuidRE.MatchString(owner) || !ident.MatchString(b.Node) || !ident.MatchString(b.Pool) || !ident.MatchString(b.ISOStorage) || !ident.MatchString(b.DiskStorage) || b.ISOStorage == "local" || b.DiskStorage == "tank-store" {
|
||||
return c, errors.New("invalid private owner binding; shared lab storage forbidden")
|
||||
}
|
||||
|
||||
@@ -13,6 +13,7 @@ import (
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
type guestReady struct {
|
||||
@@ -79,28 +80,70 @@ func readGuestJSON(ctx context.Context, c *pve.Client, b Binding, l allocation,
|
||||
}
|
||||
return nil
|
||||
}
|
||||
func (e *engine) ready(ctx context.Context, a attempt, l allocation, b Binding, cs clients, privilege string) (guestReady, error) {
|
||||
func boundedDiagnostic(text string) string {
|
||||
var out strings.Builder
|
||||
count := 0
|
||||
for _, r := range text {
|
||||
if count == 300 {
|
||||
break
|
||||
}
|
||||
if !unicode.IsPrint(r) {
|
||||
r = ' '
|
||||
}
|
||||
out.WriteRune(r)
|
||||
count++
|
||||
}
|
||||
return out.String()
|
||||
}
|
||||
|
||||
func (e *engine) ready(ctx, parent context.Context, a attempt, l allocation, b Binding, cs clients, privilege string) (guestReady, error) {
|
||||
var ready guestReady
|
||||
var lastTransport, lastGuest string
|
||||
t := time.NewTicker(2 * time.Second)
|
||||
defer t.Stop()
|
||||
for {
|
||||
call, cancel := context.WithTimeout(ctx, 20*time.Second)
|
||||
raw, err := guestScript(call, cs.runtime, b, l, "Test-OtcheReady.ps1", "-Privilege", privilege)
|
||||
cancel()
|
||||
parsed := json.Unmarshal(raw, &ready) == nil
|
||||
if err != nil && (ctx.Err() == nil || lastTransport == "") {
|
||||
lastTransport = boundedDiagnostic(err.Error())
|
||||
}
|
||||
var current guestReady
|
||||
parsed := json.Unmarshal(raw, ¤t) == nil
|
||||
if parsed {
|
||||
ready = current
|
||||
if current.Error != "" {
|
||||
lastGuest = boundedDiagnostic(current.Error)
|
||||
}
|
||||
}
|
||||
if err == nil && parsed && ready.Ready && ready.SessionID > 0 && ready.User != "" && ready.BootID != "" && shaRE.MatchString(ready.Baseline.Fingerprint) {
|
||||
return ready, nil
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
path := filepath.Join(e.root, "attempts", a.ID, "readiness-diagnostic.json")
|
||||
if parsed {
|
||||
_ = atomicJSON(path, ready)
|
||||
diagnostic := map[string]any{"guest": ready, "last_guest_error": lastGuest, "last_transport_error": lastTransport}
|
||||
if atomicJSON(path, diagnostic) == nil {
|
||||
publish, c := context.WithTimeout(context.Background(), 5*time.Second)
|
||||
_ = e.publishFile(publish, a, "prerequisite", "readiness-diagnostic.json", "application/json", path)
|
||||
c()
|
||||
}
|
||||
return ready, errors.New("Source prerequisite failed: QGA, active interactive desktop, enabled tasks, active Defender and operator-approved signed runner script policy required; Restricted blocks the runner and is never bypassed. See readiness diagnostic when available.")
|
||||
if parent.Err() != nil {
|
||||
check, done := context.WithTimeout(context.Background(), 5*time.Second)
|
||||
fenceErr := e.fence(check, a)
|
||||
done()
|
||||
if fenceErr != nil && (fenceErr.Error() == "cancelled" || fenceErr.Error() == "attempt lease lost") {
|
||||
return ready, fenceErr
|
||||
}
|
||||
return ready, parent.Err()
|
||||
}
|
||||
detail := "no guest readiness diagnostic received"
|
||||
if lastGuest != "" {
|
||||
detail = "last guest error: " + lastGuest
|
||||
} else if lastTransport != "" {
|
||||
detail = "last transport error: " + lastTransport
|
||||
}
|
||||
return ready, fmt.Errorf("Readiness not reached within prepare_timeout_seconds=%d; %s", e.cfg.PrepareTimeoutSeconds, detail)
|
||||
case <-t.C:
|
||||
}
|
||||
}
|
||||
@@ -164,8 +207,9 @@ func (e *engine) execute(ctx context.Context, a attempt) (outcome, findings, tel
|
||||
if err = e.db.QueryRow(ctx, `SELECT COALESCE(sum(bytes),0) FROM (SELECT size bytes FROM uploads WHERE owner_id=$1 UNION ALL SELECT size FROM artifacts WHERE job_id IN (SELECT id FROM jobs WHERE owner_id=$1) UNION ALL SELECT u.size+4194304 FROM media m JOIN jobs j ON j.id=m.job_id JOIN uploads u ON u.id=j.upload_id WHERE m.owner_id=$1 UNION ALL SELECT COALESCE((metadata->>'artifact_reservation')::bigint,0) FROM allocations WHERE owner_id=$1 AND state<>'deleted') usage`, a.OwnerID).Scan(&ownedBytes); err != nil {
|
||||
return outcome, findings, telemetry, cleanup, nil, err
|
||||
}
|
||||
if a.Phase == "queued" && (b.MaxOwnedArtifactBytes <= 0 || ownedBytes+3*e.cfg.MaxArtifactBytes+e.cfg.MaxVideoBytes+grubReservation(a, e.cfg.MaxArtifactBytes)+a.Size+4194304 > b.MaxOwnedArtifactBytes) {
|
||||
return outcome, findings, telemetry, cleanup, nil, errors.New("owner input/media/artifact reservation quota exceeded including retained evidence")
|
||||
reservation := 3*e.cfg.MaxArtifactBytes + e.cfg.MaxVideoBytes + grubReservation(a, e.cfg.MaxArtifactBytes) + a.Size + 4194304
|
||||
if a.Phase == "queued" && (b.MaxOwnedArtifactBytes <= 0 || ownedBytes+reservation > b.MaxOwnedArtifactBytes) {
|
||||
return outcome, findings, telemetry, cleanup, nil, fmt.Errorf("owner input/media/artifact quota exceeded including retained evidence: used %.2f GiB + requested %.2f GiB > limit %.2f GiB", float64(ownedBytes)/(1<<30), float64(reservation)/(1<<30), float64(b.MaxOwnedArtifactBytes)/(1<<30))
|
||||
}
|
||||
var l allocation
|
||||
var session *recorder.Session
|
||||
@@ -179,6 +223,15 @@ func (e *engine) execute(ctx context.Context, a attempt) (outcome, findings, tel
|
||||
}
|
||||
final, cancel := context.WithTimeout(context.Background(), time.Duration(finalSeconds)*time.Second)
|
||||
defer cancel()
|
||||
if retErr != nil && (ctx.Err() != nil || retErr.Error() == "cancelled") {
|
||||
check, done := context.WithTimeout(final, 5*time.Second)
|
||||
fenceErr := e.fence(check, a)
|
||||
done()
|
||||
if retErr.Error() == "cancelled" || (fenceErr != nil && fenceErr.Error() == "cancelled") {
|
||||
outcome = "cancelled"
|
||||
retErr = errors.New("cancelled")
|
||||
}
|
||||
}
|
||||
if err := e.leaseAlive(final, a); err != nil {
|
||||
if session != nil {
|
||||
_ = session.Close()
|
||||
@@ -265,6 +318,8 @@ func (e *engine) execute(ctx context.Context, a attempt) (outcome, findings, tel
|
||||
}
|
||||
}
|
||||
}
|
||||
} else {
|
||||
_ = e.removeMedia(final, a, b, cs)
|
||||
}
|
||||
if recovered, observed := salvageExtracted(dir, a); recovered != nil {
|
||||
report = recovered
|
||||
@@ -288,7 +343,9 @@ func (e *engine) execute(ctx context.Context, a attempt) (outcome, findings, tel
|
||||
}
|
||||
}()
|
||||
if err = e.fence(ctx, a); err != nil {
|
||||
outcome = "cancelled"
|
||||
if err.Error() == "cancelled" {
|
||||
outcome = "cancelled"
|
||||
}
|
||||
return outcome, findings, telemetry, cleanup, nil, err
|
||||
}
|
||||
if a.Phase != "queued" {
|
||||
@@ -362,7 +419,7 @@ func (e *engine) execute(ctx context.Context, a attempt) (outcome, findings, tel
|
||||
if err = cs.provisioner.Task(prep, b.Node, upid); err != nil {
|
||||
return outcome, findings, telemetry, cleanup, nil, err
|
||||
}
|
||||
ready, err := e.ready(prep, a, l, b, cs, opts.Privilege)
|
||||
ready, err := e.ready(prep, ctx, a, l, b, cs, opts.Privilege)
|
||||
if err != nil {
|
||||
return outcome, findings, telemetry, cleanup, nil, err
|
||||
}
|
||||
@@ -471,7 +528,13 @@ func (e *engine) execute(ctx context.Context, a attempt) (outcome, findings, tel
|
||||
return outcome, findings, telemetry, cleanup, nil, err
|
||||
}
|
||||
outcome, findings, telemetry, report = r.Outcome, r.Findings, r.Telemetry, r.Report
|
||||
collectCtx, cancelCollect := context.WithTimeout(ctx, time.Duration(e.cfg.CollectTimeoutSeconds)*time.Second)
|
||||
collectionStart := time.Now()
|
||||
collectionDeadline := collectionStart.Add(time.Duration(e.cfg.CollectTimeoutSeconds) * time.Second)
|
||||
if err = e.safetyDeadline(ctx, a, l, collectionDeadline.Add(30*time.Second)); err != nil {
|
||||
telemetry = "partial"
|
||||
return outcome, findings, telemetry, cleanup, report, err
|
||||
}
|
||||
collectCtx, cancelCollect := context.WithDeadline(ctx, collectionDeadline)
|
||||
defer cancelCollect()
|
||||
if err = e.phase(collectCtx, a, "collecting"); err != nil {
|
||||
return outcome, findings, telemetry, cleanup, report, err
|
||||
|
||||
@@ -237,14 +237,15 @@ func (e *engine) extractHeld(parent context.Context, a attempt, l allocation, b
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err = storageHeadroom(ctx, cs.provisioner, b, conf.MaxDiskBytes); err != nil {
|
||||
return err
|
||||
}
|
||||
conn, err := sourceLock(ctx, e.db, "nextid-"+b.Endpoint)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer unlockSource(conn, "nextid-"+b.Endpoint)
|
||||
defer func() {
|
||||
if conn != nil {
|
||||
unlockSource(conn, "nextid-"+b.Endpoint)
|
||||
}
|
||||
}()
|
||||
var next string
|
||||
if err = cs.provisioner.Do(ctx, "GET", "/cluster/nextid", nil, &next); err != nil {
|
||||
return err
|
||||
@@ -266,10 +267,23 @@ func (e *engine) extractHeld(parent context.Context, a attempt, l allocation, b
|
||||
return errors.New("extractor VMID reserved by existing allocation")
|
||||
}
|
||||
x := extractionAllocation{ID: newID(), VMID: id, State: "reserved", DiskSlot: slot, Serial: "otche" + strings.ReplaceAll(a.ID, "-", "")[:16], DiskBytes: size}
|
||||
_, err = e.db.Exec(ctx, `INSERT INTO extractor_allocations(id,attempt_id,owner_id,node,vmid,state,disk_slot,disk_serial,disk_bytes,metadata)VALUES($1,$2,$3,$4,$5,'clone_intent',$6,$7,$8,jsonb_build_object('reserved_disk_bytes',$9::bigint,'source_vmid',$10::integer,'original_volume',$11::text))`, x.ID, a.ID, a.OwnerID, b.Node, id, slot, x.Serial, size, conf.MaxDiskBytes, conf.SourceVMID, strings.SplitN(disk, ",", 2)[0])
|
||||
tx, err := conn.Begin(ctx)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
if err = e.reserveResources(ctx, tx, a, b, cs.provisioner, conf.MaxDiskBytes, 0, 0); err != nil {
|
||||
return err
|
||||
}
|
||||
_, err = tx.Exec(ctx, `INSERT INTO extractor_allocations(id,attempt_id,owner_id,node,vmid,state,disk_slot,disk_serial,disk_bytes,metadata)VALUES($1,$2,$3,$4,$5,'clone_intent',$6,$7,$8,jsonb_build_object('reserved_disk_bytes',$9::bigint,'source_vmid',$10::integer,'original_volume',$11::text))`, x.ID, a.ID, a.OwnerID, b.Node, id, slot, x.Serial, size, conf.MaxDiskBytes, conf.SourceVMID, strings.SplitN(disk, ",", 2)[0])
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err = tx.Commit(ctx); err != nil {
|
||||
return err
|
||||
}
|
||||
unlockSource(conn, "nextid-"+b.Endpoint)
|
||||
conn = nil
|
||||
var upid string
|
||||
if err = cs.provisioner.Do(ctx, "POST", pve.VMPath(b.Node, conf.SourceVMID)+"/clone", url.Values{"newid": {strconv.Itoa(id)}, "full": {"1"}, "pool": {b.Pool}, "storage": {b.DiskStorage}, "name": {"otche-extractor-" + a.ID}}, &upid); err != nil {
|
||||
return err
|
||||
|
||||
@@ -10,25 +10,24 @@ import (
|
||||
"time"
|
||||
)
|
||||
|
||||
// A genuine lease generation regression: old workers must not regain authority
|
||||
// merely because the same process reclaims an expired attempt.
|
||||
func TestLeaseGenerationFencesExpiredClaimAndAllowsCancelledCleanup(t *testing.T) {
|
||||
func workerTestDB(t *testing.T) (context.Context, *pgxpool.Pool) {
|
||||
t.Helper()
|
||||
dsn := os.Getenv("OTCHE_TEST_DATABASE_URL")
|
||||
if dsn == "" {
|
||||
t.Skip("OTCHE_TEST_DATABASE_URL required for isolated PostgreSQL regression")
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
|
||||
defer cancel()
|
||||
t.Cleanup(cancel)
|
||||
admin, err := pgxpool.New(ctx, dsn)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer admin.Close()
|
||||
t.Cleanup(admin.Close)
|
||||
schema := "test_" + strings.ReplaceAll(newID(), "-", "")
|
||||
if _, err = admin.Exec(ctx, "CREATE SCHEMA "+schema); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer admin.Exec(context.Background(), "DROP SCHEMA "+schema+" CASCADE")
|
||||
t.Cleanup(func() { _, _ = admin.Exec(context.Background(), "DROP SCHEMA "+schema+" CASCADE") })
|
||||
cfg, err := pgxpool.ParseConfig(dsn)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
@@ -38,10 +37,16 @@ func TestLeaseGenerationFencesExpiredClaimAndAllowsCancelledCleanup(t *testing.T
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer db.Close()
|
||||
t.Cleanup(db.Close)
|
||||
if err = store.Migrate(ctx, db); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return ctx, db
|
||||
}
|
||||
|
||||
// Old workers must not regain authority when the same process reclaims a lease.
|
||||
func TestLeaseGenerationFencesExpiredClaimAndAllowsCancelledCleanup(t *testing.T) {
|
||||
ctx, db := workerTestDB(t)
|
||||
owner, profile, revision, upload, job, run, aid, command := newID(), newID(), newID(), newID(), newID(), newID(), newID(), newID()
|
||||
statements := []struct {
|
||||
sql string
|
||||
@@ -56,15 +61,19 @@ func TestLeaseGenerationFencesExpiredClaimAndAllowsCancelledCleanup(t *testing.T
|
||||
{"INSERT INTO attempts(id,run_id,command_id)VALUES($1,$2,$3)", []any{aid, run, command}},
|
||||
}
|
||||
for _, s := range statements {
|
||||
if _, err = db.Exec(ctx, s.sql, s.args...); err != nil {
|
||||
if _, err := db.Exec(ctx, s.sql, s.args...); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
e := engine{db: db, cfg: Config{Concurrent: 1}, id: "same-process"}
|
||||
e := engine{db: db, cfg: Config{Concurrent: 1, Owners: map[string]Binding{owner: {MaxActiveAttempts: 1}}}, id: "same-process"}
|
||||
first, err := e.claim(ctx)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var claimedAt time.Time
|
||||
if err = db.QueryRow(ctx, "SELECT claimed_at FROM attempts WHERE id=$1", aid).Scan(&claimedAt); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
old := e
|
||||
old.id = first.Lease
|
||||
if err = old.phase(ctx, first, "provisioning"); err != nil {
|
||||
@@ -83,6 +92,10 @@ func TestLeaseGenerationFencesExpiredClaimAndAllowsCancelledCleanup(t *testing.T
|
||||
if first.Lease == second.Lease {
|
||||
t.Fatal("reclaimed attempt reused old lease generation")
|
||||
}
|
||||
var reclaimedAt time.Time
|
||||
if err = db.QueryRow(ctx, "SELECT claimed_at FROM attempts WHERE id=$1", aid).Scan(&reclaimedAt); err != nil || !reclaimedAt.Equal(claimedAt) {
|
||||
t.Fatalf("reclaim changed initial claim timestamp: %v", err)
|
||||
}
|
||||
current := e
|
||||
current.id = second.Lease
|
||||
if old.leaseAlive(ctx, first) == nil || old.phase(ctx, first, "running") == nil {
|
||||
|
||||
@@ -174,6 +174,11 @@ func (e *engine) mediaExists(ctx context.Context, c *pve.Client, b Binding, volu
|
||||
return errMediaAbsent
|
||||
}
|
||||
func (e *engine) removeMedia(ctx context.Context, a attempt, b Binding, cs clients) error {
|
||||
conn, err := sourceLock(ctx, e.db, "media-"+a.JobID+"-"+b.Node)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer unlockSource(conn, "media-"+a.JobID+"-"+b.Node)
|
||||
var active int
|
||||
if err := e.db.QueryRow(ctx, "SELECT count(*) FROM allocations l JOIN attempts a ON a.id=l.attempt_id JOIN runs r ON r.id=a.run_id WHERE r.job_id=$1 AND l.state<>'deleted'", a.JobID).Scan(&active); err != nil {
|
||||
return err
|
||||
@@ -189,7 +194,7 @@ func (e *engine) removeMedia(ctx context.Context, a attempt, b Binding, cs clien
|
||||
return nil
|
||||
}
|
||||
var volume, state string
|
||||
err := e.db.QueryRow(ctx, "SELECT volume,state FROM media WHERE job_id=$1 AND node=$2 AND owner_id=$3", a.JobID, b.Node, a.OwnerID).Scan(&volume, &state)
|
||||
err = e.db.QueryRow(ctx, "SELECT volume,state FROM media WHERE job_id=$1 AND node=$2 AND owner_id=$3", a.JobID, b.Node, a.OwnerID).Scan(&volume, &state)
|
||||
if isNoRows(err) {
|
||||
return nil
|
||||
}
|
||||
@@ -257,3 +262,68 @@ func (e *engine) removeMedia(ctx context.Context, a attempt, b Binding, cs clien
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
// Finalizers may both see another unfinished attempt. Revisit completed jobs
|
||||
// after their final phase writes, retaining all ambiguous or referenced media.
|
||||
func (e *engine) reconcileMedia(ctx context.Context) error {
|
||||
rows, err := e.db.Query(ctx, `SELECT candidate.id,m.node FROM media m
|
||||
JOIN LATERAL (SELECT a.id FROM attempts a JOIN runs r ON r.id=a.run_id
|
||||
WHERE r.job_id=m.job_id AND a.phase='finished' AND (a.lease_until IS NULL OR a.lease_until<now())
|
||||
ORDER BY a.created_at DESC LIMIT 1) candidate ON true
|
||||
WHERE m.state IN ('ready','delete_intent')
|
||||
AND NOT EXISTS(SELECT 1 FROM attempts a JOIN runs r ON r.id=a.run_id WHERE r.job_id=m.job_id AND a.phase<>'finished')
|
||||
AND NOT EXISTS(SELECT 1 FROM allocations l JOIN attempts a ON a.id=l.attempt_id JOIN runs r ON r.id=a.run_id WHERE r.job_id=m.job_id AND l.state<>'deleted')
|
||||
ORDER BY m.job_id,m.node LIMIT 8`)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
var candidates []struct{ id, node string }
|
||||
for rows.Next() {
|
||||
var candidate struct{ id, node string }
|
||||
if err = rows.Scan(&candidate.id, &candidate.node); err != nil {
|
||||
rows.Close()
|
||||
return err
|
||||
}
|
||||
candidates = append(candidates, candidate)
|
||||
}
|
||||
rows.Close()
|
||||
if err = rows.Err(); err != nil {
|
||||
return err
|
||||
}
|
||||
var failures error
|
||||
for _, candidate := range candidates {
|
||||
action, cancel := context.WithTimeout(ctx, 120*time.Second)
|
||||
generation := *e
|
||||
generation.id = newID()
|
||||
err = generation.reconcileJobMedia(action, candidate.id, candidate.node)
|
||||
cancel()
|
||||
failures = errors.Join(failures, err)
|
||||
}
|
||||
return failures
|
||||
}
|
||||
|
||||
func (e *engine) reconcileJobMedia(ctx context.Context, id, node string) error {
|
||||
tag, err := e.db.Exec(ctx, "UPDATE attempts SET lease_owner=$2,lease_until=now()+interval '150 seconds' WHERE id=$1 AND phase='finished' AND (lease_until IS NULL OR lease_until<now())", id, e.id)
|
||||
if err != nil || tag.RowsAffected() == 0 {
|
||||
return err
|
||||
}
|
||||
defer func() {
|
||||
final, cancel := context.WithTimeout(context.Background(), 5*time.Second)
|
||||
defer cancel()
|
||||
_, _ = e.db.Exec(final, "UPDATE attempts SET lease_until=NULL WHERE id=$1 AND lease_owner=$2 AND lease_until>now()", id, e.id)
|
||||
}()
|
||||
a, err := e.load(ctx, id)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
b, ok := e.cfg.Owners[a.OwnerID]
|
||||
if !ok || b.Node != node {
|
||||
return errors.New("media owner/node binding is unavailable or changed")
|
||||
}
|
||||
cs, err := b.clients()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer cs.close()
|
||||
return e.removeMedia(ctx, a, b, cs)
|
||||
}
|
||||
|
||||
+127
-19
@@ -27,27 +27,13 @@ func (e *engine) provision(ctx context.Context, a attempt, b Binding, cs clients
|
||||
if err = checkSeal(s, a.SourceRef, a.RevisionID, a.ConfigDigest); err != nil {
|
||||
return allocation{}, err
|
||||
}
|
||||
source, dg, err := sourceConfig(ctx, cs.provisioner, s)
|
||||
source, dg, diskBytes, err := sourceConfig(ctx, cs.provisioner, s)
|
||||
if err != nil {
|
||||
return allocation{}, err
|
||||
}
|
||||
if dg != a.ConfigDigest {
|
||||
return allocation{}, errors.New("source configuration drift")
|
||||
}
|
||||
if err = storageHeadroom(ctx, cs.provisioner, b, s.MaxDiskBytes); err != nil {
|
||||
return allocation{}, err
|
||||
}
|
||||
var used int64
|
||||
if err = e.db.QueryRow(ctx, `SELECT COALESCE(sum(bytes),0) FROM (SELECT COALESCE((metadata->>'reserved_disk_bytes')::bigint,0) bytes FROM allocations WHERE owner_id=$1 AND state<>'deleted' UNION ALL SELECT COALESCE((metadata->>'reserved_disk_bytes')::bigint,0) FROM extractor_allocations WHERE owner_id=$1 AND state<>'deleted') reserved`, a.OwnerID).Scan(&used); err != nil {
|
||||
return allocation{}, err
|
||||
}
|
||||
reserve := s.MaxDiskBytes
|
||||
if e.cfg.Extractor != nil {
|
||||
reserve += e.cfg.Extractor.MaxDiskBytes
|
||||
}
|
||||
if used+reserve > b.MaxOwnedDiskBytes {
|
||||
return allocation{}, errors.New("owner disk reservation quota exceeded including retained evidence/extractor")
|
||||
}
|
||||
l, err := e.getAllocation(ctx, a)
|
||||
if err == nil {
|
||||
return l, errors.New("allocation already exists: reconcile, never issue another clone")
|
||||
@@ -59,7 +45,11 @@ func (e *engine) provision(ctx context.Context, a attempt, b Binding, cs clients
|
||||
if err != nil {
|
||||
return l, err
|
||||
}
|
||||
defer unlockSource(idConn, "nextid-"+b.Endpoint)
|
||||
defer func() {
|
||||
if idConn != nil {
|
||||
unlockSource(idConn, "nextid-"+b.Endpoint)
|
||||
}
|
||||
}()
|
||||
var suggested string
|
||||
if err = cs.provisioner.Do(ctx, "GET", "/cluster/nextid", nil, &suggested); err != nil {
|
||||
return l, err
|
||||
@@ -94,11 +84,30 @@ func (e *engine) provision(ctx context.Context, a attempt, b Binding, cs clients
|
||||
}
|
||||
}
|
||||
l = allocation{ID: newID(), VMID: id, State: "reserved"}
|
||||
metadata, _ := json.Marshal(map[string]any{"reserved_disk_bytes": s.MaxDiskBytes, "artifact_reservation": e.cfg.MaxVideoBytes + 3*e.cfg.MaxArtifactBytes + grubReservation(a, e.cfg.MaxArtifactBytes), "source_ref": a.SourceRef, "source_config": source, "source_vmid": s.VMID, "cd_slot": s.CDSlot})
|
||||
_, err = e.db.Exec(ctx, "INSERT INTO allocations(id,attempt_id,owner_id,node,vmid,pool,state,metadata) VALUES($1,$2,$3,$4,$5,$6,'reserved',$7)", l.ID, a.ID, a.OwnerID, b.Node, id, b.Pool, metadata)
|
||||
memoryMiB, err := strconv.ParseInt(pve.Text(source["memory"]), 10, 64)
|
||||
if err != nil || memoryMiB <= 0 || memoryMiB > 1<<30 {
|
||||
return l, errors.New("source memory cannot be reserved safely")
|
||||
}
|
||||
memoryBytes := memoryMiB<<20 + 1<<30
|
||||
artifactBytes := e.cfg.MaxVideoBytes + 3*e.cfg.MaxArtifactBytes + grubReservation(a, e.cfg.MaxArtifactBytes)
|
||||
tx, err := idConn.Begin(ctx)
|
||||
if err != nil {
|
||||
return l, err
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
if err = e.reserveResources(ctx, tx, a, b, cs.provisioner, diskBytes, memoryBytes, artifactBytes); err != nil {
|
||||
return l, err
|
||||
}
|
||||
metadata, _ := json.Marshal(map[string]any{"reserved_disk_bytes": diskBytes, "reserved_memory_bytes": memoryBytes, "artifact_reservation": artifactBytes, "source_ref": a.SourceRef, "source_config": source, "source_vmid": s.VMID, "cd_slot": s.CDSlot})
|
||||
_, err = tx.Exec(ctx, "INSERT INTO allocations(id,attempt_id,owner_id,node,vmid,pool,state,metadata) VALUES($1,$2,$3,$4,$5,$6,'reserved',$7)", l.ID, a.ID, a.OwnerID, b.Node, id, b.Pool, metadata)
|
||||
if err != nil {
|
||||
return l, err
|
||||
}
|
||||
if err = tx.Commit(ctx); err != nil {
|
||||
return l, err
|
||||
}
|
||||
unlockSource(idConn, "nextid-"+b.Endpoint)
|
||||
idConn = nil
|
||||
l.Metadata = metadata
|
||||
if err = e.recordAction(ctx, a, l, "clone_intent", ""); err != nil {
|
||||
return l, err
|
||||
@@ -120,7 +129,7 @@ func (e *engine) provision(ctx context.Context, a attempt, b Binding, cs clients
|
||||
if err = cs.provisioner.Task(taskCtx, s.Node, upid); err != nil {
|
||||
return l, err
|
||||
}
|
||||
_, after, err := sourceConfig(taskCtx, cs.provisioner, s)
|
||||
_, after, _, err := sourceConfig(taskCtx, cs.provisioner, s)
|
||||
if err != nil || after != dg {
|
||||
return l, errors.New("source changed during full clone; disposable evidence held")
|
||||
}
|
||||
@@ -133,6 +142,105 @@ func (e *engine) provision(ctx context.Context, a attempt, b Binding, cs clients
|
||||
l.State = "cloned"
|
||||
return l, nil
|
||||
}
|
||||
|
||||
// All owner accounting and the caller's INSERT share this transaction. The
|
||||
// artifact-root lock also serializes reservations by owners on different nodes.
|
||||
func (e *engine) reserveResources(ctx context.Context, tx pgx.Tx, a attempt, b Binding, c *pve.Client, diskBytes, memoryBytes, artifactBytes int64) error {
|
||||
if _, err := tx.Exec(ctx, "SELECT pg_advisory_xact_lock(hashtextextended('owner:'||$1,731))", a.OwnerID); err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err := tx.Exec(ctx, "SELECT pg_advisory_xact_lock(hashtextextended('artifacts:'||$1,731))", e.root); err != nil {
|
||||
return err
|
||||
}
|
||||
var used int64
|
||||
if err := tx.QueryRow(ctx, `SELECT COALESCE(sum(bytes),0) FROM (SELECT COALESCE((metadata->>'reserved_disk_bytes')::bigint,0) bytes FROM allocations WHERE owner_id=$1 AND state<>'deleted' UNION ALL SELECT COALESCE((metadata->>'reserved_disk_bytes')::bigint,0) FROM extractor_allocations WHERE owner_id=$1 AND state<>'deleted') reserved`, a.OwnerID).Scan(&used); err != nil {
|
||||
return err
|
||||
}
|
||||
if used+diskBytes > b.MaxOwnedDiskBytes {
|
||||
return fmt.Errorf("owner disk reservation quota exceeded (%.2f GiB reserved of %.2f GiB)", float64(used+diskBytes)/(1<<30), float64(b.MaxOwnedDiskBytes)/(1<<30))
|
||||
}
|
||||
if err := tx.QueryRow(ctx, `SELECT COALESCE(sum(bytes),0) FROM (SELECT size bytes FROM uploads WHERE owner_id=$1 UNION ALL SELECT size FROM artifacts WHERE job_id IN (SELECT id FROM jobs WHERE owner_id=$1) UNION ALL SELECT u.size+4194304 FROM media m JOIN jobs j ON j.id=m.job_id JOIN uploads u ON u.id=j.upload_id WHERE m.owner_id=$1 UNION ALL SELECT COALESCE((metadata->>'artifact_reservation')::bigint,0) FROM allocations WHERE owner_id=$1 AND state<>'deleted') usage`, a.OwnerID).Scan(&used); err != nil {
|
||||
return err
|
||||
}
|
||||
if used+artifactBytes > b.MaxOwnedArtifactBytes {
|
||||
return fmt.Errorf("owner artifact reservation quota exceeded (%.2f GiB reserved of %.2f GiB)", float64(used+artifactBytes)/(1<<30), float64(b.MaxOwnedArtifactBytes)/(1<<30))
|
||||
}
|
||||
if err := tx.QueryRow(ctx, `SELECT COALESCE(sum((metadata->>'artifact_reservation')::bigint),0) FROM allocations WHERE state<>'deleted'`).Scan(&used); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := artifactHeadroom(e.root, e.cfg.MinArtifactFreeBytes+used+artifactBytes); err != nil {
|
||||
return err
|
||||
}
|
||||
storageOwners := []string{}
|
||||
for owner, binding := range e.cfg.Owners {
|
||||
if binding.Endpoint == b.Endpoint && binding.Node == b.Node && binding.DiskStorage == b.DiskStorage {
|
||||
storageOwners = append(storageOwners, owner)
|
||||
}
|
||||
}
|
||||
if err := tx.QueryRow(ctx, `SELECT COALESCE(sum(bytes),0) FROM (SELECT (metadata->>'reserved_disk_bytes')::bigint bytes FROM allocations WHERE owner_id::text=ANY($1) AND state IN ('reserved','clone_intent','cloning') UNION ALL SELECT (metadata->>'reserved_disk_bytes')::bigint FROM extractor_allocations WHERE owner_id::text=ANY($1) AND state IN ('reserved','clone_intent','cloning')) pending`, storageOwners).Scan(&used); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := storageHeadroom(ctx, c, b, used+diskBytes); err != nil {
|
||||
return err
|
||||
}
|
||||
var node struct {
|
||||
Memory struct {
|
||||
Total int64 `json:"total"`
|
||||
Used int64 `json:"used"`
|
||||
} `json:"memory"`
|
||||
}
|
||||
if err := c.Do(ctx, "GET", "/nodes/"+b.Node+"/status", nil, &node); err != nil {
|
||||
return err
|
||||
}
|
||||
var vms []struct {
|
||||
VMID int `json:"vmid"`
|
||||
Status string `json:"status"`
|
||||
Mem int64 `json:"mem"`
|
||||
}
|
||||
if err := c.Do(ctx, "GET", "/nodes/"+b.Node+"/qemu", nil, &vms); err != nil {
|
||||
return err
|
||||
}
|
||||
rows, err := tx.Query(ctx, `SELECT vmid,CASE WHEN state IN ('stopped','evidence_held') THEN 0 ELSE COALESCE((metadata->>'reserved_memory_bytes')::bigint,((metadata->'source_config'->>'memory')::bigint*1048576)+1073741824,0) END FROM allocations WHERE owner_id=$1 AND node=$2 AND pool=$3 AND state<>'deleted'`, a.OwnerID, b.Node, b.Pool)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
owned := make(map[int]bool)
|
||||
var reservedMemory int64
|
||||
for rows.Next() {
|
||||
var vmid int
|
||||
var memory int64
|
||||
if err = rows.Scan(&vmid, &memory); err != nil {
|
||||
rows.Close()
|
||||
return err
|
||||
}
|
||||
owned[vmid] = true
|
||||
reservedMemory += memory
|
||||
}
|
||||
rows.Close()
|
||||
if err = rows.Err(); err != nil {
|
||||
return err
|
||||
}
|
||||
var runningMemory int64
|
||||
for _, vm := range vms {
|
||||
if vm.Status == "running" && owned[vm.VMID] {
|
||||
runningMemory += vm.Mem
|
||||
}
|
||||
}
|
||||
baseline := max(int64(0), node.Memory.Used-runningMemory)
|
||||
budget := node.Memory.Total - b.MinMemoryFreeBytes - baseline
|
||||
if reservedMemory+memoryBytes > budget {
|
||||
return errors.New("node memory reservation insufficient for another clone")
|
||||
}
|
||||
var alive bool
|
||||
if err = tx.QueryRow(ctx, "SELECT lease_owner=$2 AND lease_until>now() FROM attempts WHERE id=$1", a.ID, e.id).Scan(&alive); err != nil {
|
||||
return err
|
||||
}
|
||||
if !alive {
|
||||
return errors.New("attempt lease lost")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (e *engine) prepareClone(ctx context.Context, a attempt, l allocation, b Binding, cs clients) error {
|
||||
o := e.own(a, l, b)
|
||||
state, err := cs.provisioner.Status(ctx, b.Node, l.VMID)
|
||||
|
||||
+110
-60
@@ -10,6 +10,8 @@ import (
|
||||
"time"
|
||||
)
|
||||
|
||||
const settleJobSQL = `UPDATE jobs SET status=CASE WHEN EXISTS(SELECT 1 FROM runs WHERE job_id=$1 AND status IN ('queued','running')) THEN 'running' WHEN cancel_requested THEN 'cancelled' WHEN EXISTS(SELECT 1 FROM runs WHERE job_id=$1 AND status='failed') THEN 'failed' ELSE 'completed' END,updated_at=now() WHERE id=$1`
|
||||
|
||||
func vmAbsent(ctx context.Context, c *pve.Client, node string, id int) (bool, error) {
|
||||
var vms []struct {
|
||||
VMID int `json:"vmid"`
|
||||
@@ -158,66 +160,82 @@ func (e *engine) releaseEvidence(ctx context.Context) error {
|
||||
if err = rows.Err(); err != nil {
|
||||
return err
|
||||
}
|
||||
var failures error
|
||||
for _, id := range ids {
|
||||
a, err := e.load(ctx, id)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
b, ok := e.cfg.Owners[a.OwnerID]
|
||||
if !ok {
|
||||
return errors.New("retained owner binding is unavailable")
|
||||
}
|
||||
cs, err := b.clients()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
action, cancel := context.WithTimeout(ctx, 120*time.Second)
|
||||
tag, err := e.db.Exec(action, "UPDATE attempts SET lease_owner=$2,lease_until=now()+interval '150 seconds' WHERE id=$1 AND (lease_until IS NULL OR lease_until<now())", a.ID, e.id)
|
||||
if err != nil {
|
||||
cancel()
|
||||
cs.close()
|
||||
return err
|
||||
}
|
||||
if tag.RowsAffected() == 0 {
|
||||
cancel()
|
||||
cs.close()
|
||||
continue
|
||||
}
|
||||
l, err := e.getAllocation(action, a)
|
||||
if err == nil && (l.State == "deleting" || l.State == "delete_intent") {
|
||||
if l.UPID != nil && l.State == "deleting" {
|
||||
err = cs.provisioner.Task(action, b.Node, *l.UPID)
|
||||
}
|
||||
if err == nil {
|
||||
var absent bool
|
||||
absent, err = vmAbsent(action, cs.provisioner, b.Node, l.VMID)
|
||||
if err == nil && absent {
|
||||
_, err = e.db.Exec(action, "UPDATE allocations SET state='deleted' WHERE id=$1", l.ID)
|
||||
l.State = "deleted"
|
||||
}
|
||||
}
|
||||
}
|
||||
if err == nil && l.State != "deleted" {
|
||||
err = e.stopOwned(action, a, l, b, cs)
|
||||
if err == nil {
|
||||
err = e.deleteOwned(action, a, l, b, cs)
|
||||
}
|
||||
}
|
||||
if err == nil {
|
||||
err = e.removeMedia(action, a, b, cs)
|
||||
}
|
||||
if err == nil {
|
||||
_, err = e.db.Exec(action, "UPDATE attempts SET cleanup='complete',release_requested=false,lease_until=NULL WHERE id=$1 AND lease_owner=$2 AND lease_until>now()", a.ID, e.id)
|
||||
} else {
|
||||
_, _ = e.db.Exec(action, "UPDATE attempts SET cleanup='failed',error=$2,lease_until=NULL WHERE id=$1 AND lease_owner=$3 AND lease_until>now()", a.ID, err.Error(), e.id)
|
||||
}
|
||||
generation := *e
|
||||
generation.id = newID()
|
||||
err := generation.releaseAttempt(action, id)
|
||||
cancel()
|
||||
cs.close()
|
||||
if err != nil {
|
||||
return err
|
||||
failures = errors.Join(failures, err)
|
||||
}
|
||||
return failures
|
||||
}
|
||||
|
||||
func (e *engine) releaseAttempt(ctx context.Context, id string) (retErr error) {
|
||||
tag, err := e.db.Exec(ctx, "UPDATE attempts SET lease_owner=$2,lease_until=now()+interval '150 seconds' WHERE id=$1 AND phase='finished' AND release_requested AND cleanup IN ('evidence_held','failed') AND (lease_until IS NULL OR lease_until<now())", id, e.id)
|
||||
if err != nil || tag.RowsAffected() == 0 {
|
||||
return err
|
||||
}
|
||||
defer func() {
|
||||
if retErr == nil {
|
||||
return
|
||||
}
|
||||
// The action may have timed out; persist the failure using the still-live lease.
|
||||
final, cancel := context.WithTimeout(context.Background(), 5*time.Second)
|
||||
defer cancel()
|
||||
_, err := e.db.Exec(final, `WITH failed AS (
|
||||
UPDATE attempts SET cleanup='failed',release_requested=false,lease_until=NULL
|
||||
WHERE id=$1 AND lease_owner=$2 AND lease_until>now() RETURNING id,run_id
|
||||
) INSERT INTO events(job_id,attempt_id,kind,message)
|
||||
SELECT r.job_id,f.id,'release_error','Evidence release failed; resources retained for operator inspection. Request release again after resolving the problem.'
|
||||
FROM failed f JOIN runs r ON r.id=f.run_id`, id, e.id)
|
||||
retErr = errors.Join(retErr, err)
|
||||
}()
|
||||
a, err := e.load(ctx, id)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
b, ok := e.cfg.Owners[a.OwnerID]
|
||||
if !ok {
|
||||
return errors.New("retained owner binding is unavailable")
|
||||
}
|
||||
cs, err := b.clients()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer cs.close()
|
||||
l, err := e.getAllocation(ctx, a)
|
||||
if err == nil && (l.State == "deleting" || l.State == "delete_intent") {
|
||||
if l.UPID != nil && l.State == "deleting" {
|
||||
err = cs.provisioner.Task(ctx, b.Node, *l.UPID)
|
||||
}
|
||||
if err == nil {
|
||||
var absent bool
|
||||
absent, err = vmAbsent(ctx, cs.provisioner, b.Node, l.VMID)
|
||||
if err == nil && absent {
|
||||
_, err = e.db.Exec(ctx, "UPDATE allocations SET state='deleted' WHERE id=$1", l.ID)
|
||||
l.State = "deleted"
|
||||
}
|
||||
}
|
||||
}
|
||||
return nil
|
||||
if err == nil && l.State != "deleted" {
|
||||
err = e.stopOwned(ctx, a, l, b, cs)
|
||||
if err == nil {
|
||||
err = e.deleteOwned(ctx, a, l, b, cs)
|
||||
}
|
||||
}
|
||||
if err == nil {
|
||||
err = e.removeMedia(ctx, a, b, cs)
|
||||
}
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
tag, err = e.db.Exec(ctx, "UPDATE attempts SET cleanup='complete',release_requested=false,lease_until=NULL WHERE id=$1 AND lease_owner=$2 AND lease_until>now()", a.ID, e.id)
|
||||
if err == nil && tag.RowsAffected() != 1 {
|
||||
return errors.New("attempt lease lost")
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
// Watchdog must be a separate supervised process. It does not record, dispatch,
|
||||
@@ -231,8 +249,19 @@ func Watchdog(ctx context.Context, db *pgxpool.Pool, configPath string) error {
|
||||
e := &engine{db: db, cfg: cfg, id: "watchdog-" + newID()}
|
||||
tick := time.NewTicker(5 * time.Second)
|
||||
defer tick.Stop()
|
||||
var lastPrune time.Time
|
||||
var pruneError error
|
||||
for {
|
||||
blockers := []string{}
|
||||
if time.Since(lastPrune) >= 10*time.Minute {
|
||||
prune, cancel := context.WithTimeout(ctx, 15*time.Second)
|
||||
_, pruneError = db.Exec(prune, "DELETE FROM sessions WHERE expires_at<now(); DELETE FROM login_limits WHERE reset_at<now(); DELETE FROM events WHERE job_id IS NULL AND created_at<now()-interval '14 days'")
|
||||
cancel()
|
||||
lastPrune = time.Now()
|
||||
}
|
||||
if pruneError != nil {
|
||||
blockers = append(blockers, "Watchdog expired session/login/event pruning failed")
|
||||
}
|
||||
rows, err := db.Query(ctx, `SELECT a.id FROM attempts a JOIN allocations l ON l.attempt_id=a.id WHERE (l.state NOT IN ('deleted','evidence_held','stopped') OR EXISTS(SELECT 1 FROM extractor_allocations x WHERE x.attempt_id=a.id AND x.state<>'deleted')) AND (a.lease_until IS NULL OR a.lease_until<now() OR (l.metadata->>'safety_deadline')::timestamptz<now()) ORDER BY a.created_at LIMIT 32`)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -293,11 +322,10 @@ func Watchdog(ctx context.Context, db *pgxpool.Pool, configPath string) error {
|
||||
}
|
||||
}
|
||||
if err == nil {
|
||||
_, err = db.Exec(action, "UPDATE attempts SET phase='finished',outcome='interrupted',telemetry='partial',cleanup='evidence_held',error='Independent watchdog stopped expired attempt; evidence retained',finished_at=now(),lease_until=NULL WHERE id=$1 AND lease_owner=$2 AND lease_until>now()", id, generation.id)
|
||||
_, _ = db.Exec(action, "UPDATE runs SET status='failed' WHERE id=$1", a.RunID)
|
||||
_, _ = db.Exec(action, "UPDATE jobs SET status='failed',updated_at=now() WHERE id=$1", a.JobID)
|
||||
} else {
|
||||
blockers = append(blockers, "Watchdog could not establish owned VM stop: "+err.Error())
|
||||
err = generation.finishInterrupted(action, a)
|
||||
}
|
||||
if err != nil {
|
||||
blockers = append(blockers, "Watchdog could not finish owned VM recovery: "+err.Error())
|
||||
}
|
||||
cs.close()
|
||||
cancel()
|
||||
@@ -314,3 +342,25 @@ func Watchdog(ctx context.Context, db *pgxpool.Pool, configPath string) error {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (e *engine) finishInterrupted(ctx context.Context, a attempt) error {
|
||||
tx, err := e.db.Begin(ctx)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer tx.Rollback(ctx)
|
||||
tag, err := tx.Exec(ctx, `UPDATE attempts SET phase='finished',outcome=CASE WHEN (SELECT cancel_requested FROM jobs WHERE id=$3) THEN 'cancelled' ELSE 'interrupted' END,telemetry='partial',cleanup='evidence_held',error='Independent watchdog stopped expired attempt; evidence retained',finished_at=now(),lease_until=NULL WHERE id=$1 AND lease_owner=$2 AND lease_until>now()`, a.ID, e.id, a.JobID)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if tag.RowsAffected() != 1 {
|
||||
return errors.New("attempt lease lost")
|
||||
}
|
||||
if _, err = tx.Exec(ctx, "UPDATE runs SET status=CASE WHEN (SELECT cancel_requested FROM jobs WHERE id=$2) THEN 'cancelled' ELSE 'failed' END WHERE id=$1", a.RunID, a.JobID); err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err = tx.Exec(ctx, settleJobSQL, a.JobID); err != nil {
|
||||
return err
|
||||
}
|
||||
return tx.Commit(ctx)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,170 @@
|
||||
package worker
|
||||
|
||||
import (
|
||||
"context"
|
||||
"github.com/jackc/pgx/v5/pgxpool"
|
||||
"os"
|
||||
"otche/internal/store"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func recoveryDatabase(t *testing.T) (*pgxpool.Pool, context.Context) {
|
||||
t.Helper()
|
||||
dsn := os.Getenv("OTCHE_TEST_DATABASE_URL")
|
||||
if dsn == "" {
|
||||
t.Skip("OTCHE_TEST_DATABASE_URL required for isolated PostgreSQL regression")
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
|
||||
t.Cleanup(cancel)
|
||||
admin, err := pgxpool.New(ctx, dsn)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(admin.Close)
|
||||
schema := "test_" + strings.ReplaceAll(newID(), "-", "")
|
||||
if _, err = admin.Exec(ctx, "CREATE SCHEMA "+schema); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(func() {
|
||||
cleanup, done := context.WithTimeout(context.Background(), 10*time.Second)
|
||||
defer done()
|
||||
_, _ = admin.Exec(cleanup, "DROP SCHEMA "+schema+" CASCADE")
|
||||
})
|
||||
cfg, err := pgxpool.ParseConfig(dsn)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
cfg.ConnConfig.RuntimeParams["search_path"] = schema
|
||||
db, err := pgxpool.NewWithConfig(ctx, cfg)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(db.Close)
|
||||
if err = store.Migrate(ctx, db); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return db, ctx
|
||||
}
|
||||
|
||||
func recoveryAttempt(t *testing.T, ctx context.Context, db *pgxpool.Pool) attempt {
|
||||
t.Helper()
|
||||
a := attempt{ID: newID(), RunID: newID(), JobID: newID(), OwnerID: newID(), ProfileID: newID(), RevisionID: newID(), CommandID: newID()}
|
||||
upload := newID()
|
||||
statements := []struct {
|
||||
sql string
|
||||
args []any
|
||||
}{
|
||||
{"INSERT INTO users(id,username,password_hash,role)VALUES($1,$2,'unused','operator')", []any{a.OwnerID, a.OwnerID}},
|
||||
{"INSERT INTO profiles(id,name,os,architecture)VALUES($1,'recovery-profile','Windows','x64')", []any{a.ProfileID}},
|
||||
{"INSERT INTO revisions(id,profile_id,source_ref,fingerprint,config_digest)VALUES($1,$2,'source','fp','cfg')", []any{a.RevisionID, a.ProfileID}},
|
||||
{"INSERT INTO uploads(id,owner_id,filename,size,sha256,storage_key)VALUES($1,$2,'safe.exe',1,'hash',$3)", []any{upload, a.OwnerID, upload}},
|
||||
{"INSERT INTO jobs(id,owner_id,upload_id,execution_filename,settings)VALUES($1,$2,$3,'safe.exe','{}')", []any{a.JobID, a.OwnerID, upload}},
|
||||
{"INSERT INTO runs(id,job_id,profile_id,revision_id,profile_name)VALUES($1,$2,$3,$4,'Recovery fixture')", []any{a.RunID, a.JobID, a.ProfileID, a.RevisionID}},
|
||||
{"INSERT INTO attempts(id,run_id,command_id)VALUES($1,$2,$3)", []any{a.ID, a.RunID, a.CommandID}},
|
||||
}
|
||||
for _, s := range statements {
|
||||
if _, err := db.Exec(ctx, s.sql, s.args...); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return a
|
||||
}
|
||||
|
||||
func TestSettleJobAccountsForAllRunsAndCancellation(t *testing.T) {
|
||||
db, ctx := recoveryDatabase(t)
|
||||
a := recoveryAttempt(t, ctx, db)
|
||||
second := newID()
|
||||
if _, err := db.Exec(ctx, "INSERT INTO runs(id,job_id,profile_id,revision_id,profile_name)VALUES($1,$2,$3,$4,'Second fixture')", second, a.JobID, a.ProfileID, a.RevisionID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, tc := range []struct {
|
||||
first, second string
|
||||
cancelled bool
|
||||
want string
|
||||
}{
|
||||
{"failed", "queued", false, "running"},
|
||||
{"failed", "running", true, "running"},
|
||||
{"failed", "cancelled", true, "cancelled"},
|
||||
{"failed", "completed", false, "failed"},
|
||||
{"completed", "completed", false, "completed"},
|
||||
} {
|
||||
if _, err := db.Exec(ctx, "UPDATE runs SET status=CASE WHEN id=$1 THEN $3 ELSE $4 END WHERE job_id=$2", a.RunID, a.JobID, tc.first, tc.second); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := db.Exec(ctx, "UPDATE jobs SET cancel_requested=$2 WHERE id=$1", a.JobID, tc.cancelled); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := db.Exec(ctx, settleJobSQL, a.JobID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var status string
|
||||
if err := db.QueryRow(ctx, "SELECT status FROM jobs WHERE id=$1", a.JobID).Scan(&status); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if status != tc.want {
|
||||
t.Fatalf("runs=(%s,%s) cancelled=%v: got %s, want %s", tc.first, tc.second, tc.cancelled, status, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestReleaseFailurePreservesEvidenceReasonAndProcessesNext(t *testing.T) {
|
||||
db, ctx := recoveryDatabase(t)
|
||||
first := recoveryAttempt(t, ctx, db)
|
||||
second := recoveryAttempt(t, ctx, db)
|
||||
for _, a := range []attempt{first, second} {
|
||||
if _, err := db.Exec(ctx, "UPDATE attempts SET phase='finished',cleanup='evidence_held',release_requested=true,error='original observation failure' WHERE id=$1", a.ID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
// Both fail deterministically before PVE access: no owner credentials exist.
|
||||
e := engine{db: db, cfg: Config{}, id: "release-worker"}
|
||||
if err := e.releaseEvidence(ctx); err == nil {
|
||||
t.Fatal("missing owner binding accepted")
|
||||
}
|
||||
for _, a := range []attempt{first, second} {
|
||||
var cleanup, reason string
|
||||
var requested bool
|
||||
if err := db.QueryRow(ctx, "SELECT cleanup,error,release_requested FROM attempts WHERE id=$1", a.ID).Scan(&cleanup, &reason, &requested); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if cleanup != "failed" || requested || reason != "original observation failure" {
|
||||
t.Fatalf("release failure not isolated: cleanup=%s requested=%v original reason=%q", cleanup, requested, reason)
|
||||
}
|
||||
var count int
|
||||
if err := db.QueryRow(ctx, "SELECT count(*) FROM events WHERE attempt_id=$1 AND job_id=$2 AND kind='release_error' AND length(message)<=300", a.ID, a.JobID).Scan(&count); err != nil || count != 1 {
|
||||
t.Fatalf("bounded job release event missing: count=%d error=%v", count, err)
|
||||
}
|
||||
}
|
||||
if err := e.releaseEvidence(ctx); err != nil {
|
||||
t.Fatalf("failed releases retried without administrator request: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWatchdogSettlementHonorsCancellationAndLease(t *testing.T) {
|
||||
db, ctx := recoveryDatabase(t)
|
||||
a := recoveryAttempt(t, ctx, db)
|
||||
e := engine{db: db, id: "watchdog-generation"}
|
||||
if _, err := db.Exec(ctx, "UPDATE attempts SET lease_owner=$2,lease_until=now()+interval '1 minute' WHERE id=$1", a.ID, e.id); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := db.Exec(ctx, "UPDATE jobs SET cancel_requested=true WHERE id=$1", a.JobID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stale := e
|
||||
stale.id = "expired-generation"
|
||||
if err := stale.finishInterrupted(ctx, a); err == nil {
|
||||
t.Fatal("stale watchdog settled attempt")
|
||||
}
|
||||
if err := e.finishInterrupted(ctx, a); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var outcome, runStatus, jobStatus string
|
||||
if err := db.QueryRow(ctx, "SELECT a.outcome,r.status,j.status FROM attempts a JOIN runs r ON r.id=a.run_id JOIN jobs j ON j.id=r.job_id WHERE a.id=$1", a.ID).Scan(&outcome, &runStatus, &jobStatus); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if outcome != "cancelled" || runStatus != "cancelled" || jobStatus != "cancelled" {
|
||||
t.Fatalf("watchdog lost cancellation: outcome=%s run=%s job=%s", outcome, runStatus, jobStatus)
|
||||
}
|
||||
}
|
||||
+24
-18
@@ -59,34 +59,34 @@ func atomicJSON(path string, v any) error {
|
||||
}
|
||||
return os.Rename(name, path)
|
||||
}
|
||||
func sourceConfig(ctx context.Context, c *pve.Client, s Source) (map[string]any, string, error) {
|
||||
func sourceConfig(ctx context.Context, c *pve.Client, s Source) (map[string]any, string, int64, error) {
|
||||
state, err := c.Status(ctx, s.Node, s.VMID)
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
return nil, "", 0, err
|
||||
}
|
||||
if state != "stopped" {
|
||||
return nil, "", errors.New("source must already be stopped; worker never stops a master")
|
||||
return nil, "", 0, errors.New("source must already be stopped; worker never stops a master")
|
||||
}
|
||||
cfg, err := c.Config(ctx, s.Node, s.VMID)
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
return nil, "", 0, err
|
||||
}
|
||||
if pve.Text(cfg["template"]) != "" && pve.Text(cfg["template"]) != "0" {
|
||||
return nil, "", errors.New("source must be ordinary VM, template=0")
|
||||
return nil, "", 0, errors.New("source must be ordinary VM, template=0")
|
||||
}
|
||||
if err = c.NoPending(ctx, s.Node, s.VMID, ""); err != nil {
|
||||
return nil, "", err
|
||||
return nil, "", 0, err
|
||||
}
|
||||
if pve.Text(cfg["lock"]) != "" {
|
||||
return nil, "", errors.New("source has active PVE lock")
|
||||
return nil, "", 0, errors.New("source has active PVE lock")
|
||||
}
|
||||
for k, v := range cfg {
|
||||
str := pve.Text(v)
|
||||
if strings.HasPrefix(k, "hostpci") || strings.HasPrefix(k, "usb") || k == "args" || k == "hookscript" || strings.HasPrefix(k, "virtiofs") || strings.HasPrefix(k, "unused") {
|
||||
return nil, "", errors.New("source contains unapproved passthrough, hook or unused disk")
|
||||
return nil, "", 0, errors.New("source contains unapproved passthrough, hook or unused disk")
|
||||
}
|
||||
if (k == s.CDSlot || diskHas(str, "media", "cdrom")) && !cdromMatches(str, "none") {
|
||||
return nil, "", errors.New("source has inserted or ambiguous CD media")
|
||||
return nil, "", 0, errors.New("source has inserted or ambiguous CD media")
|
||||
}
|
||||
if strings.HasPrefix(k, "net") {
|
||||
bridge := ""
|
||||
@@ -102,7 +102,7 @@ func sourceConfig(ctx context.Context, c *pve.Client, s Source) (map[string]any,
|
||||
}
|
||||
}
|
||||
if !allowed {
|
||||
return nil, "", errors.New("source inherited network is not an approved lab bridge")
|
||||
return nil, "", 0, errors.New("source inherited network is not an approved lab bridge")
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -111,32 +111,38 @@ func sourceConfig(ctx context.Context, c *pve.Client, s Source) (map[string]any,
|
||||
if !cdromMatches(pve.Text(value), "none") && (strings.HasPrefix(key, "scsi") && key != "scsihw" || strings.HasPrefix(key, "sata") || strings.HasPrefix(key, "ide") || strings.HasPrefix(key, "virtio") || strings.HasPrefix(key, "efidisk") || strings.HasPrefix(key, "tpmstate")) {
|
||||
size, sizeErr := diskSize(pve.Text(value))
|
||||
if sizeErr != nil {
|
||||
return nil, "", errors.New("source disk size cannot be reserved safely")
|
||||
return nil, "", 0, errors.New("source disk size cannot be reserved safely")
|
||||
}
|
||||
diskBytes += size
|
||||
}
|
||||
}
|
||||
if diskBytes <= 0 || diskBytes > s.MaxDiskBytes {
|
||||
return nil, "", errors.New("source current disk capacity exceeds configured full-clone reservation")
|
||||
return nil, "", 0, errors.New("source current disk capacity exceeds configured full-clone reservation")
|
||||
}
|
||||
if pve.Text(cfg["agent"]) == "" || pve.Text(cfg["agent"]) == "0" || pve.Text(cfg["vga"]) == "none" {
|
||||
return nil, "", errors.New("source requires QGA and graphical console")
|
||||
return nil, "", 0, errors.New("source requires QGA and graphical console")
|
||||
}
|
||||
delete(cfg, "digest")
|
||||
return cfg, digest(cfg), nil
|
||||
return cfg, digest(cfg), diskBytes, nil
|
||||
}
|
||||
func sourceLock(ctx context.Context, db *pgxpool.Pool, ref string) (*pgxpool.Conn, error) {
|
||||
bounded, cancel := context.WithTimeout(ctx, 30*time.Second)
|
||||
bounded, cancel := context.WithTimeout(ctx, 20*time.Minute)
|
||||
defer cancel()
|
||||
lockError := func(err error) error {
|
||||
if errors.Is(bounded.Err(), context.DeadlineExceeded) {
|
||||
return fmt.Errorf("waiting for lock %s timed out", ref)
|
||||
}
|
||||
return err
|
||||
}
|
||||
conn, err := db.Acquire(bounded)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
return nil, lockError(err)
|
||||
}
|
||||
_, err = conn.Exec(bounded, "SELECT pg_advisory_lock(hashtextextended($1,731))", ref)
|
||||
if err != nil {
|
||||
conn.Conn().Close(context.Background())
|
||||
conn.Release()
|
||||
return nil, err
|
||||
return nil, lockError(err)
|
||||
}
|
||||
return conn, nil
|
||||
}
|
||||
@@ -183,7 +189,7 @@ func ValidateSource(ctx context.Context, db *pgxpool.Pool, configPath, profileID
|
||||
return "", err
|
||||
}
|
||||
defer cs.close()
|
||||
_, dg, err := sourceConfig(ctx, cs.provisioner, s)
|
||||
_, dg, _, err := sourceConfig(ctx, cs.provisioner, s)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
|
||||
+49
-20
@@ -4,7 +4,6 @@ import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"github.com/jackc/pgx/v5"
|
||||
"github.com/jackc/pgx/v5/pgxpool"
|
||||
"otche/internal/pve"
|
||||
@@ -79,15 +78,18 @@ func Run(ctx context.Context, db *pgxpool.Pool, artifactRoot, configPath string)
|
||||
return err
|
||||
}
|
||||
if len(blockers) == 0 {
|
||||
select {
|
||||
case slots <- struct{}{}:
|
||||
a, claimErr := e.claim(ctx)
|
||||
if claimErr != nil {
|
||||
<-slots
|
||||
if !errors.Is(claimErr, pgx.ErrNoRows) {
|
||||
return claimErr
|
||||
claimSlots:
|
||||
for {
|
||||
select {
|
||||
case slots <- struct{}{}:
|
||||
a, claimErr := e.claim(ctx)
|
||||
if claimErr != nil {
|
||||
<-slots
|
||||
if !errors.Is(claimErr, pgx.ErrNoRows) {
|
||||
return claimErr
|
||||
}
|
||||
break claimSlots
|
||||
}
|
||||
} else {
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
@@ -96,13 +98,17 @@ func Run(ctx context.Context, db *pgxpool.Pool, artifactRoot, configPath string)
|
||||
claimed.id = a.Lease
|
||||
claimed.runClaim(ctx, a)
|
||||
}()
|
||||
default:
|
||||
break claimSlots
|
||||
}
|
||||
default:
|
||||
}
|
||||
}
|
||||
if err = e.releaseEvidence(ctx); err != nil {
|
||||
e.healthError(ctx, "Evidence release requires operator attention")
|
||||
}
|
||||
if err = e.reconcileMedia(ctx); err != nil {
|
||||
e.healthError(ctx, "Job media cleanup requires operator attention")
|
||||
}
|
||||
if len(blockers) == 0 {
|
||||
if err = e.qualify(ctx); err != nil {
|
||||
e.healthError(ctx, err.Error())
|
||||
@@ -115,7 +121,20 @@ func Run(ctx context.Context, db *pgxpool.Pool, artifactRoot, configPath string)
|
||||
}
|
||||
}
|
||||
}
|
||||
var healthMessages = struct {
|
||||
sync.Mutex
|
||||
last map[string]time.Time
|
||||
}{last: make(map[string]time.Time)}
|
||||
|
||||
func (e *engine) healthError(ctx context.Context, message string) {
|
||||
healthMessages.Lock()
|
||||
now := time.Now()
|
||||
if last, ok := healthMessages.last[message]; ok && now.Sub(last) < 10*time.Minute {
|
||||
healthMessages.Unlock()
|
||||
return
|
||||
}
|
||||
healthMessages.last[message] = now
|
||||
healthMessages.Unlock()
|
||||
_, _ = e.db.Exec(ctx, "INSERT INTO events(kind,message) VALUES('worker', $1)", message)
|
||||
}
|
||||
func (e *engine) claim(ctx context.Context) (attempt, error) {
|
||||
@@ -135,12 +154,20 @@ func (e *engine) claim(ctx context.Context) (attempt, error) {
|
||||
return attempt{}, pgx.ErrNoRows
|
||||
}
|
||||
var id string
|
||||
err = tx.QueryRow(ctx, `SELECT a.id FROM attempts a JOIN runs r ON r.id=a.run_id JOIN jobs j ON j.id=r.job_id WHERE a.phase<>'finished' AND (a.lease_until IS NULL OR a.lease_until<now()) AND NOT EXISTS(SELECT 1 FROM attempts b JOIN runs br ON br.id=b.run_id JOIN jobs bj ON bj.id=br.job_id WHERE bj.owner_id=j.owner_id AND b.lease_until>now() AND b.phase<>'finished') ORDER BY a.created_at FOR UPDATE OF a SKIP LOCKED LIMIT 1`).Scan(&id)
|
||||
limits := make(map[string]int, len(e.cfg.Owners))
|
||||
for owner, binding := range e.cfg.Owners {
|
||||
limits[owner] = binding.MaxActiveAttempts
|
||||
}
|
||||
bindings, err := json.Marshal(limits)
|
||||
if err != nil {
|
||||
return attempt{}, err
|
||||
}
|
||||
err = tx.QueryRow(ctx, `SELECT a.id FROM attempts a JOIN runs r ON r.id=a.run_id JOIN jobs j ON j.id=r.job_id JOIN jsonb_each_text($1::jsonb) limits ON limits.key=j.owner_id::text WHERE a.phase<>'finished' AND (a.lease_until IS NULL OR a.lease_until<now()) AND (SELECT count(*) FROM attempts b JOIN runs br ON br.id=b.run_id JOIN jobs bj ON bj.id=br.job_id WHERE bj.owner_id=j.owner_id AND b.lease_until>now() AND b.phase<>'finished') < limits.value::integer ORDER BY a.created_at,r.profile_name,r.id FOR UPDATE OF a SKIP LOCKED LIMIT 1`, bindings).Scan(&id)
|
||||
if err != nil {
|
||||
return attempt{}, err
|
||||
}
|
||||
generation := newID()
|
||||
_, err = tx.Exec(ctx, "UPDATE attempts SET lease_owner=$2,lease_until=now()+interval '30 seconds' WHERE id=$1", id, generation)
|
||||
_, err = tx.Exec(ctx, "UPDATE attempts SET lease_owner=$2,lease_until=now()+interval '30 seconds',claimed_at=COALESCE(claimed_at,now()) WHERE id=$1", id, generation)
|
||||
if err != nil {
|
||||
return attempt{}, err
|
||||
}
|
||||
@@ -220,6 +247,15 @@ func (e *engine) runClaim(parent context.Context, a attempt) {
|
||||
if len(msg) > 1500 {
|
||||
msg = msg[:1500]
|
||||
}
|
||||
if outcome == "error" || outcome == "interrupted" {
|
||||
var requested bool
|
||||
if finishErr := e.db.QueryRow(final, "SELECT cancel_requested FROM jobs WHERE id=$1", a.JobID).Scan(&requested); finishErr != nil {
|
||||
return
|
||||
}
|
||||
if requested {
|
||||
outcome = "cancelled"
|
||||
}
|
||||
}
|
||||
tag, finishErr := e.db.Exec(final, `UPDATE attempts SET phase='finished',outcome=$3,findings=$4,telemetry=$5,cleanup=$6,error=$7,report=$8,finished_at=now(),lease_until=NULL WHERE id=$1 AND lease_owner=$2 AND lease_until>now()`, a.ID, e.id, outcome, findings, telemetry, cleanup, msg, report)
|
||||
if finishErr != nil || tag.RowsAffected() == 0 {
|
||||
return
|
||||
@@ -231,7 +267,7 @@ func (e *engine) runClaim(parent context.Context, a attempt) {
|
||||
status = "failed"
|
||||
}
|
||||
_, _ = e.db.Exec(final, "UPDATE runs SET status=$2 WHERE id=$1", a.RunID, status)
|
||||
_, _ = e.db.Exec(final, `UPDATE jobs SET status=CASE WHEN EXISTS(SELECT 1 FROM runs WHERE job_id=$1 AND status IN ('queued','running')) THEN 'running' WHEN cancel_requested THEN 'cancelled' WHEN EXISTS(SELECT 1 FROM runs WHERE job_id=$1 AND status='failed') THEN 'failed' ELSE 'completed' END,updated_at=now() WHERE id=$1`, a.JobID)
|
||||
_, _ = e.db.Exec(final, settleJobSQL, a.JobID)
|
||||
}
|
||||
func (e *engine) getAllocation(ctx context.Context, a attempt) (allocation, error) {
|
||||
var l allocation
|
||||
@@ -259,10 +295,3 @@ func (e *engine) recordAction(ctx context.Context, a attempt, l allocation, stat
|
||||
}
|
||||
return err
|
||||
}
|
||||
func (e *engine) deadline(ctx context.Context, a attempt, t time.Time) error {
|
||||
tag, err := e.db.Exec(ctx, "UPDATE attempts SET deadline_at=$3 WHERE id=$1 AND lease_owner=$2", a.ID, e.id, t)
|
||||
if err == nil && tag.RowsAffected() != 1 {
|
||||
return fmt.Errorf("attempt lease lost")
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
@@ -28,6 +28,7 @@
|
||||
"min_memory_free_bytes": 12884901888,
|
||||
"max_owned_disk_bytes": 343597383680,
|
||||
"max_owned_artifact_bytes": 8589934592,
|
||||
"max_active_attempts": 1,
|
||||
"online": null
|
||||
}
|
||||
},
|
||||
|
||||
+22
-11
@@ -10,17 +10,20 @@ $manifest=Read-OtcheJson ([string]$pointer.manifest_path)
|
||||
Confirm-OtcheManifest $manifest
|
||||
$directory="$root\results\$($manifest.command_id)"
|
||||
$receipt=Read-OtcheJson "$directory\receipt.json"
|
||||
$since=([DateTime]::Parse($receipt.accepted_at)).ToUniversalTime().AddMinutes(-2)
|
||||
$since=[DateTime]::Parse([string]$receipt.accepted_at,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime().AddMinutes(-2)
|
||||
$clock=[Diagnostics.Stopwatch]::StartNew()
|
||||
$errors=New-Object 'Collections.Generic.List[string]'
|
||||
$detections=New-Object 'Collections.Generic.List[object]'
|
||||
$seen=New-Object 'Collections.Generic.HashSet[string]'
|
||||
$events=New-Object 'Collections.Generic.List[object]'
|
||||
$before=$null
|
||||
try {$before=Get-OtcheDefender} catch {$errors.Add('Defender baseline: '+$_.Exception.Message)}
|
||||
$environment=$null
|
||||
try {$environment=Get-OtcheEnvironment} catch {$errors.Add('Environment: '+$_.Exception.Message)}
|
||||
$boot=Get-OtcheBootId
|
||||
$before=$null
|
||||
try {$before=Get-OtcheDefender} catch {$errors.Add('Defender baseline: '+$_.Exception.Message)}
|
||||
if ($null -ne $before) {Write-OtcheJson "$directory\telemetry.json" ([ordered]@{updated_at=(Get-OtcheUtc);boot_id=$boot;before=$before;after=$before;drift=$false;detections=@();environment=$environment;errors=@($errors.ToArray());initial=$true})}
|
||||
$eventCursors=@{}
|
||||
$eventSince=$since.ToString('yyyy-MM-ddTHH:mm:ss.fffZ',[Globalization.CultureInfo]::InvariantCulture)
|
||||
$runtimeStatus=$null
|
||||
$candidatePaths=New-Object 'Collections.Generic.List[string]'
|
||||
$candidatePaths.Add("$root\samples\$($manifest.command_id)\$($manifest.filename)")
|
||||
@@ -35,9 +38,9 @@ function Test-ResourceMatch([string]$Text) {
|
||||
}
|
||||
function Get-Stage([DateTime]$Time) {
|
||||
$r=Read-OtcheJson "$directory\receipt.json"
|
||||
if ($null -ne $runtimeStatus -and $null -ne $runtimeStatus.report.execution.finished_at -and $Time.ToUniversalTime() -ge [DateTime]::Parse($runtimeStatus.report.execution.finished_at).ToUniversalTime()) {return 'collection'}
|
||||
if ($null -ne $r.started_at -and $Time.ToUniversalTime() -ge [DateTime]::Parse($r.started_at).ToUniversalTime()) {return 'execution'}
|
||||
if ($Time.ToUniversalTime() -lt [DateTime]::Parse($r.accepted_at).ToUniversalTime()) {return 'delivery'}
|
||||
if ($null -ne $runtimeStatus -and $null -ne $runtimeStatus.report.execution.finished_at -and $Time.ToUniversalTime() -ge [DateTime]::Parse([string]$runtimeStatus.report.execution.finished_at,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime()) {return 'collection'}
|
||||
if ($null -ne $r.started_at -and $Time.ToUniversalTime() -ge [DateTime]::Parse([string]$r.started_at,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime()) {return 'execution'}
|
||||
if ($Time.ToUniversalTime() -lt [DateTime]::Parse([string]$r.accepted_at,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime()) {return 'delivery'}
|
||||
return 'preparation'
|
||||
}
|
||||
do {
|
||||
@@ -46,8 +49,8 @@ do {
|
||||
try {$after=Get-OtcheDefender} catch {Add-CollectionError ('Defender current: '+$_.Exception.Message)}
|
||||
try {
|
||||
$raw=@(Get-MpThreatDetection -ErrorAction Stop | Sort-Object InitialDetectionTime -Descending | Select-Object -First 256)
|
||||
$threats=@(Get-MpThreat -ErrorAction Stop | Select-Object -First 256)
|
||||
if ($raw.Count -eq 256 -or $threats.Count -eq 256) {Add-CollectionError 'Defender query bound reached'}
|
||||
$threats=$null
|
||||
if ($raw.Count -eq 256) {Add-CollectionError 'Defender query bound reached'}
|
||||
foreach ($d in $raw) {
|
||||
$time=$d.InitialDetectionTime
|
||||
if (!$time -or $time.ToUniversalTime() -lt $since) {continue}
|
||||
@@ -55,20 +58,28 @@ do {
|
||||
if (@($d.Resources).Count -gt 4) {Add-CollectionError 'Detection resource bound reached'}
|
||||
if (!($resources | Where-Object {Test-ResourceMatch $_})) {continue}
|
||||
$key='threat:'+([string]$d.DetectionID)
|
||||
if (!$seen.Add($key)) {continue}
|
||||
if ($seen.Contains($key)) {continue}
|
||||
if ($detections.Count -ge 64) {Add-CollectionError 'Detection output bound reached';continue}
|
||||
# Do not query threat metadata when no new correlated detection needs it.
|
||||
if ($null -eq $threats) {$threats=@(Get-MpThreat -ErrorAction Stop | Select-Object -First 256);if ($threats.Count -eq 256) {Add-CollectionError 'Defender query bound reached'}}
|
||||
$threat=$threats | Where-Object {$_.ThreatID -eq $d.ThreatID} | Select-Object -First 1
|
||||
$name=$(if ($threat) {[string]$threat.ThreatName} else {'Threat '+$d.ThreatID})
|
||||
$detections.Add([ordered]@{name=$name; id=[string]$d.DetectionID; action=('cleaning_action={0};success={1};status={2}' -f $d.CleaningActionID,$d.ActionSuccess,$d.ThreatStatusID); resources=$resources; timestamp=$time.ToUniversalTime().ToString('o'); stage=(Get-Stage $time); source='defender'})
|
||||
[void]$seen.Add($key)
|
||||
}
|
||||
} catch {Add-CollectionError ('Defender detections: '+$_.Exception.Message)}
|
||||
foreach ($source in @(@('Microsoft-Windows-Windows Defender/Operational','defender'),@('Microsoft-Windows-AppLocker/EXE and DLL','policy'),@('Microsoft-Windows-AppLocker/MSI and Script','policy'),@('Microsoft-Windows-CodeIntegrity/Operational','policy'))) {
|
||||
try {
|
||||
$queryErrors=@()
|
||||
$batch=@(Get-WinEvent -FilterHashtable @{LogName=$source[0];StartTime=$since} -MaxEvents 200 -ErrorAction SilentlyContinue -ErrorVariable queryErrors)
|
||||
$channel=[string]$source[0]
|
||||
if (!$eventCursors.ContainsKey($channel)) {$eventCursors[$channel]=[long]0}
|
||||
$xpath="*[System[EventRecordID>$($eventCursors[$channel]) and TimeCreated[@SystemTime>='$eventSince']]]"
|
||||
# Oldest first drains a bounded backlog without skipping events beyond the cap.
|
||||
$batch=@(Get-WinEvent -LogName $channel -FilterXPath $xpath -Oldest -MaxEvents 200 -ErrorAction SilentlyContinue -ErrorVariable queryErrors)
|
||||
foreach ($queryError in $queryErrors) {if ($queryError.FullyQualifiedErrorId -notlike 'NoMatchingEventsFound*') {Add-CollectionError ('Event log '+$source[0]+': '+$queryError.Exception.Message)}}
|
||||
if ($batch.Count -eq 200) {Add-CollectionError ('Event log bound reached: '+$source[0])}
|
||||
foreach ($event in $batch) {
|
||||
$eventCursors[$channel]=[Math]::Max([long]$eventCursors[$channel],[long]$event.RecordId)
|
||||
$xml=$event.ToXml()
|
||||
if (!(Test-ResourceMatch $xml)) {continue}
|
||||
$matchedResources=@($candidatePaths | Where-Object {$xml -match [regex]::Escape($_)})
|
||||
@@ -87,7 +98,7 @@ do {
|
||||
$drift=$null -ne $before -and $null -ne $after -and $before.fingerprint -ne $after.fingerprint
|
||||
if ($drift) {Add-CollectionError 'Defender baseline drift during observation'}
|
||||
Write-OtcheJson "$directory\defender-events.json" @($events.ToArray())
|
||||
Write-OtcheJson "$directory\telemetry.json" ([ordered]@{updated_at=(Get-OtcheUtc);boot_id=$boot;before=$before;after=$after;drift=$drift;detections=@($detections.ToArray());environment=$environment;errors=@($errors.ToArray())})
|
||||
Write-OtcheJson "$directory\telemetry.json" ([ordered]@{updated_at=(Get-OtcheUtc);boot_id=$boot;before=$before;after=$after;drift=$drift;detections=@($detections.ToArray());environment=$environment;errors=@($errors.ToArray());initial=$false})
|
||||
if ([IO.File]::Exists("$directory\result.json")) {Export-OtcheGrub $manifest $directory;break}
|
||||
Start-Sleep -Seconds 3
|
||||
} while ($clock.Elapsed.TotalSeconds -lt ([int]$manifest.settings.duration_seconds+240))
|
||||
|
||||
+113
-16
@@ -29,8 +29,12 @@ Qualification uses isolated disposable clones, never the master. See Source-Setu
|
||||
#>
|
||||
#requires -Version 5.1
|
||||
#requires -RunAsAdministrator
|
||||
[CmdletBinding()]
|
||||
param([Parameter(Mandatory=$true)][System.Management.Automation.PSCredential]$Credential,[switch]$EnableAutoLogon)
|
||||
[CmdletBinding(DefaultParameterSetName='Install')]
|
||||
param(
|
||||
[Parameter(Mandatory=$true,ParameterSetName='Install')][System.Management.Automation.PSCredential]$Credential,
|
||||
[Parameter(ParameterSetName='Install')][switch]$EnableAutoLogon,
|
||||
[Parameter(Mandatory=$true,ParameterSetName='Verify')][switch]$Verify,
|
||||
[Parameter(Mandatory=$true,ParameterSetName='Update')][switch]$UpdateRunner)
|
||||
Set-StrictMode -Version 2.0
|
||||
$ErrorActionPreference='Stop'
|
||||
if ($PSVersionTable.PSEdition -ne 'Desktop' -or $PSVersionTable.PSVersion.Major -ne 5) {throw 'Windows PowerShell 5.1 is required'}
|
||||
@@ -38,6 +42,107 @@ if ([Environment]::Is64BitOperatingSystem -and ![Environment]::Is64BitProcess) {
|
||||
if ($ExecutionContext.SessionState.LanguageMode -ne 'FullLanguage') {throw 'Approved FullLanguage policy is required for fixed trusted native helpers'}
|
||||
if ((Get-ExecutionPolicy) -eq 'Restricted') {throw 'Restricted blocks installer and runner files; obtain approved organizational script policy first'}
|
||||
$root='C:\ProgramData\Otche'
|
||||
$files=@('Otche.Common.psm1','Invoke-Otche.ps1','Invoke-OtcheDispatch.ps1','Start-OtcheCommand.ps1','Collect-Otche.ps1','Test-OtcheReady.ps1')
|
||||
function Get-SourceTaskDefinitions([string]$AccountName) {
|
||||
foreach ($level in @('user','admin')) {
|
||||
@{name=('Otche-'+$level);user=$AccountName;logon='Interactive';run_level=$(if ($level -eq 'admin') {'Highest'} else {'Limited'});arguments="-NoProfile -NonInteractive -File `"$root\runner\Invoke-OtcheDispatch.ps1`" -Privilege $level";seconds=1800}
|
||||
}
|
||||
@{name='Otche-DesktopReady';user=$AccountName;logon='Interactive';run_level='Limited';arguments="-NoProfile -NonInteractive -WindowStyle Hidden -File `"$root\runner\Test-OtcheReady.ps1`" -Privilege user -DesktopProbe";seconds=15}
|
||||
@{name='Otche-Collector';user='SYSTEM';logon='ServiceAccount';run_level='Highest';arguments="-NoProfile -NonInteractive -File `"$root\runner\Collect-Otche.ps1`"";seconds=1800}
|
||||
}
|
||||
function Get-SourceSid([string]$Name) {
|
||||
if ($Name -match '^S-1-') {return ([Security.Principal.SecurityIdentifier]::new($Name)).Value}
|
||||
return ([Security.Principal.NTAccount]::new($Name)).Translate([Security.Principal.SecurityIdentifier]).Value
|
||||
}
|
||||
function Test-SourceInstall([switch]$ForUpdate) {
|
||||
$problems=New-Object 'Collections.Generic.List[string]'
|
||||
$source=$null;$prepared=$null
|
||||
try {
|
||||
$sourcePath="$root\source.json"
|
||||
if ((Get-Item -LiteralPath $sourcePath).Attributes -band [IO.FileAttributes]::ReparsePoint) {throw 'source.json is a reparse point'}
|
||||
$stream=[IO.File]::OpenRead($sourcePath)
|
||||
try {
|
||||
if ($stream.Length -gt 1048576) {throw 'source.json exceeds control limit'}
|
||||
$reader=[IO.StreamReader]::new($stream)
|
||||
try {$source=$reader.ReadToEnd()|ConvertFrom-Json} finally {$reader.Dispose()}
|
||||
} finally {$stream.Dispose()}
|
||||
$prepared=Get-LocalUser -SID ([Security.Principal.SecurityIdentifier]::new([string]$source.account_sid)) -ErrorAction Stop
|
||||
if (!$prepared.Enabled -or $prepared.SID.Value.EndsWith('-500') -or (Get-SourceSid ([string]$source.account)) -ne $prepared.SID.Value) {throw 'Prepared local account identity differs or is disabled/RID500'}
|
||||
if (!(@(Get-LocalGroupMember -SID 'S-1-5-32-544' -ErrorAction Stop)|Where-Object {$_.SID.Value -eq $prepared.SID.Value})) {throw 'Prepared account is not a local administrator'}
|
||||
if ($source.autologon -isnot [bool]) {throw 'source.json autologon must be boolean'}
|
||||
} catch {$problems.Add('Source identity: '+$_.Exception.Message);$prepared=$null}
|
||||
if ($null -ne $prepared) {
|
||||
foreach ($dir in @('','runner','control','results','samples')) {
|
||||
try {
|
||||
$path=$(if ($dir) {Join-Path $root $dir} else {$root})
|
||||
$item=Get-Item -LiteralPath $path -ErrorAction Stop
|
||||
if (!$item.PSIsContainer -or ($item.Attributes -band [IO.FileAttributes]::ReparsePoint)) {throw 'Expected an ordinary directory'}
|
||||
$acl=[IO.Directory]::GetAccessControl($path)
|
||||
if ($acl.AreAccessRulesProtected -ne ($dir -eq '')) {throw 'Directory inheritance differs'}
|
||||
$expected=New-Object 'Collections.Generic.List[string]'
|
||||
foreach ($sid in @('S-1-5-18','S-1-5-32-544',$prepared.SID.Value)) {
|
||||
$rights=$(if ($sid -eq $prepared.SID.Value) {'ReadAndExecute'} else {'FullControl'})
|
||||
$rule=[Security.AccessControl.FileSystemAccessRule]::new([Security.Principal.SecurityIdentifier]::new($sid),$rights,'ContainerInherit,ObjectInherit','None','Allow')
|
||||
$expected.Add(('{0}|{1}|3|0|Allow|{2}' -f $sid,[int]$rule.FileSystemRights,($dir -ne '')))
|
||||
}
|
||||
if ($dir -in @('results','samples')) {
|
||||
$rule=[Security.AccessControl.FileSystemAccessRule]::new($prepared.SID,'Modify','ContainerInherit,ObjectInherit','None','Allow')
|
||||
$expected.Add(('{0}|{1}|3|0|Allow|False' -f $prepared.SID.Value,[int]$rule.FileSystemRights))
|
||||
}
|
||||
$actual=@($acl.GetAccessRules($true,$true,[Security.Principal.SecurityIdentifier])|ForEach-Object {'{0}|{1}|{2}|{3}|{4}|{5}' -f $_.IdentityReference.Value,[int]$_.FileSystemRights,[int]$_.InheritanceFlags,[int]$_.PropagationFlags,$_.AccessControlType,$_.IsInherited})
|
||||
if ((($actual|Sort-Object) -join ';') -cne (($expected.ToArray()|Sort-Object) -join ';')) {throw 'Directory access rules differ'}
|
||||
} catch {$problems.Add('ACL '+$dir+': '+$_.Exception.Message)}
|
||||
}
|
||||
foreach ($definition in @(Get-SourceTaskDefinitions ([string]$source.account))) {
|
||||
try {
|
||||
$task=Get-ScheduledTask -TaskPath '\' -TaskName $definition.name -ErrorAction Stop
|
||||
$principal=$task.Principal
|
||||
if ((Get-SourceSid $principal.UserId) -ne (Get-SourceSid $definition.user) -or [string]$principal.LogonType -ne $definition.logon -or [string]$principal.RunLevel -ne $definition.run_level) {throw 'Task principal differs'}
|
||||
if ([string]$task.State -eq 'Disabled' -or ($ForUpdate -and [string]$task.State -ne 'Ready')) {throw 'Task is disabled or busy'}
|
||||
if (@($task.Actions).Count -ne 1 -or $task.Actions[0].Execute -ine "$env:WINDIR\System32\WindowsPowerShell\v1.0\powershell.exe" -or $task.Actions[0].Arguments -cne $definition.arguments -or $task.Actions[0].WorkingDirectory -ine "$root\runner") {throw 'Task action differs'}
|
||||
if (@($task.Triggers|Where-Object {$null -ne $_}).Count -ne 0) {throw 'On-demand task has unexpected triggers'}
|
||||
} catch {$problems.Add('Task '+$definition.name+': '+$_.Exception.Message)}
|
||||
}
|
||||
try {
|
||||
$key=[Microsoft.Win32.Registry]::LocalMachine.OpenSubKey('SOFTWARE\Microsoft\Windows NT\CurrentVersion\Winlogon')
|
||||
try {
|
||||
if (($key.GetValue('AutoAdminLogon','0') -eq '1') -ne [bool]$source.autologon) {throw 'AutoAdminLogon differs from source.json'}
|
||||
if ($source.autologon -and ($key.GetValue('DefaultUserName','') -ine $prepared.Name -or $key.GetValue('DefaultDomainName','') -ine $env:COMPUTERNAME)) {throw 'Winlogon account differs'}
|
||||
if ($null -ne $key.GetValue('DefaultPassword',$null)) {throw 'Plaintext Winlogon password is present'}
|
||||
} finally {if ($null -ne $key) {$key.Dispose()}}
|
||||
} catch {$problems.Add('Autologon: '+$_.Exception.Message)}
|
||||
}
|
||||
if (!$ForUpdate) {
|
||||
foreach ($name in $files) {
|
||||
try {
|
||||
$installed="$root\runner\$name";$supplied=Join-Path $PSScriptRoot $name
|
||||
if ((Get-Item -LiteralPath $installed).Attributes -band [IO.FileAttributes]::ReparsePoint) {throw 'Runner file is a reparse point'}
|
||||
if ((Get-AuthenticodeSignature -LiteralPath $installed).Status -ne 'Valid' -or (Get-AuthenticodeSignature -LiteralPath $supplied).Status -ne 'Valid') {throw 'Valid trusted Authenticode signatures required for AllSigned compatibility'}
|
||||
if ((Get-FileHash -LiteralPath $installed -Algorithm SHA256).Hash -cne (Get-FileHash -LiteralPath $supplied -Algorithm SHA256).Hash) {throw 'Installed hash differs from supplied file'}
|
||||
} catch {$problems.Add('Runner '+$name+': '+$_.Exception.Message)}
|
||||
}
|
||||
}
|
||||
return [ordered]@{ok=($problems.Count -eq 0);errors=@($problems.ToArray())}
|
||||
}
|
||||
if ($Verify -or $UpdateRunner) {
|
||||
$check=Test-SourceInstall -ForUpdate:$UpdateRunner
|
||||
if ($Verify -or !$check.ok) {$check|ConvertTo-Json -Depth 5 -Compress;if (!$check.ok) {exit 1};exit 0}
|
||||
$staged=New-Object 'Collections.Generic.List[string]'
|
||||
try {
|
||||
foreach ($name in $files) {
|
||||
$supplied=Join-Path $PSScriptRoot $name;$destination="$root\runner\$name";$temp=$destination+'.new'
|
||||
if ((Get-Item -LiteralPath $destination).Attributes -band [IO.FileAttributes]::ReparsePoint) {throw "Runner file is a reparse point: $name"}
|
||||
if ((Get-AuthenticodeSignature -LiteralPath $supplied).Status -ne 'Valid') {throw "Update requires valid trusted Authenticode signature: $name"}
|
||||
$hash=(Get-FileHash -LiteralPath $supplied -Algorithm SHA256).Hash
|
||||
[IO.File]::Copy($supplied,$temp,$false);$staged.Add($temp)
|
||||
if ((Get-FileHash -LiteralPath $temp -Algorithm SHA256).Hash -cne $hash -or (Get-AuthenticodeSignature -LiteralPath $supplied).Status -ne 'Valid' -or (Get-FileHash -LiteralPath $supplied -Algorithm SHA256).Hash -cne $hash) {throw "Staged runner changed: $name"}
|
||||
}
|
||||
foreach ($name in $files) {[IO.File]::Replace("$root\runner\$name.new","$root\runner\$name",[System.Management.Automation.Language.NullString]::Value)}
|
||||
} finally {foreach ($path in $staged) {if ([IO.File]::Exists($path)) {[IO.File]::Delete($path)}}}
|
||||
Import-Module "$root\runner\Otche.Common.psm1" -Force
|
||||
Get-OtcheBaseline|ConvertTo-Json -Depth 18 -Compress
|
||||
exit 0
|
||||
}
|
||||
if ([IO.Directory]::Exists($root)) {throw 'Otche source directory already exists; review existing setup rather than overwrite it'}
|
||||
foreach ($name in @('Otche-user','Otche-admin','Otche-Collector','Otche-DesktopReady')) {if (Get-ScheduledTask -TaskName $name -ErrorAction SilentlyContinue) {throw "Existing task $name must be reviewed, not overwritten"}}
|
||||
$user=$Credential.UserName
|
||||
@@ -52,7 +157,6 @@ if (!($admins | Where-Object {$_.SID.Value -eq $account.SID.Value})) {throw 'Exi
|
||||
$uac=[Microsoft.Win32.Registry]::LocalMachine.OpenSubKey('SOFTWARE\Microsoft\Windows\CurrentVersion\Policies\System')
|
||||
try {if ($uac.GetValue('EnableLUA',0) -ne 1) {throw 'UAC must already be enabled; installer does not change it'}} finally {$uac.Dispose()}
|
||||
if ($account.SID.Value.EndsWith('-500')) {throw 'Use an operator-approved existing non-built-in split-token account, not RID500 Administrator'}
|
||||
$files=@('Otche.Common.psm1','Invoke-Otche.ps1','Invoke-OtcheDispatch.ps1','Start-OtcheCommand.ps1','Collect-Otche.ps1','Test-OtcheReady.ps1')
|
||||
foreach ($name in $files) {
|
||||
$path=Join-Path $PSScriptRoot $name
|
||||
if (![IO.File]::Exists($path)) {throw "Required runner file missing: $name"}
|
||||
@@ -100,7 +204,6 @@ try {
|
||||
if ($EnableAutoLogon) {
|
||||
$winlogon=[Microsoft.Win32.Registry]::LocalMachine.OpenSubKey('SOFTWARE\Microsoft\Windows NT\CurrentVersion\Winlogon',$true)
|
||||
if ($winlogon.GetValue('AutoAdminLogon','0') -eq '1' -or $null -ne $winlogon.GetValue('DefaultPassword',$null)) {$winlogon.Dispose();throw 'Existing autologon/plaintext password configuration refused'}
|
||||
try {[OtcheInstall.Secret]::StoreNew($password,$Credential.Password.Length)} catch {$winlogon.Dispose();throw}
|
||||
}
|
||||
[void][IO.Directory]::CreateDirectory($root)
|
||||
# Protected inheritance: only SYSTEM/admins modify trusted control/code; account only reads.
|
||||
@@ -118,23 +221,17 @@ try {
|
||||
Import-Module "$root\runner\Otche.Common.psm1" -Force
|
||||
Write-OtcheJson "$root\source.json" @{account_sid=$account.SID.Value;account="$env:COMPUTERNAME\$user";installed_at=(Get-OtcheUtc);runner_version='1.0.0';autologon=[bool]$EnableAutoLogon} -CreateNew
|
||||
$exe="$env:WINDIR\System32\WindowsPowerShell\v1.0\powershell.exe"
|
||||
$settings=New-ScheduledTaskSettingsSet -MultipleInstances IgnoreNew -ExecutionTimeLimit ([TimeSpan]::FromMinutes(30)) -AllowStartIfOnBatteries -DontStopIfGoingOnBatteries
|
||||
foreach ($level in @('user','admin')) {
|
||||
$principal=New-ScheduledTaskPrincipal -UserId "$env:COMPUTERNAME\$user" -LogonType Interactive -RunLevel $(if ($level -eq 'admin') {'Highest'} else {'Limited'})
|
||||
$action=New-ScheduledTaskAction -Execute $exe -Argument "-NoProfile -NonInteractive -File `"$root\runner\Invoke-OtcheDispatch.ps1`" -Privilege $level" -WorkingDirectory "$root\runner"
|
||||
Register-ScheduledTask -TaskName ('Otche-'+$level) -Action $action -Principal $principal -Settings $settings | Out-Null
|
||||
foreach ($definition in @(Get-SourceTaskDefinitions "$env:COMPUTERNAME\$user")) {
|
||||
$settings=New-ScheduledTaskSettingsSet -MultipleInstances IgnoreNew -ExecutionTimeLimit ([TimeSpan]::FromSeconds($definition.seconds)) -AllowStartIfOnBatteries -DontStopIfGoingOnBatteries
|
||||
$principal=New-ScheduledTaskPrincipal -UserId $definition.user -LogonType $definition.logon -RunLevel $definition.run_level
|
||||
$action=New-ScheduledTaskAction -Execute $exe -Argument $definition.arguments -WorkingDirectory "$root\runner"
|
||||
Register-ScheduledTask -TaskName $definition.name -Action $action -Principal $principal -Settings $settings | Out-Null
|
||||
}
|
||||
$principal=New-ScheduledTaskPrincipal -UserId "$env:COMPUTERNAME\$user" -LogonType Interactive -RunLevel Limited
|
||||
$action=New-ScheduledTaskAction -Execute $exe -Argument "-NoProfile -NonInteractive -WindowStyle Hidden -File `"$root\runner\Test-OtcheReady.ps1`" -Privilege user -DesktopProbe" -WorkingDirectory "$root\runner"
|
||||
$desktopSettings=New-ScheduledTaskSettingsSet -MultipleInstances IgnoreNew -ExecutionTimeLimit ([TimeSpan]::FromSeconds(15)) -AllowStartIfOnBatteries -DontStopIfGoingOnBatteries
|
||||
Register-ScheduledTask -TaskName 'Otche-DesktopReady' -Action $action -Principal $principal -Settings $desktopSettings | Out-Null
|
||||
$principal=New-ScheduledTaskPrincipal -UserId 'SYSTEM' -LogonType ServiceAccount -RunLevel Highest
|
||||
$action=New-ScheduledTaskAction -Execute $exe -Argument "-NoProfile -NonInteractive -File `"$root\runner\Collect-Otche.ps1`"" -WorkingDirectory "$root\runner"
|
||||
Register-ScheduledTask -TaskName 'Otche-Collector' -Action $action -Principal $principal -Settings $settings | Out-Null
|
||||
$baseline=Get-OtcheBaseline
|
||||
Write-OtcheJson "$root\installed-baseline.json" $baseline -CreateNew
|
||||
if ($EnableAutoLogon) {
|
||||
try {
|
||||
[OtcheInstall.Secret]::StoreNew($password,$Credential.Password.Length)
|
||||
$winlogon.SetValue('DefaultUserName',$user,[Microsoft.Win32.RegistryValueKind]::String)
|
||||
$winlogon.SetValue('DefaultDomainName',$env:COMPUTERNAME,[Microsoft.Win32.RegistryValueKind]::String)
|
||||
$winlogon.SetValue('AutoAdminLogon','1',[Microsoft.Win32.RegistryValueKind]::String)
|
||||
|
||||
@@ -22,7 +22,7 @@ function Update-Telemetry {
|
||||
try {
|
||||
$script:telemetry=Read-OtcheJson "$directory\telemetry.json"
|
||||
if ($script:telemetry.boot_id -ne $receipt.boot_id) {throw 'Telemetry boot changed'}
|
||||
$script:lastTelemetry=[DateTime]::Parse($script:telemetry.updated_at).ToUniversalTime()
|
||||
$script:lastTelemetry=[DateTime]::Parse([string]$script:telemetry.updated_at,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime()
|
||||
$report.defender.before=$script:telemetry.before; $report.defender.after=$script:telemetry.after
|
||||
$report.defender.drift=$script:telemetry.drift; $report.defender.detections=@($script:telemetry.detections)
|
||||
$report.environment=$script:telemetry.environment
|
||||
@@ -55,7 +55,7 @@ try {
|
||||
$execution.user=$session.user; $execution.session_id=$session.session_id
|
||||
if ((Get-OtcheBootId) -ne $receipt.boot_id) {throw 'INTERRUPTED: Guest rebooted after command acceptance'}
|
||||
$readyClock=[Diagnostics.Stopwatch]::StartNew()
|
||||
while (![IO.File]::Exists("$directory\telemetry.json") -and $readyClock.Elapsed.TotalSeconds -lt 30) {Start-Sleep -Milliseconds 250}
|
||||
while (![IO.File]::Exists("$directory\telemetry.json") -and $readyClock.Elapsed.TotalSeconds -lt 90) {Start-Sleep -Milliseconds 250}
|
||||
Update-Telemetry
|
||||
if ($null -eq $report.defender.before -or !$report.defender.before.active -or $null -eq $report.environment -or $errors.Count -gt 0) {throw 'Required baseline unavailable/incomplete; source is unqualified'}
|
||||
$volume=[IO.DriveInfo]::new('C:\')
|
||||
@@ -160,7 +160,7 @@ try {
|
||||
$state.phase='collecting';Save-Status
|
||||
if (@(Get-OtcheProperty $manifest.settings 'grub_paths' @()).Count -gt 0) {
|
||||
# This runs on the already verified selected interactive token, not SYSTEM.
|
||||
$grubAfter=$(if ($null -ne $execution.started_at) {[DateTime]::Parse($execution.started_at)} else {[DateTime]::Parse($receipt.accepted_at)}).ToUniversalTime().AddSeconds([int]$manifest.settings.duration_seconds)
|
||||
$grubAfter=[DateTime]::Parse([string]$(if ($null -ne $execution.started_at) {$execution.started_at} else {$receipt.accepted_at}),[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime().AddSeconds([int]$manifest.settings.duration_seconds)
|
||||
while ([DateTime]::UtcNow -lt $grubAfter) {Start-Sleep -Milliseconds 500;Update-Telemetry;Save-Status}
|
||||
$report.grub=Get-OtcheGrubSnapshot $manifest $directory
|
||||
}
|
||||
|
||||
@@ -19,7 +19,7 @@ Import-Module "$PSScriptRoot\Otche.Common.psm1" -Force
|
||||
$marker=Read-OtcheJson 'C:\ProgramData\Otche\control\disposable-qualification.json'
|
||||
$key=[Microsoft.Win32.Registry]::LocalMachine.OpenSubKey('SOFTWARE\Microsoft\Cryptography')
|
||||
try {$machine=[string]$key.GetValue('MachineGuid')} finally {$key.Dispose()}
|
||||
if ($marker.purpose -ne 'qualification' -or $marker.machine_guid -ne $machine -or [DateTime]::Parse($marker.expires_at).ToUniversalTime() -le [DateTime]::UtcNow) {throw 'Missing, mismatched or expired disposable qualification marker'}
|
||||
if ($marker.purpose -ne 'qualification' -or $marker.machine_guid -ne $machine -or [DateTime]::Parse([string]$marker.expires_at,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime() -le [DateTime]::UtcNow) {throw 'Missing, mismatched or expired disposable qualification marker'}
|
||||
$output=[IO.Path]::GetFullPath($OutputDirectory)
|
||||
if ($output.StartsWith('C:\ProgramData\Otche\',[StringComparison]::OrdinalIgnoreCase)) {throw 'Fixtures must not be written into installed source/control/results directories'}
|
||||
if ([IO.Directory]::Exists($output)) {throw 'Use a new fixture directory; existing files are never overwritten'}
|
||||
|
||||
@@ -159,10 +159,27 @@ function Assert-OtcheDesktop([int]$SessionId) {
|
||||
[void][Otche.Native]::EnumWindows($callback,[IntPtr]::Zero)
|
||||
if ($blocked.Count) {throw ('Visible setup/sign-in host prevents desktop readiness: '+($blocked -join ', '))}
|
||||
}
|
||||
function Get-OtcheRegistryValue([string]$Path, [string]$Name) {
|
||||
$key=[Microsoft.Win32.Registry]::LocalMachine.OpenSubKey($Path)
|
||||
if ($null -eq $key) {return $null}
|
||||
try {return $key.GetValue($Name,$null)} finally {$key.Dispose()}
|
||||
}
|
||||
function Get-OtcheEnvironment {
|
||||
$key = [Microsoft.Win32.Registry]::LocalMachine.OpenSubKey('SOFTWARE\Microsoft\Windows NT\CurrentVersion')
|
||||
try { $build = '{0}.{1}.{2}' -f $key.GetValue('CurrentBuildNumber'),$key.GetValue('UBR'),$key.GetValue('BuildLabEx') } finally { $key.Dispose() }
|
||||
return [ordered]@{ os_build=$build; architecture=$(if ([Environment]::Is64BitOperatingSystem) {'x64'} else {'x86'}); powershell_version=$PSVersionTable.PSVersion.ToString(); execution_policy=[string](Get-ExecutionPolicy); runner_version=$script:OtcheVersion }
|
||||
# Registry-backed state is identical for Limited, Highest and SYSTEM tokens.
|
||||
$uac=[ordered]@{}
|
||||
foreach ($name in @('EnableLUA','ConsentPromptBehaviorAdmin','PromptOnSecureDesktop','FilterAdministratorToken')) {$uac[$name]=Get-OtcheRegistryValue 'SOFTWARE\Microsoft\Windows\CurrentVersion\Policies\System' $name}
|
||||
$security=[ordered]@{uac=$uac;smart_app_control=(Get-OtcheRegistryValue 'SYSTEM\CurrentControlSet\Control\CI\Policy' 'VerifiedAndReputablePolicyState');secure_boot=(Get-OtcheRegistryValue 'SYSTEM\CurrentControlSet\Control\SecureBoot\State' 'UEFISecureBootEnabled');device_guard=[ordered]@{EnableVirtualizationBasedSecurity=(Get-OtcheRegistryValue 'SYSTEM\CurrentControlSet\Control\DeviceGuard' 'EnableVirtualizationBasedSecurity');RequirePlatformSecurityFeatures=(Get-OtcheRegistryValue 'SYSTEM\CurrentControlSet\Control\DeviceGuard' 'RequirePlatformSecurityFeatures');HVCIEnabled=(Get-OtcheRegistryValue 'SYSTEM\CurrentControlSet\Control\DeviceGuard\Scenarios\HypervisorEnforcedCodeIntegrity' 'Enabled')}}
|
||||
return [ordered]@{ os_build=$build; architecture=$(if ([Environment]::Is64BitOperatingSystem) {'x64'} else {'x86'}); powershell_version=$PSVersionTable.PSVersion.ToString(); execution_policy=[string](Get-ExecutionPolicy); runner_version=$script:OtcheVersion; security=$security }
|
||||
}
|
||||
function Get-OtcheDeviceGuard {
|
||||
# CIM availability can depend on token rights; never hash this diagnostic snapshot.
|
||||
try {
|
||||
$state=Get-CimInstance -Namespace root\Microsoft\Windows\DeviceGuard -ClassName Win32_DeviceGuard -ErrorAction Stop
|
||||
if ($null -eq $state) {return $null}
|
||||
return [ordered]@{VirtualizationBasedSecurityStatus=(Get-OtcheProperty $state 'VirtualizationBasedSecurityStatus');SecurityServicesConfigured=@(Get-OtcheProperty $state 'SecurityServicesConfigured' @());SecurityServicesRunning=@(Get-OtcheProperty $state 'SecurityServicesRunning' @());CodeIntegrityPolicyEnforcementStatus=(Get-OtcheProperty $state 'CodeIntegrityPolicyEnforcementStatus');UsermodeCodeIntegrityPolicyEnforcementStatus=(Get-OtcheProperty $state 'UsermodeCodeIntegrityPolicyEnforcementStatus')}
|
||||
} catch {return $null}
|
||||
}
|
||||
function Get-OtcheDefender {
|
||||
$status = Get-MpComputerStatus -ErrorAction Stop
|
||||
@@ -181,6 +198,7 @@ function Get-OtcheDefender {
|
||||
function Get-OtcheBaseline {
|
||||
$errors = New-Object 'Collections.Generic.List[string]'; $defender=$null; $environment=$null
|
||||
try { $environment=Get-OtcheEnvironment } catch { $errors.Add($_.Exception.Message) }
|
||||
if ($null -ne $environment -and $environment.security.uac.EnableLUA -ne 1) {$errors.Add('UAC EnableLUA must be 1')}
|
||||
try { $defender=Get-OtcheDefender; if (!$defender.active) { $errors.Add('Defender is not active') } } catch { $errors.Add($_.Exception.Message) }
|
||||
try {
|
||||
$null=Get-MpThreatDetection -ErrorAction Stop
|
||||
@@ -199,7 +217,7 @@ function Get-OtcheBaseline {
|
||||
$policy=@(Get-ExecutionPolicy -List | ForEach-Object { [ordered]@{scope=[string]$_.Scope; policy=[string]$_.ExecutionPolicy} })
|
||||
if ((Get-ExecutionPolicy) -eq 'Restricted') { $errors.Add('Restricted blocks all script files, including signed runners') }
|
||||
$fingerprint=Get-OtcheTextHash ([ordered]@{environment=$environment; defender=$defender; scripts=$scripts; policies=$policy} | ConvertTo-Json -Depth 12 -Compress)
|
||||
return [ordered]@{fingerprint=$fingerprint; qualified=($errors.Count -eq 0); errors=@($errors.ToArray()); defender=$defender; environment=$environment; scripts=$scripts; policies=$policy}
|
||||
return [ordered]@{fingerprint=$fingerprint; qualified=($errors.Count -eq 0); errors=@($errors.ToArray()); defender=$defender; environment=$environment; scripts=$scripts; policies=$policy; device_guard_runtime=(Get-OtcheDeviceGuard)}
|
||||
}
|
||||
function Get-OtcheBootId { return (Get-CimInstance Win32_OperatingSystem -ErrorAction Stop).LastBootUpTime.ToUniversalTime().ToString('o') }
|
||||
function Get-OtchePE([string]$Path) {
|
||||
|
||||
@@ -8,6 +8,13 @@ Prerequisites and trust
|
||||
- UAC must already be enabled. Limited and Highest tasks use the same account's
|
||||
split interactive token. The user task is never elevated; admin is never silently
|
||||
downgraded. Actual account SID, console session and TokenElevation are checked.
|
||||
'user' means the filtered (Limited) token of this prepared local administrator,
|
||||
NOT a separate standard account. An account such as OtcheUser may exist harmlessly
|
||||
but is not used by the runner. True standard-user semantics require a separate
|
||||
source/profile whose console account is a standard user; this installer does not
|
||||
support that model. With the default ConsentPromptBehaviorAdmin=5, eligible Windows
|
||||
components can auto-elevate; Limited is not a guarantee against UAC bypasses. The
|
||||
installer records policy but never changes UAC or other protection settings.
|
||||
- Restricted blocks ALL script files, including signed installer and signed runner.
|
||||
First obtain an independently approved organizational script-execution policy.
|
||||
Recommended: AllSigned, with every supplied .ps1 and .psm1 signed by an approved
|
||||
@@ -26,6 +33,28 @@ Prerequisites and trust
|
||||
overwriting it. LSA secrets remain recoverable by SYSTEM/admins: protect master,
|
||||
clones, exports and backups. Installation failures require operator inspection;
|
||||
installer does not destructively rollback accounts, registry or existing evidence.
|
||||
- Maintenance (run elevated under the existing approved policy; no credential):
|
||||
.\Install-OtcheSource.ps1 -Verify
|
||||
Read-only JSON {ok,errors[]} checks directory ACL shape, local account SID,
|
||||
all four task principals/actions/working directories, on-demand/enabled state,
|
||||
Winlogon autologon/account values, and installed/supplied runner signatures and
|
||||
SHA256 equality. Both copies must have Valid trusted Authenticode signatures for
|
||||
AllSigned compatibility, even if the current policy is RemoteSigned. Without
|
||||
configured autologon, DefaultUserName is not required. This cannot verify the
|
||||
password stored in LSA or replace the real desktop/clone qualification check.
|
||||
.\Install-OtcheSource.ps1 -UpdateRunner
|
||||
Use the reviewed, newly signed installer bundle beside the six runner files on a
|
||||
source explicitly taken out of publication for maintenance. Stops on source/ACL/
|
||||
task/autologon drift or a busy task, before copying anything. Requires Valid
|
||||
trusted signatures for all six replacement files; stages all as .new, checks
|
||||
their hashes, then uses atomic per-file replacement preserving existing ACLs.
|
||||
No task, account, registry, autologon secret or directory ACL is changed. The JSON
|
||||
output is the NEW Get-OtcheBaseline, not qualification. Do not dispatch during
|
||||
maintenance. Replacement is atomic per file, not a six-file transaction: if an
|
||||
I/O failure interrupts replacement, correct the failure and rerun -UpdateRunner,
|
||||
then -Verify. Existing installed-baseline.json is installation evidence, not
|
||||
rewritten by an update. Re-sign changed files, reboot/check the desktop as needed,
|
||||
and requalify disposable clones before republishing the source revision.
|
||||
- No disk is initialized or formatted and no evidence disk is assumed. Artifacts
|
||||
use existing NTFS C:. The full source remains an ordinary stopped VM template=0.
|
||||
Operator reboots and verifies a real active console desktop before stopping it.
|
||||
@@ -53,6 +82,14 @@ Each request has a new nonce and 12-second deadline; bounded response must match
|
||||
nonce, boot ID, configured account and active console session. The task is limited
|
||||
to 15 seconds and runs without opening a console window. UserOOBEBroker residency
|
||||
and arbitrary WWAHost applications are not treated as unfinished setup.
|
||||
Deadline timestamps are parsed with invariant RoundtripKind and normalized to UTC;
|
||||
the guest may use any Windows time zone (UTC is no longer a prerequisite). The clock
|
||||
still must be correct. Probe responses use results\desktop-probe-<nonce>.json; the
|
||||
caller deletes its consumed response and checks at most 128 probe files per request
|
||||
for stale responses older than 60 seconds. Session/task/desktop failures return
|
||||
ready=false with error and baseline=null, without performing a full baseline. Only
|
||||
a successful desktop probe pays for one baseline; -BaselineOnly still collects it
|
||||
without probing sample-owned UI.
|
||||
The normal readiness gate runs before media/dispatch; the actual runner repeats
|
||||
desktop/token checks once before launch. During observation, -BootOnly returns only
|
||||
the boot ID without launching a desktop probe. Post-run -BaselineOnly collects the
|
||||
@@ -90,8 +127,22 @@ Results: C:\ProgramData\Otche\results\<command UUID>\
|
||||
readable active prerequisites, NOT a published qualified revision. Worker compares
|
||||
observed fingerprint with externally qualified revision; drift or unreadable
|
||||
required telemetry is unqualified. Baseline fingerprints include full OS build,
|
||||
runner hashes/signatures, PS policy and Defender versions/preferences. Intelligence
|
||||
updates can change fingerprints; requalify rather than silently accept drift.
|
||||
runner hashes/signatures, PS policy, Defender versions/preferences and security
|
||||
state. environment.security contains uac.{EnableLUA,ConsentPromptBehaviorAdmin,
|
||||
PromptOnSecureDesktop,FilterAdministratorToken}, smart_app_control (CI\Policy
|
||||
VerifiedAndReputablePolicyState), secure_boot (SecureBoot\State
|
||||
UEFISecureBootEnabled), and device_guard.{EnableVirtualizationBasedSecurity,
|
||||
RequirePlatformSecurityFeatures,HVCIEnabled}. Absent registry values are null;
|
||||
EnableLUA other than 1 makes the baseline unqualified. Secure Boot uses the
|
||||
user-readable registry value, never an elevated-only Confirm-SecureBootUEFI call.
|
||||
The Win32_DeviceGuard CIM snapshot is baseline.device_guard_runtime (null if
|
||||
unavailable): VirtualizationBasedSecurityStatus, SecurityServicesConfigured,
|
||||
SecurityServicesRunning, CodeIntegrityPolicyEnforcementStatus and
|
||||
UsermodeCodeIntegrityPolicyEnforcementStatus. This token-dependent diagnostic is
|
||||
NOT fingerprinted; registry DeviceGuard/HVCI configuration is. Thus these new
|
||||
fingerprint fields do not differ solely because a token is Limited vs Highest.
|
||||
Intelligence or security-state updates can change fingerprints; requalify rather
|
||||
than silently accept drift.
|
||||
|
||||
Dispatch and evidence semantics
|
||||
- EXE/SCR/PE COM direct; DOS16/non-PE COM explicitly incompatible. PE x86/x64 chooses
|
||||
@@ -142,6 +193,16 @@ Dispatch and evidence semantics
|
||||
detections and Defender/AppLocker/CodeIntegrity events by time and sample path.
|
||||
Ordinary operational/audit events are not declared blocks. No SmartScreen result
|
||||
is invented from a dialog/exit. Crash/status exit differs from policy evidence.
|
||||
Collector publishes telemetry.json with initial=true immediately after the first
|
||||
successful Defender snapshot (after=before, empty detections), before querying
|
||||
detections/event logs or taking a second snapshot. This removes those queries
|
||||
from the start gate (normally targeting 1-3 seconds, not a timing guarantee).
|
||||
Invoke-Otche permits up to 90 seconds instead of 30 and still fails closed on
|
||||
missing/inactive/error telemetry. Subsequent snapshots set initial=false.
|
||||
Event channels use independent EventRecordID cursors plus the original UTC time
|
||||
filter, oldest-first bounded batches, avoiding repeated XML work. Threat detection
|
||||
queries still run every loop; threat-name metadata is queried only when a new
|
||||
correlated detection needs it. Collection loops sleep three seconds after work.
|
||||
- Detection lists cap 64 (4 resources x 1 KiB), event lists 128, queries 200/channel,
|
||||
Defender queries 256, event XML 4 KiB,
|
||||
control JSON 1 MiB, output JSON 2 MiB, stdout/stderr 1 MiB, exported MSI log 4 MiB.
|
||||
|
||||
@@ -10,26 +10,24 @@ Import-Module "$PSScriptRoot\Otche.Common.psm1" -Force
|
||||
if ($BootOnly) { @{boot_id=(Get-OtcheBootId)}|ConvertTo-Json -Compress;exit 0 }
|
||||
if ($DesktopProbe) {
|
||||
$request=Read-OtcheJson 'C:\ProgramData\Otche\control\desktop-probe.json' 4096
|
||||
if ([string]$request.nonce -notmatch '^[0-9a-f]{32}$' -or [DateTime]::UtcNow -gt [DateTime]::Parse($request.deadline) -or [DateTime]::Parse($request.deadline) -gt [DateTime]::UtcNow.AddSeconds(15)) {throw 'Desktop probe is invalid or expired'}
|
||||
$deadline=[DateTime]::Parse([string]$request.deadline,[Globalization.CultureInfo]::InvariantCulture,[Globalization.DateTimeStyles]::RoundtripKind).ToUniversalTime()
|
||||
if ([string]$request.nonce -notmatch '^[0-9a-f]{32}$' -or [DateTime]::UtcNow -gt $deadline -or $deadline -gt [DateTime]::UtcNow.AddSeconds(15)) {throw 'Desktop probe is invalid or expired'}
|
||||
$probe=[ordered]@{nonce=$request.nonce;boot_id=(Get-OtcheBootId);session_id=0;user='';ready=$false;error=''}
|
||||
try {$session=Get-OtcheSession 'user' -Current;$probe.session_id=$session.session_id;$probe.user=$session.user;$probe.ready=$true} catch {$probe.error=$_.Exception.Message}
|
||||
Write-OtcheJson 'C:\ProgramData\Otche\results\desktop-probe.json' $probe
|
||||
Write-OtcheJson ('C:\ProgramData\Otche\results\desktop-probe-'+$request.nonce+'.json') $probe
|
||||
exit 0
|
||||
}
|
||||
$result=[ordered]@{ready=$false;session_id=$null;user='';privilege=$Privilege;boot_id=$null;baseline=$null;defender=$null;environment=$null;error=''}
|
||||
try {
|
||||
$session=Get-OtcheSession $Privilege
|
||||
$result.session_id=$session.session_id; $result.user=$session.user; $result.boot_id=Get-OtcheBootId
|
||||
$baseline=Get-OtcheBaseline
|
||||
$result.baseline=@{fingerprint=$baseline.fingerprint;qualified=$baseline.qualified;errors=$baseline.errors}
|
||||
$result.defender=$baseline.defender; $result.environment=$baseline.environment
|
||||
$task=Get-ScheduledTask -TaskName ('Otche-'+$Privilege) -ErrorAction Stop
|
||||
$source=Read-OtcheJson 'C:\ProgramData\Otche\source.json'
|
||||
$taskSid=$(if ($task.Principal.UserId -match '^S-1-') {([Security.Principal.SecurityIdentifier]::new($task.Principal.UserId)).Value} else {([Security.Principal.NTAccount]::new($task.Principal.UserId)).Translate([Security.Principal.SecurityIdentifier]).Value})
|
||||
if ($taskSid -ne $source.account_sid -or [string]$task.Principal.LogonType -ne 'Interactive' -or [string]$task.Principal.RunLevel -ne $(if ($Privilege -eq 'admin') {'Highest'} else {'Limited'})) {throw 'Prepared InteractiveToken task principal has drifted'}
|
||||
if ([string]$task.State -eq 'Disabled') {throw 'Prepared runner task is disabled'}
|
||||
$collector=Get-ScheduledTask -TaskName 'Otche-Collector' -ErrorAction Stop
|
||||
if ([string]$collector.Principal.UserId -notin @('SYSTEM','S-1-5-18') -or [string]$collector.State -eq 'Disabled') {throw 'SYSTEM telemetry collector is not prepared'}
|
||||
if ([string]$collector.Principal.UserId -notin @('SYSTEM','S-1-5-18') -or [string]$collector.Principal.LogonType -ne 'ServiceAccount' -or [string]$collector.Principal.RunLevel -ne 'Highest' -or [string]$collector.State -eq 'Disabled') {throw 'SYSTEM telemetry collector is not prepared'}
|
||||
$exe="$env:WINDIR\System32\WindowsPowerShell\v1.0\powershell.exe"
|
||||
$runnerArguments='-NoProfile -NonInteractive -File "C:\ProgramData\Otche\runner\Invoke-OtcheDispatch.ps1" -Privilege '+$Privilege
|
||||
$collectorArguments='-NoProfile -NonInteractive -File "C:\ProgramData\Otche\runner\Collect-Otche.ps1"'
|
||||
@@ -40,18 +38,24 @@ try {
|
||||
$desktopSid=$(if ($desktopTask.Principal.UserId -match '^S-1-') {$desktopTask.Principal.UserId} else {([Security.Principal.NTAccount]::new($desktopTask.Principal.UserId)).Translate([Security.Principal.SecurityIdentifier]).Value})
|
||||
$desktopArguments='-NoProfile -NonInteractive -WindowStyle Hidden -File "C:\ProgramData\Otche\runner\Test-OtcheReady.ps1" -Privilege user -DesktopProbe'
|
||||
if ($desktopSid -ne $source.account_sid -or [string]$desktopTask.Principal.LogonType -ne 'Interactive' -or [string]$desktopTask.Principal.RunLevel -ne 'Limited' -or [string]$desktopTask.State -ne 'Ready' -or @($desktopTask.Actions).Count -ne 1 -or $desktopTask.Actions[0].Execute -ine $exe -or $desktopTask.Actions[0].Arguments -cne $desktopArguments) {throw 'Prepared desktop readiness task has drifted or is busy'}
|
||||
$staleBefore=[DateTime]::UtcNow.AddSeconds(-60)
|
||||
Get-ChildItem -LiteralPath 'C:\ProgramData\Otche\results' -Filter 'desktop-probe-*.json' -File | Select-Object -First 128 | Where-Object {$_.LastWriteTimeUtc -lt $staleBefore -and $_.Name -cmatch '^desktop-probe-[0-9a-f]{32}\.json$'} | ForEach-Object {Remove-Item -LiteralPath $_.FullName -Force -ErrorAction SilentlyContinue}
|
||||
$nonce=[guid]::NewGuid().ToString('N');$deadline=[DateTime]::UtcNow.AddSeconds(12)
|
||||
$probePath='C:\ProgramData\Otche\results\desktop-probe-'+$nonce+'.json'
|
||||
Write-OtcheJson 'C:\ProgramData\Otche\control\desktop-probe.json' @{nonce=$nonce;deadline=$deadline.ToString('o')}
|
||||
Start-ScheduledTask -TaskName 'Otche-DesktopReady'
|
||||
$probe=$null
|
||||
do {
|
||||
Start-Sleep -Milliseconds 200
|
||||
try {$candidate=Read-OtcheJson 'C:\ProgramData\Otche\results\desktop-probe.json' 4096;if ($candidate.nonce -ceq $nonce) {$probe=$candidate;break}} catch {}
|
||||
try {$candidate=Read-OtcheJson $probePath 4096;if ($candidate.nonce -ceq $nonce) {$probe=$candidate;Remove-Item -LiteralPath $probePath -Force;break}} catch {}
|
||||
} while ([DateTime]::UtcNow -lt $deadline)
|
||||
if ($null -eq $probe) {throw 'Fresh interactive desktop probe timed out'}
|
||||
if (!$probe.ready -or [DateTime]::UtcNow -gt $deadline) {throw ('Interactive desktop unavailable or probe expired: '+$probe.error)}
|
||||
if ($probe.boot_id -cne $result.boot_id -or $probe.session_id -ne $session.session_id -or $probe.user -cne $session.user) {throw 'Desktop probe boot/account/session identity mismatch'}
|
||||
}
|
||||
$baseline=Get-OtcheBaseline
|
||||
$result.baseline=@{fingerprint=$baseline.fingerprint;qualified=$baseline.qualified;errors=$baseline.errors;device_guard_runtime=$baseline.device_guard_runtime}
|
||||
$result.defender=$baseline.defender; $result.environment=$baseline.environment
|
||||
$result.ready=$baseline.qualified
|
||||
if (!$result.ready) {$result.error=$baseline.errors -join '; '}
|
||||
} catch {$result.error=$_.Exception.Message}
|
||||
|
||||
Reference in New Issue
Block a user