Otche deployment

Portable deployment of the Otche Windows/Microsoft Defender observation service. The service accepts an immutable file, schedules one Run per qualified Windows profile, records the disposable VM console before delivery, and keeps results and Grub ZIP downloads private to their owner/admin. Observations are not a certificate of file safety.

Architecture and canonical repositories

Clone these repositories as siblings. There is no monorepo, duplicated backend module, submodule, or generated-source export workflow.

Repository Canonical ownership
otche-frontend React/TypeScript UI, npm lockfile, Vite, nginx and web image
otche-backend One shared Go module; API, worker, watchdog and CLI; embedded PostgreSQL schema; unsigned Windows runner; recorder/extractor and trusted operator tools; API/config contracts
otche-deploy Compose, nonsecret environment/TLS/systemd examples and this guide

API, worker and watchdog are separate processes and containers built from the same backend Go module. They share code, not runtime credentials. PostgreSQL is on an internal network with no published port. API has only the database secret and private artifact volume; only worker/watchdog mount operator-supplied PVE configuration. Web has neither. Worker includes FFmpeg/xorriso; Windows runs the signed PowerShell controller. Sample execution requires independently prepared stopped sources and real isolation/source qualification; none is enabled by the default panel-only startup.

Long-running containers are nonroot, read-only, drop capabilities and use no-new-privileges. Named volumes/tmpfs are the only writable application storage. One terminating storage-init container sets artifact-volume ownership with only CHOWN/FOWNER. There is no privileged container, host network, Docker socket mount or production database credential in this repository.

Prerequisites and exact clone/build commands

Use a dedicated Linux deployment host with Docker Engine, Compose v2+, Git, Python 3 and sufficient build/data space. Docker Desktop Linux containers can be used for loopback development; Windows-host bind-mounted secret permissions must be checked explicitly. Linux ownership commands below target a Linux host. Source-only builds need Go 1.26 / Node.js 22, but Docker supplies them. Obtain trusted images and verify dependencies under your organization's policy; no license is invented by this publication.

mkdir otche-workspace
cd otche-workspace
git clone https://git.qomar.pw/otche/otche-backend.git
git clone https://git.qomar.pw/otche/otche-frontend.git
git clone https://git.qomar.pw/otche/otche-deploy.git
cd otche-deploy
cp .env.example .env

For a loopback-only development panel, edit .env to:

PUBLIC_ORIGIN=http://127.0.0.1:8088
ALLOW_INSECURE_HTTP=true
BIND_ADDRESS=127.0.0.1
WEB_PORT=8088

Use a dedicated project name throughout these commands. Compose build contexts are ../otche-backend and ../otche-frontend, relative to this repository, not the invoking shell's arbitrary current directory.

Generate local secrets without printing them

The following creates new secrets only; it refuses existing files. Run in otche-deploy on Linux. Preserve the restrictive directory/file ownership. PostgreSQL alpine uses UID70; backend uses UID10001. A trusted administrator performs this once; private files must never be pasted into Git, .env, command arguments, issue reports or screenshots.

Start a root administration shell in this directory and keep every subsequent secret-generation, Compose, bootstrap and operational command in that same root shell. This is intentional: a Docker-group-only caller cannot traverse the root-owned mode0700 secrets directory. Do not solve that by making secrets world-readable. Container file ownership remains UID70/UID10001.

sudo -s
# Confirm: id -u prints 0; pwd ends in otche-deploy.

When returning for future administration, enter this directory and repeat sudo -s; finish with exit after the work. The source clone/npm developer commands in the other repositories do not require root.

python3 - <<'PY'
import os, pathlib, secrets
os.umask(0o077)
d = pathlib.Path('secrets')
d.mkdir(mode=0o700, exist_ok=False)
pw = secrets.token_hex(32)
values = {
    'postgres-password': (pw, 70),
    'database-url': ('postgres://otche:' + pw + '@postgres:5432/otche?sslmode=disable', 10001),
    'bootstrap-password': (secrets.token_urlsafe(32), 0),
}
for name, (value, uid) in values.items():
    p = d / name
    with p.open('x') as f:
        f.write(value + '\n')
    os.chown(p, uid, uid)
    os.chmod(p, 0o600)
PY

The DSN's sslmode=disable applies only to this isolated Docker database network. External PostgreSQL requires verified TLS. Compose local file secrets are bind mounts: do not rely on ignored Compose uid/mode metadata to fix host ownership.

docker compose -p otche-local config --quiet
docker compose -p otche-local --profile execution build
docker compose -p otche-local up -d --wait

The explicit build covers API and worker image targets plus frontend, but the normal up does not activate the execution profile. migrate applies the embedded idempotent schema with an advisory lock; init/migrate then exit successfully. Worker absence is an expected fail-closed readiness blocker for panel-only use.

Bootstrap the first account via stdin, not a password argument. The protected file is read by the privileged shell; the password is not echoed:

docker compose -p otche-local run --rm -T api user-create --username administrator --role admin --password-stdin < secrets/bootstrap-password

Retrieve the generated password through your approved protected secret-manager workflow to log in at the exact PUBLIC_ORIGIN; never expose it in a shared terminal. Replace/remove the temporary bootstrap copy after secure handoff under your retention policy. Passwords must be 14–72 bytes; there is no default account/password. The UI can create further admin/operator accounts. Missing profiles/worker/proofs are genuine empty or blocked states, not demonstration data.

Production HTTPS

Production requires an exact https:// origin and ALLOW_INSECURE_HTTP=false. examples/nginx-tls.conf and examples/compose.tls.yaml demonstrate direct TLS termination in the existing nonroot web container; substitute your actual DNS hostname, approved certificate chain and private key. The files themselves are not shipped. Mount only server leaf chain/key into web, never a CA private key. nginx needs to read them as UID101; protect their directory and retain appropriate file ACLs/ownership. Keep the bind loopback if using a separate trusted proxy, or explicitly bind your approved interface for direct TLS. Do not expose API/PostgreSQL ports or use wildcard certificate-verification bypasses.

Set .env COMPOSE_FILE=compose.yaml:examples/compose.tls.yaml, PUBLIC_ORIGIN=https://your-approved-hostname, WEB_PORT=443, BIND_ADDRESS=<approved-interface-address>. Edit the example server_name and provide secrets/tls/panel-chain.pem and secrets/tls/panel-key.pem. Relative paths in the override are relative to the first Compose file (repository root). Verify trusted chain, SAN, validity and private-key match independently; install trust by explicit operator policy only. docker compose ... exec -T web nginx -t checks syntax, not certificate trust.

examples/otche.service is the existing-deployment production example: /opt/otche-workspace/otche-deploy with sibling clones, prebuilt images, explicit project otche, and external environment/override/pin files under /srv/otche/deployment. Adapt the project and paths together to your inspected deployment. After the first successful pinned activation below, review/copy it into systemd only on the intended deployment host. Its stop action stops the app, not the host and not volumes. Do not use it on a PVE node merely because that node exists.

Enable Windows execution deliberately

Follow backend CONFIG.md and Source-Setup.txt. Start with ../otche-backend/provisioning/runtime-config.example.json, replacing example IDs/paths with reviewed resources in protected secrets/worker/config.json. This configuration is mounted as /run/otche/config.json; private proof/token/CA paths must match container paths. Seals live on writable /var/lib/otche/source-seals/, not the read-only config mount.

Supply separately scoped provisioner/runtime/recorder/uploader/housekeeping credentials; runtime mounts are UID10001-readable and operator-protected. Authenticate all five through bindings-sync; establish genuine owner-isolation evidence and its expiry. Source preparation requires reviewed signed PS1/PSM1, approved execution policy and signing trust, a dedicated local split-token account, real interactive desktop/autologon, active Defender/UAC/Secure Boot and QGA. Never distribute signing private keys or passwords here. Do not run samples or EICAR on source VMs. Existing protected VMIDs 7000/7001 and shared-storage exclusions are deliberate code safety restrictions.

docker compose -p otche-local --profile execution run --rm worker bindings-sync
docker compose -p otche-local --profile execution run --rm worker source-maintenance --source-ref win11-stopped-revision
# Operator performs approved maintenance and graceful shutdown separately.
docker compose -p otche-local --profile execution run --rm worker source-validate --profile-id ACTUAL_PROFILE_UUID --source-ref win11-stopped-revision
docker compose -p otche-local --profile execution up -d worker watchdog

Admin queues qualification, reviews actual disposable EICAR/benign/video evidence and publishes the exact passed revision. Optional online and read-only extractor support require their own real matching proofs. Do not manufacture passed:true, extend proof expiry by editing dates, remove safety gates, enable broad host privileges or replace private storage with shared aliases. Empty installs intentionally cannot execute files.

Strict per-job network allowlists

Online execution uses the backend's operator-installed provisioning/otche-network.py on the approved PVE node, not a privileged API/worker container and not the cluster-wide firewall switch. Use a real canonical backend checkout at a reviewed commit, the adjacent executable otche-network-dhcp helper, and the supplied otche-network.service. Stage an operator-owned mode0600 /etc/otche-network.json from network-gateway.example.json, with a dedicated server certificate/key, client CA, exact owner/pool/client-CN mapping, existing uplink bridge, private sandbox pool, explicit site networks and protected addresses (including public NAT/hairpin addresses). Do not reuse panel/PVE private keys or mount the broker key into the panel. Worker/watchdog receive only their own mTLS client key, server CA and real probe proof in the existing private mount.

Requirements: Python3, iproute2, nftables, BusyBox with udhcpc, dnsmasq and PVE tools. The dedicated namespace obtains a WAN lease through a fixed BusyBox DHCP hook; host routes, DNS and AppArmor policy remain unchanged. Existing bridge-netfilter legacy switches must be off; the broker refuses incompatible state instead of weakening host filtering. No shared source NIC, direct LAN attachment, IPv6 fallback or DNS relay is permitted. Each disposable clone receives an isolated bridge, exact DHCP assignment and a kernel-enforced 45-second lease. /0 authorizes public IPv4 only; local CIDRs remain separate opt-ins and protected infrastructure always wins.

Check python3 provisioning/otche-network.py --config /etc/otche-network.json --check-config before starting the service. --print-rules namespace and --print-rules bridge support syntax inspection in a disposable namespace; they do not prove runtime isolation. Verify server/client certificate identities, actual allowed/denied destinations, DNS, IPv6, source spoofing, cross-clone isolation, lease expiry and restart failure behavior before creating any online proof or enabling profile admission. Preserve evidence privately. A partially initialized namespace or changed configuration is deliberately refused: stop the service, inspect its recorded inode/ownership markers and exact resources, and reconcile only verified owned empty state. Never fake initialization success or flush unrelated host rules.

The application deployment still follows the pinned updater below. Activate worker online configuration only after real network proof, bindings-sync and paired offline/online EICAR/benign qualification. A healthy broker alone is not authorization to execute. Stop admission and drain active attempts before updating broker code or configuration; restart revokes existing leases. Losing the worker, broker or renewal authorization must close traffic, including established connections.

Automatic proof revalidation on the PVE host

revalidate.py and examples/otche-revalidation.{service,timer} renew both owner ACL/storage/QGA isolation and online network proofs from new disposable resources. They do not requalify Windows revisions, change source VMs, edit old expiry dates, or deploy application images. Use reviewed clean backend/deploy checkouts on the host and pin their full commits in a root-owned mode0600 /etc/otche-revalidation.json, based on examples/revalidation.example.json. Adapt the existing control CT, Compose project/container/network names, worker-config mount, owner and stopped source reference; example identities are not live inventory. The five owner credentials and network client certificate are read from the existing private control-CT mount and staged only in a private run directory. No credentials belong in the scheduler config or Git.

The persistent timer checks every fifteen minutes. The first run and an unconfirmed proof publication require a full cycle; thereafter work starts when either proof has eight hours or less remaining. Successful actual probes issue at most 24 hours of validity; busy deployments defer to the next tick rather than interrupt a job. PostgreSQL SHARE locks cover the idle observation and freezing the exact worker container: unfinished attempts, queued/running qualifications, retained allocations/extractors/media all prevent testing. The API stays available; submissions during the check wait in the queue. The idle worker heartbeat is temporarily stale. ExecStopPost thaws only the exact recorded container, including after a timeout. Do not run updates, source maintenance, restores or another operator probe concurrently.

The cycle full-clones the stopped source, verifies its configuration is unchanged, runs fresh owner allow/deny probes, exercises isolated real kernel packet tests plus live Windows networking through the actual broker, and verifies expiry/restart revocation. A temporary DHCP macvlan target supplies the LAN positive control; it never adds host-root addresses or forwards traffic. Configure reachable public TCP/DNS positive controls and explicit protected PVE/panel/router targets. Any missing positive control fails the check rather than inventing a denial result. Only after verified cleanup are both evidence hashes validated, new proof files atomically replaced and bindings-sync run in the existing approved worker image. Application repositories/images, accounts, data volumes, source seals and qualifications are untouched.

Enable the timer only after a successful supervised python3 revalidate.py --config /etc/otche-revalidation.json --force (force bypasses the time window, not the idle gate). Install the supplied units into /etc/systemd/system, run systemctl daemon-reload, then systemctl enable --now otche-revalidation.timer. Inspect systemctl list-timers otche-revalidation.timer, journalctl -u otche-revalidation.service and /var/lib/otche-revalidation/last-success.json; verify the API's actual binding expiries/readiness. Keep run directories and content-addressed evidence private and include them in operator retention/backups.

Failure never extends old proof validity. Ambiguous/incomplete resource cleanup leaves incomplete.json and stops subsequent cycles for operator inspection; use the exact run journals, never broad VM/storage deletion. Existing proof expiry continues to close admission normally. The scheduler does not heal changed source pins, unsafe broker state or expired credentials. --recover-worker only reverses its own recorded freeze; it does not reset evidence, repair networking or bypass safety gates.

Parallel execution sizing

The worker configuration's top-level concurrent is a global cap; each binding's max_active_attempts defaults to 1 and must be between 1 and 16 and no greater than concurrent. To allow one owner's nine-profile Job to run concurrently, set both caps to at least 9 only after checking resource and isolation capacity. Runs retain the manually authored profile-name snapshot; parallel scheduling does not change it.

Windows resources and worker-container resources are separate budgets. Nine 8 GiB / 4-vCPU guests reserve 72 GiB RAM and 36 virtual CPUs; vCPU overcommit is not guaranteed throughput. Leave headroom for the hypervisor, source VMs and other workloads. The worker runs an FFmpeg encoder per active attempt: .env / the protected external environment controls WORKER_CPUS (default 2.0), WORKER_MEM_LIMIT (default 1536m) and WORKER_PIDS_LIMIT (default 128). An unvalidated starting point for a nine-run load test, not a capacity guarantee, is WORKER_CPUS=8.0, WORKER_MEM_LIMIT=3072m, WORKER_PIDS_LIMIT=512. Measure throttling, peak memory/PIDs, recording completeness and host load before production use. Raising caps alone does not supply resources.

The binding's max_owned_disk_bytes must cover concurrent clones' full per-source max_disk_bytes reservations plus retained evidence and other owned allocations. Nine sources reserving 80 GiB each require at least 720 GiB of quota before evidence headroom, and adequate real storage capacity independently of thin provisioning. Held/failed-cleanup evidence continues to consume capacity until safely released. Do not weaken ownership, qualification, lease or evidence-retention checks to fit a quota.

Operation, backup and updates

Use docker compose -p otche-local ps -a and service logs for diagnosis (logs may contain private job metadata; do not publish them). Authenticate and inspect /api/v1/admin/health; panel-only reports database ready and an unavailable worker, not full execution readiness. Login/session endpoints are described by the backend API contract. Browser assets and /api are same origin; login requires matching Origin.

Every service uses Docker's local log driver with max-size=10m, max-file=5; logs still contain private metadata and require host-level access control. The API runs otche healthcheck every 10 seconds (3-second probe timeout, six retries, 10-second start period); it checks the loopback health endpoint without opening a CLI database pool. Web waits for API health and probes HTTP port 8080, falling back to HTTPS port 8443 for the supplied TLS overrides. Certificate verification is bypassed only for that loopback HTTPS liveness probe, never for clients or PVE. compose up --wait now waits for API/web health, but that does not prove worker heartbeats, qualifications or execution readiness: verify authenticated admin health separately. Runtime Docker DNS resolution keeps /api/ routing intact across API recreation; gzip and immutable hashed-asset caching apply to web responses without duplicating the API's security headers.

Back up PostgreSQL, artifacts/source seals and protected config/keys as a consistent private set using the procedure below. Drain admission/execution before updates; update.py fetches and builds reviewed commits before advancing checkouts under its API-stop fence. After restoring, reconcile allocations and source/proof identity before admitting execution. Never use down -v, volume pruning or broad VM deletion on an existing deployment.

No uploaded files, video, ZIP, database dump, live inventory, operational IP addresses, certificates, runtime logs or acceptance proofs are published. Keep those in ignored protected locations, not alongside tracked sources.

Consistent backup (Linux root Bash shell)

Coordinate a maintenance window, pause submissions and let execution drain. Do not run this concurrently with updates, source maintenance, administrator writes or backup restores. Adapt the inspected existing project/paths below; no example value is a production host or credential. The subshell restores the same original containers on refusal/failure, never newly built images. It stops worker/watchdog only after the database is quiet so background cleanup cannot change artifacts during the dump. Prepare approved helper images before starting the fence.

(
  set -euo pipefail
  umask 077
  cd /opt/otche-workspace/otche-deploy
  PROJECT=EXISTING_PROJECT
  dc() { docker compose --env-file /srv/otche/deployment/deploy.env -p "$PROJECT" \
    -f compose.yaml -f /srv/otche/deployment/compose.production.yaml --profile execution "$@"; }
  BACKUP="/srv/otche-backups/$(date -u +%Y%m%dT%H%M%SZ)"
  mkdir -p "$BACKUP"
  api=$(dc ps -q api)
  test -n "$api"
  ARTIFACT_VOLUME=$(docker inspect --format '{{range .Mounts}}{{if eq .Destination "/var/lib/otche"}}{{.Name}}{{end}}{{end}}' "$api")
  test -n "$ARTIFACT_VOLUME"
  resume="$api"
  trap 'for id in $resume; do docker start "$id"; done' EXIT
  docker stop --time 30 "$api"
  active=$(dc exec -T postgres psql -U otche -d otche -At -v ON_ERROR_STOP=1 -c \
    "SELECT (SELECT count(*) FROM jobs WHERE status IN ('queued','running')) + (SELECT count(*) FROM qualifications WHERE status IN ('queued','running'));")
  test "$active" = 0  # Otherwise the trap restarts the SAME API and this backup is refused.
  executors=$(dc ps -q worker watchdog)
  resume="$executors $api"
  for id in $executors; do docker stop --time 60 "$id"; done
  dc exec -T postgres pg_dump -U otche -d otche -Fc > "$BACKUP/database.dump"
  docker run --rm --network none --read-only --cap-drop ALL --security-opt no-new-privileges \
    --user 10001:10001 --mount "type=volume,src=$ARTIFACT_VOLUME,dst=/artifacts,readonly" \
    debian:bookworm-slim tar -C /artifacts -cpf - . > "$BACKUP/artifacts.tar"
  # Include the protected config, proof files, credentials, TLS and approved commit pins.
  tar -C /srv/otche -cpf "$BACKUP/private-config.tar" deployment secrets tls
  dc ps -a -q | xargs docker inspect --format '{{.Name}} {{.Image}}' > "$BACKUP/container-images.txt"
  (cd "$BACKUP"; sha256sum database.dump artifacts.tar private-config.tar container-images.txt > SHA256SUMS)
  touch "$BACKUP/COMPLETE"
)

A failed run may leave partial files: accept a backup only with COMPLETE, matching hashes and a successful restore drill. The dump and tar contain private data and potentially malicious samples; encrypt and transfer through the approved backup channel, apply retention, and keep decryption keys separately. Protect signing keys using the signing system's own backup procedure; never copy them into runtime containers. Record source-VM backups/seals and operator proof dependencies too: the artifact/database backup alone does not back up PVE guests. Do not restore old proof expiry dates as if they were current attestations.

Restore drill (new isolated volumes only)

Use a separate Linux Docker host with no route or credentials to production PVE, the approved source/images corresponding to the backup pins, and a root-only working directory. Never point the drill at existing production volumes, do not start worker/watchdog, and do not publish database ports. Verify archive hashes before extracting anything. The following restores to newly created, separately named volumes; volume inspect guards against accidentally reusing an earlier drill. Use the same PostgreSQL major version as the dump (17 here).

set -euo pipefail
umask 077
BACKUP=/srv/otche-backups/REVIEWED_BACKUP_DIRECTORY
test -f "$BACKUP/COMPLETE"
(cd "$BACKUP"; sha256sum -c SHA256SUMS)
# Place a copy of the original PostgreSQL password in this protected directory,
# readable only by UID70, using the approved secret-manager workflow; never print it.
test -f /srv/otche-restore/postgres-password
for volume in otche-drill-db otche-drill-artifacts; do
  if docker volume inspect "$volume" >/dev/null 2>&1; then
    echo "Refusing existing drill volume: $volume" >&2; exit 1
  fi
  docker volume create "$volume"
done
docker network create --internal otche-drill-private
docker run -d --name otche-drill-db --network otche-drill-private \
  --user 70:70 --read-only --cap-drop ALL --security-opt no-new-privileges \
  --tmpfs /tmp --tmpfs /var/run/postgresql \
  --mount type=volume,src=otche-drill-db,dst=/var/lib/postgresql/data \
  --mount type=bind,src=/srv/otche-restore/postgres-password,dst=/run/secrets/postgres-password,readonly \
  -e POSTGRES_DB=otche -e POSTGRES_USER=otche \
  -e POSTGRES_PASSWORD_FILE=/run/secrets/postgres-password postgres:17-alpine
# Wait for readiness (bounded); any failure below stops the drill.
for i in {1..30}; do
  if docker exec otche-drill-db pg_isready -U otche -d otche; then break; fi
  sleep 2
done
docker exec otche-drill-db pg_isready -U otche -d otche
docker exec -i otche-drill-db pg_restore -U otche -d otche --exit-on-error --single-transaction < "$BACKUP/database.dump"
docker run --rm -i --network none --read-only --cap-drop ALL --cap-add CHOWN --cap-add FOWNER \
  --cap-add DAC_OVERRIDE --security-opt no-new-privileges --user 0:0 \
  --mount type=volume,src=otche-drill-artifacts,dst=/artifacts \
  debian:bookworm-slim tar -C /artifacts -xpf - < "$BACKUP/artifacts.tar"
docker exec otche-drill-db psql -U otche -d otche -v ON_ERROR_STOP=1 -c \
  "SELECT count(*) AS jobs FROM jobs; SELECT count(*) AS artifacts FROM artifacts; SELECT count(*) AS held_allocations FROM allocations WHERE state <> 'deleted';"
docker run --rm -i --network none --read-only --cap-drop ALL --security-opt no-new-privileges \
  --user 10001:10001 --mount type=volume,src=otche-drill-artifacts,dst=/artifacts,readonly \
  debian:bookworm-slim tar -C /artifacts --compare -f - < "$BACKUP/artifacts.tar"

Check expected owner/account/job counts against the backup record and verify selected artifact sizes/hashes against restored database metadata without opening/executing samples. For an application-level drill, point the pinned API image at the isolated database and restored artifact volume, use a private loopback-only panel, and verify authentication plus an authorized artifact download; never mount worker credentials or enable execution. Keep the restored configuration archive sealed unless needed. Record the tested backup ID, image IDs and results privately. Remove only these exact drill containers/network/volumes after explicitly confirming they are disposable, not production resources.

For a real restore, retain the API-stop fence and a pre-restore backup, restore into new replacement volumes rather than overwriting evidence, then deliberately update the external volume mapping. Verify data and image/pin continuity first; reconcile every allocation with actual owned VMs, source seals and fresh isolation/qualification proofs before re-enabling worker/watchdog and admitting jobs. A database restore is not authority to delete retained VMs or reuse stale leases.

Updating an existing deployment from Git

All application deployments must use actual canonical Git clones, not tar/ZIP copies of an operational workspace. Review source, test the affected behavior, commit and push to the owning repository, then select three full approved commit IDs. API/worker/watchdog share the selected backend commit. Never automatically deploy an unreviewed branch tip.

Keep the production environment, TLS proxy config and Compose override outside the Git workspace, for example under /srv/otche/deployment/ (administrator-owned directory mode0700). Keep secrets in /srv/otche/secrets/ and TLS keys in the existing protected location. examples/compose.production.yaml demonstrates absolute private mounts and explicit external existing volumes; replace its example project/volume names with the verified existing deployment values. External volumes deliberately fail if missing instead of creating an empty replacement database. The production environment can set COMPOSE_PROFILES=execution; that does not bypass worker qualification gates.

For a deployment with sibling clones under /opt/otche-workspace, run in the root administration shell:

cd /opt/otche-workspace/otche-deploy
git fetch origin main
# Do not merge any repository before the updater's quiescence fence.
python3 update.py --backend APPROVED_BACKEND_COMMIT --frontend APPROVED_FRONTEND_COMMIT --deploy APPROVED_DEPLOY_COMMIT --env-file /srv/otche/deployment/deploy.env --override /srv/otche/deployment/compose.production.yaml --project EXISTING_PROJECT
# After coordinating paused new submissions and zero active jobs/qualifications:
python3 update.py --backend APPROVED_BACKEND_COMMIT --frontend APPROVED_FRONTEND_COMMIT --deploy APPROVED_DEPLOY_COMMIT --env-file /srv/otche/deployment/deploy.env --override /srv/otche/deployment/compose.production.yaml --project EXISTING_PROJECT --activate --confirm-quiescent

Use literal reviewed 40-character commits, not the uppercase explanatory placeholders. The script refuses dirty/untracked work, detached/private branches, wrong origins, non-fast-forward/divergent targets and commits not reachable from public main. The update sequence is:

  1. Validate all three repositories and the exact project/external-volume identities. Fetch only outside --verify-only; verify-only requires HEAD to equal each approved pin, without network access or builds.
  2. Create detached sibling worktrees in /var/tmp at the three targets, resolve the target Compose configuration with the same protected external environment/override, and recheck it against the running data volumes. A generated temporary override points every build context explicitly at the matching worktree and assigns temporary image tags. Build-only mode never moves the real checkouts, retags runtime images, restarts services or changes approved pins. Temporary worktrees are removed/pruned in finally; a killed process may leave only out-of-workspace staging directories/worktree metadata, never untracked deployment files. Inspect those paths before manually removing abandoned staging data.
  3. With --activate --confirm-quiescent, capture old container image IDs and gracefully stop the existing API container to fence new admissions. Query PostgreSQL for queued/running jobs and qualifications. Active work or a failed query restarts the same old API container and refuses activation: old checkouts, runtime image tags and pins remain intact, so startup verification still passes. Existing workers and terminal held evidence are untouched on refusal.
  4. Only after zero active work, recheck checkout cleanliness/ancestry, fast-forward the three reviewed commits, and promote the captured built image IDs to the exact runtime tags resolved by target Compose (normally PROJECT-api, PROJECT-migrate, PROJECT-worker, PROJECT-watchdog, PROJECT-web, all :latest). No rebuild is needed: the temporary worktrees used those exact commits. Start with --profile execution up --no-build --wait, including worker/watchdog; migration must complete before the dependent API/executors start. This deliberately activates execution containers, though all qualification/isolation gates still apply.
  5. Explicitly recreate web and wait for health again. Compare each built service container's actual image ID with the captured build ID, require healthy/running services and a successful exited migration, then atomically write approved-commits.env. Missing/mismatched images or unhealthy services never approve pins. A failure after the fence may leave a partial rollout and advanced checkouts; the script prints the old container image IDs for manual recovery and never deletes old images or resets source. Retain the API admission fence until the operator has inspected migration compatibility and established a consistent recovery state; do not blindly downgrade a migrated database.

Temporary build tags include a digest of all three selected commits, not only the deployment revision. Activation force-recreates every built service so Compose cannot silently retain an older image after a retag. For recovery after a recorded partial rollout, the existing API may already be intentionally stopped: its artifact-volume identity is still checked and it remains stopped if validation or quiescence fails. A previously running API is restored on pre-activation refusal; an already fenced API is never reopened automatically. Inspect all checkout/container/image identities before selecting the exact recovery commits; do not edit approved pins to hide a mismatch.

Coordinate the maintenance window and prohibit concurrent deployments/config edits. If upgrading from the old updater, do not first merge its deploy commit: after fetching and reviewing it, execute the reviewed script directly from Git while staying in the deploy checkout, e.g. in Bash with set -o pipefail: git show APPROVED_DEPLOY_COMMIT:update.py | python3 - --backend APPROVED_BACKEND_COMMIT --frontend APPROVED_FRONTEND_COMMIT --deploy APPROVED_DEPLOY_COMMIT --env-file /srv/otche/deployment/deploy.env --override /srv/otche/deployment/compose.production.yaml --project EXISTING_PROJECT (append the two activation flags for the second invocation). This uses the current directory to locate sibling checkouts without advancing them.

Keep systemd's WorkingDirectory and explicit --env-file, -p, -f arguments aligned with these same clones/private files. After activation compare actual HEADs with approved public commits, verify authenticated health and prior accounts/uploads/jobs/artifacts/seals/held evidence, and perform real acceptance before reopening submissions. Logs and the exact host-specific rollout record remain private; do not put credentials or live inventory into commit messages or this public guide.

Successful activation atomically writes approved-commits.env beside the external environment file (mode0600), containing only the three approved nonsecret commit IDs. The systemd example loads these pins and runs update.py --verify-only before startup: exact clean checkouts, expected origins/main ancestry and existing external volumes are checked without network access or builds. It refuses drift instead of silently starting different code. On a failed/partial update inspect the retained containers/checkouts/pins before recovery; never edit the pin file to disguise unknown source. Startup itself never pulls newer code.

S
Description
Portable deployment for the separate Otche backend and frontend repositories
Readme
196 KiB
Languages
Python 100%