Build in detached worktrees and advance pins only after the fence

- update.py builds target commits in temporary detached worktrees with
  temporary image tags; checkouts and runtime tags move only after the API
  fence and quiescence check pass, so refused/build-only runs keep
  --verify-only passing.
- Activation includes the execution profile (worker/watchdog restart under
  the fence) and verifies every service container runs the built image
  before writing approved-commits.env.
- Compose: bounded local logging, API healthcheck via `otche healthcheck`,
  web healthcheck and healthy API dependency, env-driven worker limits.
- TLS example: request-time upstream resolution, gzip, immutable assets,
  headers not duplicated on /api.
- README: new update flow, sizing for parallel runs, fenced backup and
  restore drill.
This commit is contained in:
omar
2026-09-24 17:55:13 +03:00
parent 55a36c7607
commit ffc462cf28
5 changed files with 283 additions and 51 deletions
+4
View File
@@ -8,3 +8,7 @@ WEB_PORT=8088
MAX_UPLOAD_BYTES=536870912
MAX_OWNER_BYTES=8589934592
MAX_QUEUED_JOBS=20
# Worker/FFmpeg container limits; size separately from Windows guest resources.
WORKER_CPUS=2.0
WORKER_MEM_LIMIT=1536m
WORKER_PIDS_LIMIT=128
+111 -3
View File
@@ -117,14 +117,114 @@ docker compose -p otche-local --profile execution up -d worker watchdog
Admin queues qualification, reviews actual disposable EICAR/benign/video evidence and publishes the exact passed revision. Optional online and read-only extractor support require their own real matching proofs. Do not manufacture `passed:true`, extend proof expiry by editing dates, remove safety gates, enable broad host privileges or replace private storage with shared aliases. Empty installs intentionally cannot execute files.
### Parallel execution sizing
The worker configuration's top-level `concurrent` is a global cap; each binding's `max_active_attempts` defaults to 1 and must be between 1 and 16 and no greater than `concurrent`. To allow one owner's nine-profile Job to run concurrently, set both caps to at least 9 **only after** checking resource and isolation capacity. Runs retain the manually authored profile-name snapshot; parallel scheduling does not change it.
Windows resources and worker-container resources are separate budgets. Nine 8 GiB / 4-vCPU guests reserve 72 GiB RAM and 36 virtual CPUs; vCPU overcommit is not guaranteed throughput. Leave headroom for the hypervisor, source VMs and other workloads. The worker runs an FFmpeg encoder per active attempt: `.env` / the protected external environment controls `WORKER_CPUS` (default `2.0`), `WORKER_MEM_LIMIT` (default `1536m`) and `WORKER_PIDS_LIMIT` (default `128`). An **unvalidated starting point for a nine-run load test**, not a capacity guarantee, is `WORKER_CPUS=8.0`, `WORKER_MEM_LIMIT=3072m`, `WORKER_PIDS_LIMIT=512`. Measure throttling, peak memory/PIDs, recording completeness and host load before production use. Raising caps alone does not supply resources.
The binding's `max_owned_disk_bytes` must cover concurrent clones' full per-source `max_disk_bytes` reservations **plus retained evidence and other owned allocations**. Nine sources reserving 80 GiB each require at least 720 GiB of quota before evidence headroom, and adequate real storage capacity independently of thin provisioning. Held/failed-cleanup evidence continues to consume capacity until safely released. Do not weaken ownership, qualification, lease or evidence-retention checks to fit a quota.
## Operation, backup and updates
Use `docker compose -p otche-local ps -a` and service logs for diagnosis (logs may contain private job metadata; do not publish them). Authenticate and inspect `/api/v1/admin/health`; panel-only reports database ready and an unavailable worker, not full execution readiness. Login/session endpoints are described by the backend API contract. Browser assets and `/api` are same origin; login requires matching Origin.
Back up PostgreSQL, artifact volume, source seals/config, keys and credentials as a consistent **private** set. Drain admission/execution before coordinated updates; pull intended reviewed commits in each sibling clone, build, then recreate the application without removing data. After restoring, reconcile allocations and source/proof identity before admitting execution. Never use `down -v`, volume pruning or broad VM deletion on an existing deployment. For a newly created disposable smoke project only, `docker compose -p YOUR_DISPOSABLE_PROJECT down -v --remove-orphans` removes that project's containers/volumes after checking its name.
Every service uses Docker's `local` log driver with `max-size=10m`, `max-file=5`; logs still contain private metadata and require host-level access control. The API runs `otche healthcheck` every 10 seconds (3-second probe timeout, six retries, 10-second start period); it checks the loopback health endpoint without opening a CLI database pool. Web waits for API health and probes HTTP port 8080, falling back to HTTPS port 8443 for the supplied TLS overrides. Certificate verification is bypassed **only for that loopback HTTPS liveness probe**, never for clients or PVE. `compose up --wait` now waits for API/web health, but that does not prove worker heartbeats, qualifications or execution readiness: verify authenticated admin health separately. Runtime Docker DNS resolution keeps `/api/` routing intact across API recreation; gzip and immutable hashed-asset caching apply to web responses without duplicating the API's security headers.
Back up PostgreSQL, artifacts/source seals and protected config/keys as a consistent **private** set using the procedure below. Drain admission/execution before updates; `update.py` fetches and builds reviewed commits before advancing checkouts under its API-stop fence. After restoring, reconcile allocations and source/proof identity before admitting execution. Never use `down -v`, volume pruning or broad VM deletion on an existing deployment.
No uploaded files, video, ZIP, database dump, live inventory, operational IP addresses, certificates, runtime logs or acceptance proofs are published. Keep those in ignored protected locations, not alongside tracked sources.
### Consistent backup (Linux root Bash shell)
Coordinate a maintenance window, pause submissions and let execution drain. Do not run this concurrently with updates, source maintenance, administrator writes or backup restores. Adapt the **inspected existing** project/paths below; no example value is a production host or credential. The subshell restores the same original containers on refusal/failure, never newly built images. It stops worker/watchdog only after the database is quiet so background cleanup cannot change artifacts during the dump. Prepare approved helper images before starting the fence.
```bash
(
set -euo pipefail
umask 077
cd /opt/otche-workspace/otche-deploy
PROJECT=EXISTING_PROJECT
dc() { docker compose --env-file /srv/otche/deployment/deploy.env -p "$PROJECT" \
-f compose.yaml -f /srv/otche/deployment/compose.production.yaml --profile execution "$@"; }
BACKUP="/srv/otche-backups/$(date -u +%Y%m%dT%H%M%SZ)"
mkdir -p "$BACKUP"
api=$(dc ps -q api)
test -n "$api"
ARTIFACT_VOLUME=$(docker inspect --format '{{range .Mounts}}{{if eq .Destination "/var/lib/otche"}}{{.Name}}{{end}}{{end}}' "$api")
test -n "$ARTIFACT_VOLUME"
resume="$api"
trap 'for id in $resume; do docker start "$id"; done' EXIT
docker stop --time 30 "$api"
active=$(dc exec -T postgres psql -U otche -d otche -At -v ON_ERROR_STOP=1 -c \
"SELECT (SELECT count(*) FROM jobs WHERE status IN ('queued','running')) + (SELECT count(*) FROM qualifications WHERE status IN ('queued','running'));")
test "$active" = 0 # Otherwise the trap restarts the SAME API and this backup is refused.
executors=$(dc ps -q worker watchdog)
resume="$executors $api"
for id in $executors; do docker stop --time 60 "$id"; done
dc exec -T postgres pg_dump -U otche -d otche -Fc > "$BACKUP/database.dump"
docker run --rm --network none --read-only --cap-drop ALL --security-opt no-new-privileges \
--user 10001:10001 --mount "type=volume,src=$ARTIFACT_VOLUME,dst=/artifacts,readonly" \
debian:bookworm-slim tar -C /artifacts -cpf - . > "$BACKUP/artifacts.tar"
# Include the protected config, proof files, credentials, TLS and approved commit pins.
tar -C /srv/otche -cpf "$BACKUP/private-config.tar" deployment secrets tls
dc ps -a -q | xargs docker inspect --format '{{.Name}} {{.Image}}' > "$BACKUP/container-images.txt"
(cd "$BACKUP"; sha256sum database.dump artifacts.tar private-config.tar container-images.txt > SHA256SUMS)
touch "$BACKUP/COMPLETE"
)
```
A failed run may leave partial files: accept a backup only with `COMPLETE`, matching hashes and a successful restore drill. The dump and tar contain private data and potentially malicious samples; encrypt and transfer through the approved backup channel, apply retention, and keep decryption keys separately. Protect signing keys using the signing system's own backup procedure; never copy them into runtime containers. Record source-VM backups/seals and operator proof dependencies too: the artifact/database backup alone does not back up PVE guests. Do not restore old proof expiry dates as if they were current attestations.
### Restore drill (new isolated volumes only)
Use a separate Linux Docker host with **no route or credentials to production PVE**, the approved source/images corresponding to the backup pins, and a root-only working directory. Never point the drill at existing production volumes, do not start worker/watchdog, and do not publish database ports. Verify archive hashes before extracting anything. The following restores to newly created, separately named volumes; `volume inspect` guards against accidentally reusing an earlier drill. Use the same PostgreSQL major version as the dump (17 here).
```bash
set -euo pipefail
umask 077
BACKUP=/srv/otche-backups/REVIEWED_BACKUP_DIRECTORY
test -f "$BACKUP/COMPLETE"
(cd "$BACKUP"; sha256sum -c SHA256SUMS)
# Place a copy of the original PostgreSQL password in this protected directory,
# readable only by UID70, using the approved secret-manager workflow; never print it.
test -f /srv/otche-restore/postgres-password
for volume in otche-drill-db otche-drill-artifacts; do
if docker volume inspect "$volume" >/dev/null 2>&1; then
echo "Refusing existing drill volume: $volume" >&2; exit 1
fi
docker volume create "$volume"
done
docker network create --internal otche-drill-private
docker run -d --name otche-drill-db --network otche-drill-private \
--user 70:70 --read-only --cap-drop ALL --security-opt no-new-privileges \
--tmpfs /tmp --tmpfs /var/run/postgresql \
--mount type=volume,src=otche-drill-db,dst=/var/lib/postgresql/data \
--mount type=bind,src=/srv/otche-restore/postgres-password,dst=/run/secrets/postgres-password,readonly \
-e POSTGRES_DB=otche -e POSTGRES_USER=otche \
-e POSTGRES_PASSWORD_FILE=/run/secrets/postgres-password postgres:17-alpine
# Wait for readiness (bounded); any failure below stops the drill.
for i in {1..30}; do
if docker exec otche-drill-db pg_isready -U otche -d otche; then break; fi
sleep 2
done
docker exec otche-drill-db pg_isready -U otche -d otche
docker exec -i otche-drill-db pg_restore -U otche -d otche --exit-on-error --single-transaction < "$BACKUP/database.dump"
docker run --rm -i --network none --read-only --cap-drop ALL --cap-add CHOWN --cap-add FOWNER \
--cap-add DAC_OVERRIDE --security-opt no-new-privileges --user 0:0 \
--mount type=volume,src=otche-drill-artifacts,dst=/artifacts \
debian:bookworm-slim tar -C /artifacts -xpf - < "$BACKUP/artifacts.tar"
docker exec otche-drill-db psql -U otche -d otche -v ON_ERROR_STOP=1 -c \
"SELECT count(*) AS jobs FROM jobs; SELECT count(*) AS artifacts FROM artifacts; SELECT count(*) AS held_allocations FROM allocations WHERE state <> 'deleted';"
docker run --rm -i --network none --read-only --cap-drop ALL --security-opt no-new-privileges \
--user 10001:10001 --mount type=volume,src=otche-drill-artifacts,dst=/artifacts,readonly \
debian:bookworm-slim tar -C /artifacts --compare -f - < "$BACKUP/artifacts.tar"
```
Check expected owner/account/job counts against the backup record and verify selected artifact sizes/hashes against restored database metadata without opening/executing samples. For an application-level drill, point the pinned API image at the isolated database and restored artifact volume, use a private loopback-only panel, and verify authentication plus an authorized artifact download; **never mount worker credentials or enable execution**. Keep the restored configuration archive sealed unless needed. Record the tested backup ID, image IDs and results privately. Remove only these exact drill containers/network/volumes after explicitly confirming they are disposable, not production resources.
For a real restore, retain the API-stop fence and a pre-restore backup, restore into new replacement volumes rather than overwriting evidence, then deliberately update the external volume mapping. Verify data and image/pin continuity first; reconcile every allocation with actual owned VMs, source seals and fresh isolation/qualification proofs before re-enabling worker/watchdog and admitting jobs. A database restore is not authority to delete retained VMs or reuse stale leases.
## Updating an existing deployment from Git
All application deployments must use actual canonical Git clones, not tar/ZIP copies of an operational workspace. Review source, test the affected behavior, commit and push to the owning repository, then select three **full approved commit IDs**. API/worker/watchdog share the selected backend commit. Never automatically deploy an unreviewed branch tip.
@@ -136,13 +236,21 @@ For a deployment with sibling clones under `/opt/otche-workspace`, run in the ro
```sh
cd /opt/otche-workspace/otche-deploy
git fetch origin main
git merge --ff-only APPROVED_DEPLOY_COMMIT
# Do not merge any repository before the updater's quiescence fence.
python3 update.py --backend APPROVED_BACKEND_COMMIT --frontend APPROVED_FRONTEND_COMMIT --deploy APPROVED_DEPLOY_COMMIT --env-file /srv/otche/deployment/deploy.env --override /srv/otche/deployment/compose.production.yaml --project EXISTING_PROJECT
# After coordinating paused new submissions and zero active jobs/qualifications:
python3 update.py --backend APPROVED_BACKEND_COMMIT --frontend APPROVED_FRONTEND_COMMIT --deploy APPROVED_DEPLOY_COMMIT --env-file /srv/otche/deployment/deploy.env --override /srv/otche/deployment/compose.production.yaml --project EXISTING_PROJECT --activate --confirm-quiescent
```
Use the literal reviewed 40-character commits, not the uppercase explanatory placeholders. The script refuses dirty/untracked work, detached/private branches, wrong origins, non-fast-forward/divergent targets and commits not reachable from public main. It validates all three repositories before advancing any, verifies the exact currently mounted persistent volumes, builds, and by default **does not restart anything**. After building, activation gracefully stops only the existing API container to fence new admissions, then checks the database for queued/running jobs and qualifications. If work is active or that check fails, it restarts the SAME old API container and refuses activation; existing workers and held terminal evidence are not stopped/released. Successful activation recreates web explicitly after the API update because nginx resolves its upstream at startup. Coordinate the maintenance window with users/operators; this is a short API interruption, not a new application maintenance framework. A failed command stops without reset/force or data removal; inspect before retrying. Source checkouts may already have advanced when a later build fails, while the previous containers remain running.
Use literal reviewed 40-character commits, not the uppercase explanatory placeholders. The script refuses dirty/untracked work, detached/private branches, wrong origins, non-fast-forward/divergent targets and commits not reachable from public main. The update sequence is:
1. Validate all three repositories and the exact project/external-volume identities. Fetch only outside `--verify-only`; verify-only requires HEAD to equal each approved pin, without network access or builds.
2. Create detached **sibling worktrees in `/var/tmp`** at the three targets, resolve the target Compose configuration with the same protected external environment/override, and recheck it against the running data volumes. A generated temporary override points every build context explicitly at the matching worktree and assigns temporary image tags. Build-only mode never moves the real checkouts, retags runtime images, restarts services or changes approved pins. Temporary worktrees are removed/pruned in `finally`; a killed process may leave only out-of-workspace staging directories/worktree metadata, never untracked deployment files. Inspect those paths before manually removing abandoned staging data.
3. With `--activate --confirm-quiescent`, capture old container image IDs and gracefully stop the **existing API container** to fence new admissions. Query PostgreSQL for queued/running jobs and qualifications. Active work or a failed query restarts the **same old API container** and refuses activation: old checkouts, runtime image tags and pins remain intact, so startup verification still passes. Existing workers and terminal held evidence are untouched on refusal.
4. Only after zero active work, recheck checkout cleanliness/ancestry, fast-forward the three reviewed commits, and promote the captured built image IDs to the exact runtime tags resolved by target Compose (normally `PROJECT-api`, `PROJECT-migrate`, `PROJECT-worker`, `PROJECT-watchdog`, `PROJECT-web`, all `:latest`). No rebuild is needed: the temporary worktrees used those exact commits. Start with **`--profile execution up --no-build --wait`**, including worker/watchdog; migration must complete before the dependent API/executors start. This deliberately activates execution containers, though all qualification/isolation gates still apply.
5. Explicitly recreate web and wait for health again. Compare each built service container's actual image ID with the captured build ID, require healthy/running services and a successful exited migration, then atomically write `approved-commits.env`. Missing/mismatched images or unhealthy services never approve pins. A failure after the fence may leave a partial rollout and advanced checkouts; the script prints the old container image IDs for **manual** recovery and never deletes old images or resets source. Retain the API admission fence until the operator has inspected migration compatibility and established a consistent recovery state; do not blindly downgrade a migrated database.
Coordinate the maintenance window and prohibit concurrent deployments/config edits. If upgrading from the old updater, do not first merge its deploy commit: after fetching and reviewing it, execute the reviewed script directly from Git while staying in the deploy checkout, e.g. in Bash with `set -o pipefail`: `git show APPROVED_DEPLOY_COMMIT:update.py | python3 - --backend APPROVED_BACKEND_COMMIT --frontend APPROVED_FRONTEND_COMMIT --deploy APPROVED_DEPLOY_COMMIT --env-file /srv/otche/deployment/deploy.env --override /srv/otche/deployment/compose.production.yaml --project EXISTING_PROJECT` (append the two activation flags for the second invocation). This uses the current directory to locate sibling checkouts without advancing them.
Keep systemd's WorkingDirectory and explicit `--env-file`, `-p`, `-f` arguments aligned with these same clones/private files. After activation compare actual HEADs with approved public commits, verify authenticated health and prior accounts/uploads/jobs/artifacts/seals/held evidence, and perform real acceptance before reopening submissions. Logs and the exact host-specific rollout record remain private; do not put credentials or live inventory into commit messages or this public guide.
+29 -4
View File
@@ -1,6 +1,12 @@
name: otche
x-logging: &logging
driver: local
options:
max-size: "10m"
max-file: "5"
services:
postgres:
logging: *logging
image: postgres:17-alpine
user: "70:70"
read_only: true
@@ -22,6 +28,7 @@ services:
networks: [private]
mem_limit: 768m
storage-init:
logging: *logging
image: debian:bookworm-slim
user: "0:0"
command: [sh, -ec, "chown 10001:10001 /var/lib/otche && chmod 0700 /var/lib/otche"]
@@ -31,6 +38,7 @@ services:
cap_add: [CHOWN, FOWNER]
security_opt: [no-new-privileges:true]
migrate:
logging: *logging
build: {context: ../otche-backend, target: api}
command: [migrate]
environment: {DATABASE_URL_FILE: /run/secrets/database-url}
@@ -42,6 +50,7 @@ services:
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
api:
logging: *logging
build: {context: ../otche-backend, target: api}
environment:
DATABASE_URL_FILE: /run/secrets/database-url
@@ -56,6 +65,12 @@ services:
depends_on:
migrate: {condition: service_completed_successfully}
storage-init: {condition: service_completed_successfully}
healthcheck:
test: ["CMD", "otche", "healthcheck"]
interval: 10s
timeout: 3s
retries: 6
start_period: 10s
restart: unless-stopped
networks: [private, edge]
read_only: true
@@ -65,6 +80,7 @@ services:
mem_limit: 256m
cpus: 1.0
worker:
logging: *logging
build: {context: ../otche-backend, target: worker}
command: [worker]
profiles: [execution]
@@ -85,11 +101,12 @@ services:
tmpfs: ["/tmp:size=128m,noexec,nosuid"]
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
mem_limit: 1536m
cpus: 2.0
pids_limit: 128
mem_limit: ${WORKER_MEM_LIMIT:-1536m}
cpus: ${WORKER_CPUS:-2.0}
pids_limit: ${WORKER_PIDS_LIMIT:-128}
stop_grace_period: 60s
watchdog:
logging: *logging
build: {context: ../otche-backend, target: worker}
command: [watchdog]
profiles: [execution]
@@ -108,9 +125,17 @@ services:
mem_limit: 128m
cpus: 0.25
web:
logging: *logging
build: {context: ../otche-frontend}
ports: ["${BIND_ADDRESS:-127.0.0.1}:${WEB_PORT:-8088}:8080"]
depends_on: [api]
depends_on:
api: {condition: service_healthy}
healthcheck:
test: ["CMD-SHELL", "wget -q --spider http://127.0.0.1:8080/ || wget -q --spider --no-check-certificate https://127.0.0.1:8443/"]
interval: 10s
timeout: 3s
retries: 6
start_period: 10s
restart: unless-stopped
networks: [edge]
read_only: true
+19 -8
View File
@@ -11,14 +11,15 @@ server {
ssl_session_timeout 10m;
ssl_session_tickets off;
client_max_body_size 0;
add_header Strict-Transport-Security "max-age=31536000" always;
add_header X-Content-Type-Options nosniff always;
add_header Referrer-Policy same-origin always;
add_header X-Frame-Options DENY always;
add_header Content-Security-Policy "default-src 'self'; script-src 'self'; style-src 'self'; img-src 'self' data:; media-src 'self'; connect-src 'self'; object-src 'none'; base-uri 'self'; frame-ancestors 'none'; form-action 'self'" always;
resolver 127.0.0.11 valid=10s ipv6=off;
gzip on;
gzip_proxied any;
gzip_types application/json application/javascript text/css;
gzip_vary on;
location /api/ {
proxy_pass http://api:8080;
set $api_upstream http://api:8080;
proxy_pass $api_upstream;
add_header Strict-Transport-Security "max-age=31536000" always;
proxy_http_version 1.1;
proxy_set_header Host $http_host;
proxy_set_header X-Forwarded-Proto https;
@@ -29,10 +30,20 @@ server {
}
location /assets/ {
try_files $uri =404;
expires 1y;
add_header Cache-Control "public, max-age=31536000, immutable";
add_header Strict-Transport-Security "max-age=31536000" always;
add_header X-Content-Type-Options nosniff always;
add_header Referrer-Policy same-origin always;
add_header X-Frame-Options DENY always;
add_header Content-Security-Policy "default-src 'self'; script-src 'self'; style-src 'self'; img-src 'self' data:; media-src 'self'; connect-src 'self'; object-src 'none'; base-uri 'self'; frame-ancestors 'none'; form-action 'self'" always;
}
location / {
try_files $uri $uri/ /index.html;
expires -1;
add_header Strict-Transport-Security "max-age=31536000" always;
add_header X-Content-Type-Options nosniff always;
add_header Referrer-Policy same-origin always;
add_header X-Frame-Options DENY always;
add_header Content-Security-Policy "default-src 'self'; script-src 'self'; style-src 'self'; img-src 'self' data:; media-src 'self'; connect-src 'self'; object-src 'none'; base-uri 'self'; frame-ancestors 'none'; form-action 'self'" always;
}
}
+120 -36
View File
@@ -1,12 +1,14 @@
#!/usr/bin/env python3
"""Update reviewed sibling Git commits, build, and optionally activate a quiet deployment."""
import argparse
from contextlib import contextmanager
import json
import os
from pathlib import Path
import re
import subprocess
import sys
import tempfile
def run(args, cwd=None, capture=True):
@@ -15,6 +17,69 @@ def run(args, cwd=None, capture=True):
return result.stdout.strip() if capture else None
@contextmanager
def staged_repositories(repositories):
with tempfile.TemporaryDirectory(prefix='otche-update-', dir='/var/tmp') as temporary:
root = Path(temporary)
added = []
try:
for directory, revision in repositories:
target = root / directory.name
run(['git', 'worktree', 'add', '--detach', str(target), revision], directory, False)
added.append((directory, target))
yield root
finally:
failures = []
for directory, target in reversed(added):
for command in (['git', 'worktree', 'remove', '--force', str(target)],
['git', 'worktree', 'prune']):
if subprocess.run(command, cwd=directory, check=False).returncode:
failures.append(str(target))
if failures:
raise RuntimeError('Temporary worktree cleanup failed: ' + ', '.join(failures))
def check_volumes(parser, compose, config, project, running):
if config['name'] != project:
parser.error('Resolved Compose project identity changed')
for logical in ('postgres-data', 'artifacts'):
volume = config['volumes'][logical]
if not volume.get('external') or not volume.get('name'):
parser.error('Existing deployment requires explicit external ' + logical + ' volume')
run(['docker', 'volume', 'inspect', volume['name']])
if running:
for service, destination, logical in (('postgres', '/var/lib/postgresql/data', 'postgres-data'),
('api', '/var/lib/otche', 'artifacts')):
container = run(compose + ['ps', '-q', service])
if not container:
parser.error('Expected running ' + service + '; inspect existing deployment before updating')
mounts = json.loads(run(['docker', 'inspect', container]))[0]['Mounts']
if not any(m.get('Name') == config['volumes'][logical]['name'] and m['Destination'] == destination for m in mounts):
parser.error('Persistent volume identity would change for ' + service)
def image_name(config, service):
return config['services'][service].get('image') or config['name'] + '-' + service
def verify_images(parser, compose, images):
for service, expected in images.items():
containers = run(compose + ['ps', '-a', '-q', service]).splitlines()
if not containers:
parser.error('Missing updated service container: ' + service)
for container in containers:
actual = json.loads(run(['docker', 'inspect', container]))[0]
if actual['Image'] != expected:
parser.error('Container image mismatch for ' + service + '; approved pins NOT changed')
state = actual['State']
if service == 'migrate':
ready = state['Status'] == 'exited' and state['ExitCode'] == 0
else:
ready = state['Running'] and state.get('Health', {}).get('Status', 'healthy') == 'healthy'
if not ready:
parser.error('Updated service is not ready: ' + service + '; approved pins NOT changed')
def main():
parser = argparse.ArgumentParser(description=__doc__)
for repo in ('backend', 'frontend', 'deploy'):
@@ -61,56 +126,75 @@ def main():
run(['git', 'merge-base', '--is-ancestor', 'HEAD', revision], directory)
run(['git', 'merge-base', '--is-ancestor', revision, 'origin/main'], directory)
repositories.append((directory, revision))
# Validate every repository before advancing any; only fast-forward reviewed commits.
if not args.verify_only:
for directory, revision in repositories:
run(['git', 'merge', '--ff-only', revision], directory, False)
compose = ['docker', 'compose', '--env-file', str(args.env_file), '-p', args.project,
'-f', str(deploy / 'compose.yaml'), '-f', str(args.override)]
config = json.loads(run(compose + ['--profile', 'execution', 'config', '--format', 'json']))
if config['name'] != args.project:
parser.error('Resolved Compose project identity changed')
for logical in ('postgres-data', 'artifacts'):
volume = config['volumes'][logical]
if not volume.get('external') or not volume.get('name'):
parser.error('Existing deployment requires explicit external ' + logical + ' volume')
run(['docker', 'volume', 'inspect', volume['name']], capture=True)
'-f', str(deploy / 'compose.yaml'), '-f', str(args.override), '--profile', 'execution']
config = json.loads(run(compose + ['config', '--format', 'json']))
check_volumes(parser, compose, config, args.project, not args.verify_only)
if args.verify_only:
print('Exact clean repository pins and external persistent volumes verified')
return
# Existing data sources must match the currently running services before any build/activation.
for service, destination, logical in (('postgres', '/var/lib/postgresql/data', 'postgres-data'),
('api', '/var/lib/otche', 'artifacts')):
container = run(compose + ['ps', '-q', service])
if not container:
parser.error('Expected running ' + service + '; inspect existing deployment before updating')
mounts = json.loads(run(['docker', 'inspect', container]))[0]['Mounts']
if not any(m.get('Name') == config['volumes'][logical]['name'] and m['Destination'] == destination for m in mounts):
parser.error('Persistent volume identity would change for ' + service)
run(compose + ['--profile', 'execution', 'build'], capture=False)
# Build reviewed targets without moving checkouts or replacing runtime image tags.
with staged_repositories(repositories) as staging:
staged_compose = ['docker', 'compose', '--env-file', str(args.env_file), '-p', args.project,
'-f', str(staging / 'otche-deploy' / 'compose.yaml'), '-f', str(args.override),
'--profile', 'execution']
target_config = json.loads(run(staged_compose + ['config', '--format', 'json']))
check_volumes(parser, compose, target_config, args.project, True)
built = {service: {'image': 'otche-update-' + args.project + '-' + service + ':' + args.deploy}
for service, settings in target_config['services'].items() if 'build' in settings}
for service, settings in built.items():
if service not in ('api', 'migrate', 'worker', 'watchdog', 'web'):
parser.error('Unknown build service requires explicit reviewed source mapping: ' + service)
source = 'otche-frontend' if service == 'web' else 'otche-backend'
settings['build'] = {'context': str(staging / source)}
build_override = staging / 'build-images.json'
build_override.write_text(json.dumps({'services': built}))
staged_compose += ['-f', str(build_override)]
run(staged_compose + ['build'], capture=False)
images = {service: run(['docker', 'image', 'inspect', '--format', '{{.Id}}', settings['image']])
for service, settings in built.items()}
if args.activate:
old_images = {}
for service in images:
containers = run(compose + ['ps', '-a', '-q', service]).splitlines()
old_images[service] = [json.loads(run(['docker', 'inspect', container]))[0]['Image']
for container in containers]
# Stop only API admission, after build; workers keep their existing containers.
old_api = run(compose + ['ps', '-q', 'api'])
run(['docker', 'stop', '--time', '30', old_api], capture=False)
query = "SELECT (SELECT count(*) FROM jobs WHERE status IN ('queued','running')) + (SELECT count(*) FROM qualifications WHERE status IN ('queued','running'));"
try:
active = run(compose + ['exec', '-T', 'postgres', 'psql', '-U', 'otche', '-d', 'otche', '-At', '-c', query])
if active != '0':
parser.error('Active work remains; refusing activation and restarting the SAME old API container')
for directory, revision in repositories:
if run(['git', 'status', '--porcelain', '--untracked-files=all'], directory):
parser.error(str(directory) + ' changed during build; refusing activation')
run(['git', 'merge-base', '--is-ancestor', 'HEAD', revision], directory)
except BaseException:
run(['docker', 'start', old_api], capture=False)
raise
if active != '0':
run(['docker', 'start', old_api], capture=False)
parser.error('Active work remains; restarted the SAME old API container, workers untouched')
run(compose + ['up', '-d', '--no-build', '--wait', '--wait-timeout', '120'], capture=False)
# nginx resolves its upstream at startup; do not rely on API IP reuse.
run(compose + ['up', '-d', '--no-deps', '--no-build', '--force-recreate', 'web'], capture=False)
pin_file = args.env_file.parent / 'approved-commits.env'
temporary = pin_file.with_suffix('.env.new')
with temporary.open('x') as output:
os.chmod(temporary, 0o600)
for component in ('backend', 'frontend', 'deploy'):
output.write('OTCHE_' + component.upper() + '_COMMIT=' + getattr(args, component) + '\n')
os.replace(temporary, pin_file)
try:
# Only the successful fence permits advancing source and runtime image tags.
for directory, revision in repositories:
run(['git', 'merge', '--ff-only', revision], directory, False)
for service, image in images.items():
run(['docker', 'image', 'tag', image, image_name(target_config, service)], capture=False)
run(compose + ['up', '-d', '--no-build', '--wait', '--wait-timeout', '120'], capture=False)
# Recreate web explicitly even when its image is unchanged, then await health.
run(compose + ['up', '-d', '--no-deps', '--no-build', '--force-recreate', '--wait', '--wait-timeout', '120', 'web'], capture=False)
verify_images(parser, compose, images)
pin_file = args.env_file.parent / 'approved-commits.env'
temporary = pin_file.with_suffix('.env.new')
with temporary.open('x') as output:
os.chmod(temporary, 0o600)
for component in ('backend', 'frontend', 'deploy'):
output.write('OTCHE_' + component.upper() + '_COMMIT=' + getattr(args, component) + '\n')
os.replace(temporary, pin_file)
except BaseException:
print('Activation failed; previous container image IDs for manual recovery (not deleted):', file=sys.stderr)
print(json.dumps(old_images, indent=2), file=sys.stderr)
raise
print(json.dumps({'activated': args.activate, 'project': args.project,
'commits': {d.name: r for d, r in repositories}}, indent=2))
if args.activate: