On-call runbook
Quick-reference playbook for the most common PagerDuty / notification alerts against a production TRUSCA stack. Each scenario lists:
- Symptom — what triggered the page
- Customer impact — what users can/cannot do right now
- Diagnose — exact commands to run (host + container)
- Recover — ordered remediation steps
- Escalate — when to wake the portal dev team
The PagerDuty alert names below are examples, not something this repository
wires up on its own: it publishes the underlying data (/metrics, Slack/Teams
webhooks) but ships no Prometheus server, Alertmanager, or paging
integration. See Alerting for an example rule file that
produces alerts with these same names from an already-running Prometheus, and
for which scenarios below have no metric behind them yet.
All commands assume docker-compose V1 (hyphen) and a bash host shell.
# Replace EMAIL/PASSWORD with the super-admin you created at install.
EMAIL=admin@example.com
PASSWORD=...
ACCESS_TOKEN=$(curl -fsS -X POST "https://<your-host>/api/auth/login" \
-H "Content-Type: application/json" \
-d "{\"email\":\"$EMAIL\",\"password\":\"$PASSWORD\"}" | jq -r '.access_token')
Scenario 1 — Trivy DB stale or missing
Symptom
PagerDuty: TrustedOSSVulnDbStale (example rule, see Alerting) or TRUSCA Trivy DB missing on worker. The /admin/health → Vulnerability data card's Trivy DB panel shows the same freshness classification.
Customer impact
- New scans CAN still be queued —
cdxgen+ scancode still produce SBOMs and licence findings. - New CVE detections stop landing until the DB refresh succeeds.
- Existing
vulnerability_findingsrows are unchanged — the gap is forward-only.
Diagnose
# 0. If METRICS_ENABLED, this is what actually fired the alert - confirms the
# page before you go looking at the worker.
curl -fsS "https://<your-host>/metrics" \
| grep -E 'trusca_vuln_db_(last_update_timestamp_seconds|refresh_interval_hours)'
# 1. Is the DB on disk?
docker-compose -f docker-compose.yml exec worker \
ls -lh /var/lib/trivy/db/
# 2. DB metadata (Created timestamp)
docker-compose -f docker-compose.yml exec worker \
cat /var/lib/trivy/db/metadata.json
# 3. Recent download / refresh logs
docker-compose -f docker-compose.yml logs --tail=500 worker | grep trivy_db
docker-compose -f docker-compose.yml logs --tail=500 beat | grep trivy_db_refresh
# 4. Outbound HTTPS to ghcr.io reachable?
docker-compose -f docker-compose.yml exec worker \
curl -fsS https://ghcr.io/v2/ -o /dev/null -w "%{http_code}\n"
Recover (in order)
- Force a one-shot refresh (preferred — single command, no restart):
docker-compose -f docker-compose.yml exec worker \celery -A apps.backend.tasks.celery_app call tasks.trivy_db.refreshsleep 30docker-compose -f docker-compose.yml exec worker \cat /var/lib/trivy/db/metadata.json | jq '.Created'
- Wipe + re-download (if metadata is corrupted):
The boot-timedocker-compose -f docker-compose.yml exec worker \rm -rf /var/lib/trivy/dbdocker-compose -f docker-compose.yml restart worker
trivy --download-db-onlyruns and re-populates the directory within 1–3 minutes. - Mirror fallback (if
ghcr.iois unreachable from the worker): pointTRIVY_DB_REPOSITORYat your internal mirror — see Vulnerability data — Air-gapped operation.
After recovery, the automatic re-match beat picks up missed CVEs against existing scans on its next cycle — no operator action.
Escalate
- If two refresh attempts fail with the same error, OR
- If the internal mirror itself reports
unauthorizeddespite recenttrivy registry login, OR - If
metadata.jsonexists butResultson a spot scan is empty across multiple ecosystems (suggests a schema mismatch).
Page the portal dev team with: worker logs (docker-compose logs --tail=2000 worker), the metadata.json content, and the output of trivy --version from inside the worker.
Scenario 2 — Auto-backup failed for 3 days
Symptom
PagerDuty: TrustedOSSAutoBackupNotSucceeding (example rule, see Alerting).
Customer impact
- All in-portal data is at risk if the host crashes (no recent backup to restore from). Plan downstream tasks (compliance freezes, etc.) accordingly until a fresh backup lands.
Diagnose
# 0. If METRICS_ENABLED, this is what actually fired the alert.
curl -fsS "https://<your-host>/metrics" \
| grep -A2 '^trusca_task_runs_24h{outcome="success",task="trustedoss.backup.run"'
# 1. Celery Beat schedule heartbeat
docker-compose logs --tail=500 beat | grep daily-auto-backup
# 2. Worker logs for backup task runs
docker-compose logs --tail=2000 worker | grep -E 'backup\.(completed|failed)' | tail -20
# 3. Most recent backup row + status
curl -fsS "https://<your-host>/v1/admin/backup/list" \
-H "Authorization: Bearer $ACCESS_TOKEN" | jq '.items[0:5]'
# 4. Disk free on the backup volume (BACKUPS_ROOT is mounted at
# /opt/trustedoss/backups in the backend container)
docker-compose -f docker-compose.yml exec backend df -h /opt/trustedoss/backups
Recover
-
Manual trigger (UI:
/admin/backup→ Run manual backup now, or):curl -fsS -X POST "https://<your-host>/v1/admin/backup/trigger" \-H "Authorization: Bearer $ACCESS_TOKEN" -
If manual also fails — run the host backup script directly:
scripts/backup.shis a host script: it shells out todocker-compose ... execforpg_dumpand tars the workspace mount, so run it on the host (not inside a container). It writes toBACKUP_DIRwhen set, otherwisebackups/<stamp>under the repo root (mounted at/opt/trustedoss/backups).# From the deploy directory on the host (where docker-compose.yml + .env live).BACKUP_DIR=backups/debug-$(date +%Y%m%d-%H%M%S) bash scripts/backup.sh --no-prune 2>&1.env not found→ run from the deploy directory, or the install is incomplete.- Server version mismatch →
postgresql-client-17missing in the postgres image (regression — escalate). - Disk full → see Scenario 4.
Escalate
- If
bash scripts/backup.shfails for non-disk, non-permission reasons, OR - If the most recent successful backup is older than 7 days (auto-purge window — restore options narrowing).
Scenario 3 — Scan stuck in running for ≥ 4 hours
Symptom
PagerDuty: TRUSCA scan running > 4h for project X. No rule in Alerting produces this yet, since /metrics publishes a count of scans by status, not how long any one of them has been running, so this page currently has to come from your own query against the API or database rather than /metrics.
Customer impact
- That project: blocked from new scans (one-running-at-a-time).
- Other projects: unaffected unless worker concurrency = 1 (default 2).
Diagnose
# 1. Which stage is it stuck at?
curl -fsS "https://<your-host>/v1/scans/<scan_id>" \
-H "Authorization: Bearer $ACCESS_TOKEN" | jq '.progress_payload, .latest_log_frame'
# 2. Celery active tasks
docker-compose exec worker celery -A apps.backend.tasks.celery_app inspect active
# 3. Worker process tree (look for orphaned subprocesses)
docker-compose exec worker ps -ef | grep -E 'cdxgen|ort|trivy'
Recover
- Force-cancel the scan (preferred — no worker-wide impact):
curl -fsS -X POST "https://<your-host>/v1/admin/scans/<scan_id>/cancel" \-H "Authorization: Bearer $ACCESS_TOKEN"
- If cancel doesn't release the task (worker truly hung):
Other in-flight scans on the same worker will be marked failed and require manual re-run.# Last resort — kills all in-flight tasks on this worker.docker-compose restart worker
Escalate
- If the same project hangs at the same stage twice in a row (suggests a content-side issue — large git history, malformed lockfile, or
trivy sbomtimeout). Page portal dev team with<scan_id>and the last 200 lines ofworkerlogs filtered to that task.
Scenario 4 — Host disk ≥ 95%
Symptom
PagerDuty: TrustedOSSWorkspaceDiskCritical (example rule, see Alerting; covers only the workspace mount; a host-wide disk alert needs a node-level exporter, since /metrics does not publish that).
Customer impact
- In-flight scans continue. New scans are blocked at the
DISK_HARD_LIMIT_PCTthreshold (default 95%) —/admin/scansshows them as queued indefinitely.
Diagnose
# 0. If METRICS_ENABLED, this is the workspace-mount half of what fired.
curl -fsS "https://<your-host>/metrics" | grep trusca_workspace_disk_used_ratio
# 1. Host-wide
df -h /opt/trustedoss
docker system df
# 2. Per-card breakdown via the portal
curl -fsS "https://<your-host>/v1/admin/disk" \
-H "Authorization: Bearer $ACCESS_TOKEN" | jq
# 3. Workspace breakdown (most common offender)
docker-compose exec worker du -sh /workspace/* | sort -h | tail -10
# 4. Postgres database size
docker-compose exec postgres psql -U trustedoss -d trustedoss \
-c "SELECT pg_size_pretty(pg_database_size('trustedoss'));"
Recover
- Workspace cleanup (almost always the answer):
docker-compose exec worker find /workspace -mindepth 1 -mtime +30 -delete
- Postgres bloat (if
pg_database_size> 2 GB and growth is recent): VACUUM the heavy tables.docker-compose exec postgres psql -U trustedoss -d trustedoss \-c "VACUUM FULL audit_logs, vulnerability_findings;" - Trivy DB volume (if
/admin/diskshowstrivy_dbat fault): the Trivy DB is ~500 MB and should not grow further; if it has, prune the cache and re-download (docker-compose -f docker-compose.yml exec worker rm -rf /var/lib/trivy/db && docker-compose restart worker). - Temporary threshold raise (only as a stop-gap, NOT a fix):
# Edit .env: DISK_HARD_LIMIT_PCT=98docker-compose up -d backend worker
Escalate
- After workspace cleanup, disk still > 90%, OR
- Postgres growth is from
audit_logsdoubling every 24 hours (root cause needed — possibly a runaway integration emitting events).
Scenario 5 - Queue backlog alert fired
Symptom
A Slack/Teams message titled "Queue backlog alert" for trustedoss.scan or trustedoss.default (the existing notification channels, this is not a new integration). Requires QUEUE_BACKLOG_ALERT_ENABLED + QUEUE_BACKLOG_METRICS_ENABLED to both be on; see Environment variables - Queue backlog alert and Docker Compose - Scan capacity.
Customer impact
trustedoss.scan: new scans queue behind existing ones and take longer to start. Nothing fails outright, this is a capacity signal, not an error.trustedoss.default: notifications, backups, audit export, and ticket webhooks are delayed. If sustained, treat it as a possible stuck worker rather than pure overload (see Diagnose).
Diagnose
# 1. Current backlog and oldest-queued-scan wait (needs METRICS_ENABLED too)
curl -fsS "https://<your-host>/metrics" | grep -E 'trusca_broker_queue_backlog|trusca_scan_queue_wait_seconds'
# 2. Are the worker-scan replicas actually up and consuming?
docker-compose -f docker-compose.yml ps worker-scan
docker-compose -f docker-compose.yml exec worker-scan celery -A apps.backend.tasks.celery_app inspect active
# 3. Recent scan throughput - are scans finishing, or piling up?
curl -fsS "https://<your-host>/v1/admin/scans?status=queued" \
-H "Authorization: Bearer $ACCESS_TOKEN" | jq '.total'
Recover
- This is arrival rate over capacity (the common case): scale the worker service the alert named.
worker-scanfor scan throughput,worker-defaultfor everything else; scaling the wrong one does nothing (see the capacity guide linked above):docker-compose -f docker-compose.yml up -d --scale worker-scan=4 - This is a stuck worker, not real load: if
celery inspect activeshows a task that has been running far longer than a normal scan (compare againstSCAN_HARD_TIME_LIMIT_SECONDS), follow Scenario 3's recovery steps for that scan first. Clearing the stuck task frees the slot without permanently adding capacity you don't need. - Confirm it clears: the alert re-fires on a cooldown (
QUEUE_BACKLOG_ALERT_COOLDOWN_SECONDS, default 1h) while still breached, and stops once the backlog drops back at or under its threshold on a later beat tick (checked every 5 minutes).
Escalate
- If scaling
worker-scandoes not bring the backlog down within one scan-duration cycle (suggests the bottleneck is elsewhere: disk, Postgres, or the broker itself), OR - If the alert keeps re-firing (past its cooldown) for the same queue across multiple days.
Scenario 6 - A worker restarts on boot with task_registry.empty
Symptom
A worker container will not stay up. Its log ends with a line naming
task_registry.empty, or on the process's own stderr:
FATAL task_registry.empty: this worker registered none of the portal's tasks
Under Compose the service restarts in a loop; under Kubernetes the pod reports
CrashLoopBackOff. This has to page from your orchestrator's own restart-count
signal (Kubernetes, cAdvisor, kube-state-metrics), not from /metrics: a
worker crash-looping on boot never reaches the code that endpoint lives in
long enough to be scraped. See Alerting.
The container exits with 78. If the log shows nothing at all, that number
alone is the diagnosis: the worker stopped because it had none of the portal's
tasks. A plain 1 is what the process reports for almost any other reason it
dies, so 78 is reserved for this one and every step of the refusal is written
to survive an output failure.
Customer impact
Scans queued for that worker do not run. They stay queued rather than disappearing, which is the point of the worker stopping: a worker with no tasks that keeps running takes those messages, cannot execute them, and discards them, and a healthy worker started afterwards finds nothing left to do.
If another worker of the same kind is healthy it will pick the work up and there is no customer impact at all, only a restarting container.
Diagnose
The worker's task list is built from an include list in the application, so this is a deployment or packaging fault rather than a configuration one. It means either that the list is empty or that no module on it imported.
# 1. What the worker says on the way down
docker-compose -f docker-compose.yml logs --tail=50 worker-scan | grep -i task_registry
# 2. Whether the modules import at all in that image
docker-compose -f docker-compose.yml run --rm --entrypoint python worker-scan \
-c "from tasks.celery_app import celery_app; celery_app.loader.import_default_modules(); \
print(len([t for t in celery_app.tasks if t.startswith('trustedoss.')]))"
The second command prints the count the guard would see. A healthy image prints a number in the twenties or thirties. Zero, or a traceback, is the fault.
Recover
- A mismatched or partial image: pull the tagged image again and recreate the service. An image built from an incomplete tree is the usual cause.
- A module that fails to import: the second command above prints the traceback. That is a defect in the release rather than something to configure around; roll back to the previous tag.
Escalate
Immediately if the second command shows a traceback on an official release image: every worker of that kind in every deployment on that tag is affected.
Scenario 7 - Redis readiness field reports degraded
Symptom
GET /health/ready keeps returning 200, but its body carries
"redis": "degraded" instead of "ok". This is not itself a PagerDuty
trigger, since nothing pages on it today - the field lives on /health/ready,
not /metrics, so it is also not in Alerting's example rule
file - but it shows up while diagnosing something else (a support ticket
about slow or missing rate limiting, or a sweep from Scenario 5's
queue-backlog investigation that turns out to be Redis-shaped), or an
operator adds their own monitor on the field and it fires.
Customer impact
None of the request path's controls fail closed on Redis, so nothing stops working. What degrades silently:
- Rate limiting (
core/ratelimit.py) and the login-guess throttle (core/login_throttle.py) both fail open: every request is allowed as if it had never been rate-limited, for as long as Redis stays unreachable. - The WebSocket connection registry (
core/ws_registry.py) also fails open: a new connection is admitted uncapped rather than refused, and the per-user/global connection caps are not enforced for as long as Redis stays unreachable. Live scan-progress streaming itself is unaffected either way, since scan state lives in Postgres. - Anything else in the deployment that reads Redis (Celery's broker/result backend) is a separate failure with its own symptoms. A queue backlog alert (Scenario 5) is the more likely page for that, not this field.
redis is a live ping and can read "ok" on a check that happens to land
between failures. When present, the response also carries a
redis_fail_open object ({"ratelimit": {...}, "login_throttle": {...}, "ws_registry": {...}}, each entry shaped {"count": N, "last_degraded_at": "<ISO 8601>", "worker_pid": P}), one entry per control that has actually
fallen back to fail-open on the request path, independent of this call's
own ping. Its absence means none of the three controls has degraded since
the process started; a non-empty last_degraded_at older
than the current redis: "ok" reading is the signature of an intermittent
problem (a flaky network path, a misconfigured REDIS_URL password that
some commands reject and others do not) rather than the outage having
already ended.
This state is per-process: with UVICORN_WORKERS greater than 1, each
worker keeps its own copy, and one HTTP response only reflects whichever
worker answered it. worker_pid names that worker, so polling the endpoint
a few times and seeing the field appear, disappear, and reappear under
different worker_pid values means several workers have each independently
hit the outage, not that it started and stopped repeatedly. Nothing here
aggregates counts across workers; add up count per worker_pid by hand if
that total is needed.
Diagnose
# 1. Confirm the field and read the schema-readiness side too (this endpoint
# is unauthenticated - no token needed).
curl -fsS https://<your-host>/health/ready | jq
# 2. Is Redis actually unreachable from the backend container?
docker-compose -f docker-compose.yml exec backend python -c \
"from redis import Redis; from core.config import redis_url; print(Redis.from_url(redis_url(), socket_timeout=1).ping())"
# 3. Is the Redis container itself up?
docker-compose -f docker-compose.yml ps redis
docker-compose -f docker-compose.yml logs --tail=100 redis
Recover
- Redis container down or restarting: bring it back
(
docker-compose -f docker-compose.yml up -d redis) and confirm step 2 above returnsTrue./health/ready'sredisfield flips to"ok"on the next poll: there is no cache to clear or service to restart on the backend side, the field is read live on every request. - Redis up but unreachable from
backend(network policy, DNS, wrongREDIS_URL): fix the network path or the URL and redeploy the backend config; this is the sameREDIS_URLthe rate limiter and login throttle use, so fixing it restores all three at once. - Redis up and reachable, field still
degraded: the probe uses a 1s connect/socket timeout (core/readiness.py), so a Redis under enough load to answerPINGslower than that will also read as degraded. Check Redis's own latency/CPU before assuming a network fault.
Escalate
Only if degraded persists for an extended period (hours) with no
corresponding Redis outage found in steps 2-3 above. That is a bug in the
probe itself, not an infrastructure incident.
Standard escalation form
When paging the portal dev team, attach:
- Scenario number (1-5) and PagerDuty alert URL.
- Portal version:
docker-compose -f docker-compose.yml exec backend python -c "from main import app; print(app.version)" - Last 2000 lines of the relevant container:
docker-compose logs --tail=2000 <svc> - For Trivy DB issues: the worker's
/var/lib/trivy/db/metadata.jsoncontent anddocker-compose logs --tail=500 worker | grep trivy_db. - For scan issues:
<scan_id>and/v1/scans/<scan_id>full JSON.
See also
- Hardening - the checklist to work through before this runbook is needed for real traffic.
- Postgres sizing and connection tuning - the connection-budget model behind the
connection_budget.over_max_connectionswarning. - Alerting - the example Prometheus rules behind the PagerDuty alert names above, and which scenarios have none yet.
- Vulnerability data (Trivy DB) — DB lifecycle and troubleshooting.
- Backup and restore — backup retention + restore flow.
- Disk and health — disk threshold model + Health dashboard.
- Docker Compose - Scan capacity - slot capacity formula and scaling
worker-scan/worker-default.