The 2026-09-12 CT 101 502 incident fix (api-watchdog: detects the API container hanging while docker still reports it healthy, and restarts it) only ever landed in argus-appliance's docker/docker-compose.yml, the source-build dev compose. It never reached this repo, which is what update.sh/install.sh on every real Docker-based install actually use -- so the mitigation never shipped to a real install. Ported the service (built locally from a bundled Dockerfile, since there is no released, cosign-verified argus-api-watchdog image yet) plus its build context.
api-watchdog
Recovers from a real production incident class (2026-09-11/12, CT 101): the api container's
process hung/wedged internally while docker ps still reported it Up -- it never actually
exited, so Docker's own restart: unless-stopped policy (which only reacts to a container
exiting) never fired. The container's HEALTHCHECK (wget --spider http://localhost:8080/health)
correctly kept detecting the hang the entire time -- docker inspect's .State.Health.Status
genuinely flipped to unhealthy -- but nothing was watching that status and acting on it, so the
outage lasted roughly 11 hours until a human happened to notice and ran docker restart by hand.
What it does
Polls docker-socket-proxy (the same already-existing, scoped Docker API proxy the api service
itself uses -- see docker/docker-socket-proxy/README.md) for argus-api's health status every
WATCHDOG_INTERVAL_SECONDS (default 30s). After WATCHDOG_UNHEALTHY_THRESHOLD (default 5)
consecutive unhealthy readings (~2.5 minutes sustained, not one blip), it issues
POST /containers/argus-api/restart through the same proxy and resets its counter. A healthy
reading at any point resets the counter immediately.
Deliberately talks to docker-socket-proxy, never a raw /var/run/docker.sock mount -- zero new
grants needed (CONTAINERS=1/ALLOW_RESTARTS=1 are already on for the api service's own
diagnostics recovery actions), and this container never gets docker.sock access at all.
Why not just shorten retries/rely on restart: unless-stopped alone
restart: unless-stopped restarts a container Docker itself observes as exited -- it has no
concept of "unhealthy," and this incident's whole failure mode was a container that never
exited. Docker Engine does not ship a built-in "restart on unhealthy" action; some setups solve
this with the third-party willfarrell/autoheal image, but that requires mounting the real
docker.sock into a new container. Reusing the already-present, already-audited
docker-socket-proxy instead keeps this fix inside ARGUS's existing security boundary.
Tuning
WATCHDOG_TARGET_CONTAINER(defaultargus-api)WATCHDOG_INTERVAL_SECONDS(default30, matches the api healthcheck's own interval)WATCHDOG_UNHEALTHY_THRESHOLD(default5) -- consecutive unhealthy checks before restarting
Verification
Sandbox-tested (docker-compose.local-sandbox.yml) by temporarily forcing the api service's
healthcheck to always fail (test: ["CMD", "false"]), confirming: the watchdog logs 5 consecutive
unhealthy checks, then a real POST .../restart through the proxy, and argus-api's container
start time visibly advances. Reverted the forced-failure healthcheck override afterward -- this
watchdog is not itself part of that sandbox scenario's cleanup.