Answer status on the socket instead of queuing it behind the worker
The worker runs one job at a time, and status was one of its jobs, so the Docker healthcheck waited behind the startup scan (54 seconds of it after the last deploy) and timed out at 5. Any scan or download longer than three probes would have had a working daemon marked unhealthy. The socket now answers status straight away; everything else still queues. A test fills the queue and checks status comes back anyway. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TAC7sLVqfKmY6rsTLXzNgk
This commit is contained in:
10
CLAUDE.md
10
CLAUDE.md
@@ -46,7 +46,10 @@ docker tag mirror.gcr.io/library/rust:1-slim-bookworm rust:1-slim-bookworm
|
||||
Run those again now and then, or the local copies go stale.
|
||||
|
||||
The healthcheck runs `ipx status` against the control socket, so `(healthy)` in `docker ps` means
|
||||
the worker is alive, not just the web port. The container restarts on its own after a reboot.
|
||||
the daemon answers there and can read its database, not just that the web port is up. The socket
|
||||
answers `status` itself instead of queuing it behind the worker's current job, so a long scan or
|
||||
download does not fail the check; it also means a worker stuck on one job would still pass. The
|
||||
container restarts on its own after a reboot.
|
||||
|
||||
Before the container, ipx ran by hand in code-server, with its files in `/config/.config/ipx/` and
|
||||
`/config/.local/share/ipx/`. Those are still there and the container does not read them. If you run
|
||||
@@ -114,8 +117,9 @@ Non-trivial logic leaves one runnable check behind. Pure functions (`merge_polic
|
||||
watch the shutdown channel itself; the daemon ignored SIGTERM for exactly this reason.
|
||||
* Only one daemon per socket. Removing the socket file defeats the guard and you get two daemons
|
||||
fighting over the database, with the stale one still holding the port.
|
||||
* `/api/settings` answering `200` does **not** mean the worker is alive — it is a different task.
|
||||
Probe the control socket (`ipx status`) to check that.
|
||||
* `/api/settings` answering `200` does **not** mean the daemon is well — the web server is a
|
||||
different task. `ipx status` checks the control socket and the database; to see the worker
|
||||
getting through its jobs, watch for `scan complete` in the log.
|
||||
* **Every `ipx` command runs `migrate()` when it opens the database**, the healthcheck's
|
||||
`ipx status` included. A migration that rewrites a big table (`DROP COLUMN`) takes seconds on
|
||||
production, and a command run meanwhile fails with `migrating schema`. It changes nothing; wait
|
||||
|
||||
Reference in New Issue
Block a user