Files
ipx/.claude/skills/ipx-prod-check/SKILL.md
rays ec8fd5dd86 Sleep until the next feed is due instead of scanning every minute (#114)
The daemon ticked every 60 s and ran a scan pass each time: the sweep, then a check-state query
per feed (about 180) to find which were due. In the six hours before, 293 of 362 passes found
nothing due. Now, after each pass, it works out when the earliest feed is due (due_at, shared
with the scan's own check, over Db::http_states, one query) and sleeps until then: at least
30 s, so a feed that never gets a check time cannot spin it, and at most 10 minutes, so what no
command announces, ipx add or a shorter schedule, is picked up. Commands still wake it at once,
and the first pass after starting runs straight away, as the tick's did. The scan reads every
feed's state in one query too.

The prod-check skill says what to expect now: tens of scans in six hours, and pending as the
real queue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 19:56:13 +00:00

8.3 KiB

name, description
name description
ipx-prod-check Look for problems in production ipx (the iPX container on Tower) from its logs in Loki and its traces in Tempo, and file what is found as Gitea issues. Use when asked to check on production, look for issues or errors in ipx, see why something is slow or failing in production, review the logs, or investigate a report about the live site.

Checking production ipx

Production logs one JSON object a line (IPX_LOG_FORMAT=json), which Alloy ships to Loki under {container="iPX"}, and sends traces to Tempo tagged resource.deployment.environment.name="production". Anything without that tag is a daemon run by hand, not production. Loki and Tempo publish no query port, so use the script beside this file:

Q=.claude/skills/ipx-prod-check/query.py
$Q logs    '<LogQL>'           [since]   # lines, newest first (500 at most)
$Q metric  '<LogQL metric>'    [since]   # $range becomes `since`
$Q traces  '<TraceQL>'         [since]   # slowest first
$Q trace   <trace id>                    # one trace as a tree, with its log lines

since is 1h, 24h, 7d. Default to 24h; widen it to see whether something is new.

What the log carries

Every line has timestamp, level, message and target; lines inside a span have span (the innermost: {"name":"feed","feed":"x"}). Loki's | json flattens it to span_name, span_feed. Lines inside a traced span, requests and scans, also carry trace_id and span_id: give the trace_id to $Q trace to see the whole request or scan. (From 2026-09-29 16:30 UTC; before that, lines had no trace id, the access log's time was ms, and each event was logged twice, words under ipx::scan and fields under ipx::io.)

target fields what
ipx::http method, path, route, status, duration_ms one per web request; route is the pattern, empty for an unrouted path
ipx::scan ev and the event's own: feed, new, downloaded, failed, bytes, msg, url, feeds, reason; on a failure error.type and, from an HTTP error, http.response.status_code the daemon's events, one line each, in words; warnings are feed and download failures
ipx::io ev, feeds, pending, downloaded on the status reply commands arriving (-> {...}) and the healthcheck's answer
ipx message, sometimes fields start-up, shutdown, account and config messages

Events (ev): feed_start, feed_done (new, downloaded, failed, torrents), feed_skip (not due, routine), feed_error (msg), download_done (bytes), download_error (msg, url), torrent_deferred, reaped, scan_done (feeds checked), reap_done, status (feeds, pending, downloaded: the healthcheck's, every 30s), error (msg).

error.type is the HTTP status (404, 503) or one of dns, redirect_loop, timeout, tls, not_a_feed, site_message, connect, parse, other; Loki's | json names it error_type.

Filter on the text before | json where you can (|= "\"ev\":\"feed_error\""): it is much cheaper than parsing every line. Lines before 2026-09-29 14:00 UTC are text, not JSON, and | json | __error__="" drops them.

The checks

Run these, then read the lines behind whatever stands out. Most of the time is in the reading: a count says something happened, the lines and traces say why.

  1. Warnings and errors, grouped. What went wrong, how often, and since when.

    $Q metric 'sum by (target, message) (count_over_time({container="iPX"} | json | __error__="" | level=~"WARN|ERROR" [$range]))'
    

    Feed and download failures name the feed in the message; group them in the next check instead.

  2. Failing feeds and downloads.

    $Q metric 'sum by (feed, error_type) (count_over_time({container="iPX"} |= "\"ev\":\"feed_error\"" | json | __error__="" [$range]))' 7d
    $Q metric 'sum by (feed, error_type) (count_over_time({container="iPX"} |= "\"ev\":\"download_error\"" | json | __error__="" [$range]))' 7d
    

    Tell the publisher's problems from ipx's. A 404, 410, DNS failure or 503 from the feed's own server is the publisher (worth saying, since the feed may have moved; one issue for a feed that has been dead for days, not for a 503 once). A parse error on a feed that loads in a browser, a redirect loop ipx should follow, or the same failure on many feeds at once is ipx.

  3. Server errors and slow requests.

    $Q metric 'sum by (method, route, status) (count_over_time({container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | status >= 500 [$range]))'
    $Q metric 'topk(10, quantile_over_time(0.95, {container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | route != "" | route != "/api/events" | unwrap duration_ms [$range]) by (method, route))'
    

    Any 5xx is worth a look. 401s are people signing in, not a problem unless one address is hammering. For a slow route, find its traces (check 5) and see which span holds the time.

  4. Is the worker keeping up? Scans should finish regularly, the queue should drain, and the daemon should not be restarting on its own.

    $Q metric 'sum(count_over_time({container="iPX"} |= "\"ev\":\"scan_done\"" [$range]))' 6h
    $Q logs '{container="iPX"} |= "\"ev\":\"status\"" | json | line_format "{{.timestamp}} pending={{.pending}} downloaded={{.downloaded}}"' 6h
    $Q logs '{container="iPX"} |= "daemon started"' 7d
    

    A daemon started not matched by a deploy (see git log and the image's build time) is a crash or an OOM kill: check docker inspect iPX -f '{{.State.OOMKilled}} {{.RestartCount}}' and the lines just before it. A pending count that only grows means downloads are not keeping up or not running.

    Since 2026-10-02 the daemon sleeps until the next feed is due (at most 10 minutes) instead of scanning every minute (#114), so expect tens of scans in 6 hours, not 360, nearly all with feeds above 0; none at all for over 10 minutes means the worker is stuck. And pending is the real queue (#113): files a scan will download on its own. Back-catalogue files are held, listed but not counted, so it is usually 0 or a handful.

  5. Slow and failed traces.

    $Q traces '{resource.deployment.environment.name="production" && duration > 5s}'
    $Q traces '{resource.deployment.environment.name="production" && status = error}'
    $Q trace <id>
    

    Scans (scan) are long by nature, since they fetch many feeds one after another: look for one feed or fetch span holding most of it, or a download far slower than its size explains. A web request over a second is worth a look; the trace shows whether the time is in the handler or a scan it waited on.

Also check the container itself, since Loki cannot see a daemon that is not running:

docker ps --filter name=iPX --format '{{.Status}}'
docker logs --since 10m iPX 2>&1 | tail -5

docker logs and docker exec iPX ipx ... are fine. Do not query production's Postgres directly: ask the user if a question needs the database.

What to do with what you find

Follow the repository's rules in CLAUDE.md: every problem found gets a Gitea issue, with a closed stdin and a timeout on tea. Before filing, list the open issues and do not file one twice; comment on the existing issue with the new evidence instead.

R="--login git.sdf1.net --repo rays/ipx"
t() { timeout 30 /src/tea "$@" < /dev/null; }
t issues list $R --state open
t issues create $R -t "<what is wrong, as the user would notice it>" -L bug -d "<what, where, since when, how often, the query or trace id that shows it>"

Put in each issue what would let someone pick it up cold: the LogQL or TraceQL that shows it, a trace id, the first time it was seen and how often. A publisher's dead feed is worth one issue saying so (the user may want to unsubscribe or find its new address); a single 503 is not.

Finish with a short report to the user: what is healthy, what is wrong (with the issue numbers), and anything you could not tell from logs and traces alone. Do not fix things unless asked; the check is for finding them.

When the checks come back empty

Check that there is data before concluding all is well: $Q metric 'sum(count_over_time({container="iPX"} [1h]))' 1h should be in the hundreds or more. Nothing at all means Alloy is not shipping (it can take a few minutes to pick up a container after a deploy), or the container is down.