A skill for checking production from Loki and Tempo
.claude/skills/ipx-prod-check: what production's JSON log and traces carry, the queries that find trouble (warnings grouped, failing feeds and downloads, 5xx and slow routes, whether the worker keeps up, slow and failed traces), how to tell a publisher's dead feed from an ipx bug, and filing what is found as issues per CLAUDE.md. query.py beside it runs the LogQL and TraceQL through a throwaway container on the monitoring network, since Loki and Tempo publish no query port. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
127
.claude/skills/ipx-prod-check/SKILL.md
Normal file
127
.claude/skills/ipx-prod-check/SKILL.md
Normal file
@@ -0,0 +1,127 @@
|
|||||||
|
---
|
||||||
|
name: ipx-prod-check
|
||||||
|
description: Look for problems in production ipx (the iPX container on Tower) from its logs in Loki and its traces in Tempo, and file what is found as Gitea issues. Use when asked to check on production, look for issues or errors in ipx, see why something is slow or failing in production, review the logs, or investigate a report about the live site.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Checking production ipx
|
||||||
|
|
||||||
|
Production logs one JSON object a line (`IPX_LOG_FORMAT=json`), which Alloy ships to Loki under
|
||||||
|
`{container="iPX"}`, and sends traces to Tempo tagged
|
||||||
|
`resource.deployment.environment.name="production"`. Anything without that tag is a daemon run by
|
||||||
|
hand, not production. Loki and Tempo publish no query port, so use the script beside this file:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
Q=.claude/skills/ipx-prod-check/query.py
|
||||||
|
$Q logs '<LogQL>' [since] # lines, newest first (500 at most)
|
||||||
|
$Q metric '<LogQL metric>' [since] # $range becomes `since`
|
||||||
|
$Q traces '<TraceQL>' [since] # slowest first
|
||||||
|
$Q trace <trace id> # one trace as a tree, with its log lines
|
||||||
|
```
|
||||||
|
|
||||||
|
`since` is `1h`, `24h`, `7d`. Default to `24h`; widen it to see whether something is new.
|
||||||
|
|
||||||
|
## What the log carries
|
||||||
|
|
||||||
|
Every line has `timestamp`, `level`, `message` and `target`; lines inside a span have `span` (the
|
||||||
|
innermost: `{"name":"feed","feed":"x"}`). Loki's `| json` flattens it to `span_name`, `span_feed`.
|
||||||
|
|
||||||
|
| target | fields | what |
|
||||||
|
|---|---|---|
|
||||||
|
| `ipx::http` | `method`, `path`, `route`, `status`, `ms` | one per web request; `route` is the pattern, empty for an unrouted path |
|
||||||
|
| `ipx::io` | `ev` and the event's own: `feed`, `new`, `downloaded`, `failed`, `bytes`, `msg`, `url`, `feeds`, `pending`, `reason` | the daemon's events as they go on the wire |
|
||||||
|
| `ipx::scan` | message only | the same events in words; warnings are feed and download failures |
|
||||||
|
| `ipx` | message, sometimes fields | start-up, shutdown, account and config messages |
|
||||||
|
|
||||||
|
Events (`ev`): `feed_start`, `feed_done` (new, downloaded, failed, torrents), `feed_skip` (not due,
|
||||||
|
routine), `feed_error` (msg), `download_done` (bytes), `download_error` (msg, url),
|
||||||
|
`torrent_deferred`, `reaped`, `scan_done` (feeds checked), `reap_done`, `status` (feeds, pending,
|
||||||
|
downloaded: the healthcheck's, every 30s), `error` (msg).
|
||||||
|
|
||||||
|
Filter on the text before `| json` where you can (`|= "\"target\":\"ipx::io\""`): it is much
|
||||||
|
cheaper than parsing every line. Lines before 2026-09-29 14:00 UTC are text, not JSON, and
|
||||||
|
`| json | __error__=""` drops them.
|
||||||
|
|
||||||
|
## The checks
|
||||||
|
|
||||||
|
Run these, then read the lines behind whatever stands out. Most of the time is in the reading:
|
||||||
|
a count says something happened, the lines and traces say why.
|
||||||
|
|
||||||
|
1. **Warnings and errors, grouped.** What went wrong, how often, and since when.
|
||||||
|
```
|
||||||
|
$Q metric 'sum by (target, message) (count_over_time({container="iPX"} | json | __error__="" | level=~"WARN|ERROR" [$range]))'
|
||||||
|
```
|
||||||
|
Feed and download failures name the feed in the message; group them in the next check instead.
|
||||||
|
2. **Failing feeds and downloads.**
|
||||||
|
```
|
||||||
|
$Q metric 'sum by (feed, msg) (count_over_time({container="iPX"} |= "\"target\":\"ipx::io\"" | json | __error__="" | ev="feed_error" [$range]))' 7d
|
||||||
|
$Q metric 'sum by (feed, msg) (count_over_time({container="iPX"} |= "\"target\":\"ipx::io\"" | json | __error__="" | ev="download_error" [$range]))' 7d
|
||||||
|
```
|
||||||
|
Tell the publisher's problems from ipx's. A 404, 410, DNS failure or 503 from the feed's own
|
||||||
|
server is the publisher (worth saying, since the feed may have moved; one issue for a feed
|
||||||
|
that has been dead for days, not for a 503 once). A parse error on a feed that loads in a
|
||||||
|
browser, a redirect loop ipx should follow, or the same failure on many feeds at once is ipx.
|
||||||
|
3. **Server errors and slow requests.**
|
||||||
|
```
|
||||||
|
$Q metric 'sum by (method, route, status) (count_over_time({container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | status >= 500 [$range]))'
|
||||||
|
$Q metric 'topk(10, quantile_over_time(0.95, {container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | route != "" | route != "/api/events" | unwrap ms [$range]) by (method, route))'
|
||||||
|
```
|
||||||
|
Any 5xx is worth a look. 401s are people signing in, not a problem unless one address is
|
||||||
|
hammering. For a slow route, find its traces (check 5) and see which span holds the time.
|
||||||
|
4. **Is the worker keeping up?** Scans should finish regularly, the queue should drain, and the
|
||||||
|
daemon should not be restarting on its own.
|
||||||
|
```
|
||||||
|
$Q metric 'sum(count_over_time({container="iPX"} |= "\"ev\":\"scan_done\"" [$range]))' 6h
|
||||||
|
$Q logs '{container="iPX"} |= "\"ev\":\"status\"" | json | line_format "{{.timestamp}} pending={{.pending}} downloaded={{.downloaded}}"' 6h
|
||||||
|
$Q logs '{container="iPX"} |= "daemon started"' 7d
|
||||||
|
```
|
||||||
|
A `daemon started` not matched by a deploy (see `git log` and the image's build time) is a
|
||||||
|
crash or an OOM kill: check `docker inspect iPX -f '{{.State.OOMKilled}} {{.RestartCount}}'`
|
||||||
|
and the lines just before it. A pending count that only grows means downloads are not
|
||||||
|
keeping up or not running.
|
||||||
|
5. **Slow and failed traces.**
|
||||||
|
```
|
||||||
|
$Q traces '{resource.deployment.environment.name="production" && duration > 5s}'
|
||||||
|
$Q traces '{resource.deployment.environment.name="production" && status = error}'
|
||||||
|
$Q trace <id>
|
||||||
|
```
|
||||||
|
Scans (`scan`) are long by nature, since they fetch many feeds one after another: look for one
|
||||||
|
`feed` or `fetch` span holding most of it, or a `download` far slower than its size explains.
|
||||||
|
A web request over a second is worth a look; the trace shows whether the time is in the
|
||||||
|
handler or a scan it waited on.
|
||||||
|
|
||||||
|
Also check the container itself, since Loki cannot see a daemon that is not running:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
docker ps --filter name=iPX --format '{{.Status}}'
|
||||||
|
docker logs --since 10m iPX 2>&1 | tail -5
|
||||||
|
```
|
||||||
|
|
||||||
|
`docker logs` and `docker exec iPX ipx ...` are fine. Do not query production's Postgres
|
||||||
|
directly: ask the user if a question needs the database.
|
||||||
|
|
||||||
|
## What to do with what you find
|
||||||
|
|
||||||
|
Follow the repository's rules in CLAUDE.md: **every problem found gets a Gitea issue**, with a
|
||||||
|
closed stdin and a timeout on `tea`. Before filing, list the open issues and do not file one
|
||||||
|
twice; comment on the existing issue with the new evidence instead.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
R="--login git.sdf1.net --repo rays/ipx"
|
||||||
|
t() { timeout 30 /src/tea "$@" < /dev/null; }
|
||||||
|
t issues list $R --state open
|
||||||
|
t issues create $R -t "<what is wrong, as the user would notice it>" -L bug -d "<what, where, since when, how often, the query or trace id that shows it>"
|
||||||
|
```
|
||||||
|
|
||||||
|
Put in each issue what would let someone pick it up cold: the LogQL or TraceQL that shows it, a
|
||||||
|
trace id, the first time it was seen and how often. A publisher's dead feed is worth one issue
|
||||||
|
saying so (the user may want to unsubscribe or find its new address); a single 503 is not.
|
||||||
|
|
||||||
|
Finish with a short report to the user: what is healthy, what is wrong (with the issue numbers),
|
||||||
|
and anything you could not tell from logs and traces alone. Do not fix things unless asked; the
|
||||||
|
check is for finding them.
|
||||||
|
|
||||||
|
## When the checks come back empty
|
||||||
|
|
||||||
|
Check that there is data before concluding all is well: `$Q metric 'sum(count_over_time({container="iPX"} [1h]))' 1h`
|
||||||
|
should be in the hundreds or more. Nothing at all means Alloy is not shipping (it can take a few
|
||||||
|
minutes to pick up a container after a deploy), or the container is down.
|
||||||
89
.claude/skills/ipx-prod-check/query.py
Executable file
89
.claude/skills/ipx-prod-check/query.py
Executable file
@@ -0,0 +1,89 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Ask production's Loki or Tempo a question, from anywhere that can run docker on Tower.
|
||||||
|
|
||||||
|
query.py logs '<LogQL log query>' [since] lines, newest first
|
||||||
|
query.py metric '<LogQL metric query>' [since] one value per series; $range is `since`
|
||||||
|
query.py traces '<TraceQL query>' [since] matching traces, slowest first
|
||||||
|
query.py trace <trace id> one trace's spans, as a tree
|
||||||
|
|
||||||
|
`since` is 1h, 24h, 7d and the like (default 24h). Loki and Tempo publish no query port on the
|
||||||
|
host, so each call runs a throwaway alpine container on the monitoring project's network.
|
||||||
|
"""
|
||||||
|
import json, subprocess, sys, time, urllib.parse
|
||||||
|
|
||||||
|
NET = "monitoring_default"
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
out = subprocess.run(["docker", "run", "--rm", "--network", NET, "alpine", "wget", "-qO-", url],
|
||||||
|
capture_output=True, text=True, timeout=120)
|
||||||
|
if out.returncode:
|
||||||
|
sys.exit(f"query failed: {out.stderr.strip() or out.stdout.strip()}\n{url}")
|
||||||
|
return json.loads(out.stdout)
|
||||||
|
|
||||||
|
|
||||||
|
def seconds(since):
|
||||||
|
return int(since[:-1]) * {"m": 60, "h": 3600, "d": 86400}[since[-1]]
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
if len(sys.argv) < 3:
|
||||||
|
sys.exit(__doc__)
|
||||||
|
kind, q = sys.argv[1], sys.argv[2]
|
||||||
|
since = sys.argv[3] if len(sys.argv) > 3 else "24h"
|
||||||
|
now = time.time()
|
||||||
|
start = now - seconds(since)
|
||||||
|
enc = urllib.parse.quote
|
||||||
|
if kind == "logs":
|
||||||
|
r = fetch(f"http://loki:3100/loki/api/v1/query_range?query={enc(q)}&limit=500"
|
||||||
|
f"&start={int(start * 1e9)}&end={int(now * 1e9)}&direction=backward")
|
||||||
|
lines = [(ts, line) for s in r["data"]["result"] for ts, line in s["values"]]
|
||||||
|
for ts, line in sorted(lines, reverse=True):
|
||||||
|
print(line)
|
||||||
|
print(f"-- {len(lines)} line(s){' (limit reached)' if len(lines) >= 500 else ''}", file=sys.stderr)
|
||||||
|
elif kind == "metric":
|
||||||
|
r = fetch(f"http://loki:3100/loki/api/v1/query?query={enc(q.replace('$range', since))}&time={int(now * 1e9)}")
|
||||||
|
rows = sorted(r["data"]["result"], key=lambda s: -float(s["value"][1]))
|
||||||
|
for s in rows:
|
||||||
|
labels = {k: v for k, v in s["metric"].items() if k not in ("container", "compose_project", "service_name")}
|
||||||
|
print(f"{s['value'][1]:>12} {json.dumps(labels) if labels else ''}")
|
||||||
|
print(f"-- {len(rows)} series", file=sys.stderr)
|
||||||
|
elif kind == "traces":
|
||||||
|
r = fetch(f"http://tempo:3200/api/search?q={enc(q)}&start={int(start)}&end={int(now)}&limit=100")
|
||||||
|
traces = sorted(r.get("traces", []), key=lambda t: -t.get("durationMs", 0))
|
||||||
|
for t in traces:
|
||||||
|
when = time.strftime("%m-%d %H:%M:%S", time.gmtime(int(t["startTimeUnixNano"]) / 1e9))
|
||||||
|
print(f"{t.get('durationMs', 0):>8}ms {when}Z {t['traceID']} {t.get('rootTraceName', '')}")
|
||||||
|
print(f"-- {len(traces)} trace(s)", file=sys.stderr)
|
||||||
|
elif kind == "trace":
|
||||||
|
r = fetch(f"http://tempo:3200/api/traces/{q}")
|
||||||
|
spans = [sp for b in r.get("batches", r.get("resourceSpans", []))
|
||||||
|
for ss in b.get("scopeSpans", b.get("instrumentationLibrarySpans", [])) for sp in ss["spans"]]
|
||||||
|
kids = {}
|
||||||
|
for sp in spans:
|
||||||
|
kids.setdefault(sp.get("parentSpanId", ""), []).append(sp)
|
||||||
|
|
||||||
|
def show(sp, depth):
|
||||||
|
ms = (int(sp["endTimeUnixNano"]) - int(sp["startTimeUnixNano"])) / 1e6
|
||||||
|
attrs = {a["key"]: next(iter(a["value"].values()), None) for a in sp.get("attributes", [])
|
||||||
|
if not a["key"].startswith(("code.", "thread.")) and a["key"] not in ("busy_ns", "idle_ns", "target")}
|
||||||
|
err = " ERROR" if sp.get("status", {}).get("code") in (2, "STATUS_CODE_ERROR") else ""
|
||||||
|
print(f"{' ' * depth}{sp['name']} {ms:.0f}ms{err} {json.dumps(attrs) if attrs else ''}")
|
||||||
|
for e in sp.get("events", []):
|
||||||
|
msg = next((a["value"].get("stringValue") for a in e.get("attributes", []) if a["key"] == "message"), e.get("name"))
|
||||||
|
# A scan logs a skip for every feed not due; they bury what happened.
|
||||||
|
if '"ev":"feed_skip"' in (msg or ""):
|
||||||
|
continue
|
||||||
|
print(f"{' ' * depth} - {msg}")
|
||||||
|
for k in sorted(kids.get(sp["spanId"], []), key=lambda s: int(s["startTimeUnixNano"])):
|
||||||
|
show(k, depth + 1)
|
||||||
|
|
||||||
|
ids = {sp["spanId"] for sp in spans}
|
||||||
|
for root in [sp for sp in spans if sp.get("parentSpanId", "") not in ids]:
|
||||||
|
show(root, 0)
|
||||||
|
else:
|
||||||
|
sys.exit(__doc__)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
Reference in New Issue
Block a user