Trace ids, failure kinds and one line per event in the JSON log (#91)
From Dash0's structured logging guide, what applies here: - Each JSON line inside a traced span ends with its trace_id and span_id, so a line in Loki leads to its trace in Tempo; the access log is written inside its request's span so it has one too. The JSON formatter takes no extra fields, so WithTrace appends them to the object it writes. - A feed or download failure carries error.type (the HTTP status, or dns, redirect_loop, timeout, ...) and http.response.status_code, from failure_kind beside explain_failure, so failures group by kind without a regex over msg. - Each event was logged twice: words under ipx::scan and fields under ipx::io. It is now one line under ipx::scan with both; the wire copy is at debug, for the admin page's Daemon I/O tab, and out of production's log. The healthcheck's status reply stays under ipx::io. - The access log's ms is duration_ms. The dashboard and the prod-check skill follow. - error fields are Display with the anyhow chain everywhere, not a mix of Debug and Display. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -24,12 +24,16 @@ $Q trace <trace id> # one trace as a tree, with its log lin
|
||||
|
||||
Every line has `timestamp`, `level`, `message` and `target`; lines inside a span have `span` (the
|
||||
innermost: `{"name":"feed","feed":"x"}`). Loki's `| json` flattens it to `span_name`, `span_feed`.
|
||||
Lines inside a traced span, requests and scans, also carry `trace_id` and `span_id`: give the
|
||||
`trace_id` to `$Q trace` to see the whole request or scan. (From 2026-09-29 16:30 UTC; before that,
|
||||
lines had no trace id, the access log's time was `ms`, and each event was logged twice, words under
|
||||
`ipx::scan` and fields under `ipx::io`.)
|
||||
|
||||
| target | fields | what |
|
||||
|---|---|---|
|
||||
| `ipx::http` | `method`, `path`, `route`, `status`, `ms` | one per web request; `route` is the pattern, empty for an unrouted path |
|
||||
| `ipx::io` | `ev` and the event's own: `feed`, `new`, `downloaded`, `failed`, `bytes`, `msg`, `url`, `feeds`, `pending`, `reason` | the daemon's events as they go on the wire |
|
||||
| `ipx::scan` | message only | the same events in words; warnings are feed and download failures |
|
||||
| `ipx::http` | `method`, `path`, `route`, `status`, `duration_ms` | one per web request; `route` is the pattern, empty for an unrouted path |
|
||||
| `ipx::scan` | `ev` and the event's own: `feed`, `new`, `downloaded`, `failed`, `bytes`, `msg`, `url`, `feeds`, `reason`; on a failure `error.type` and, from an HTTP error, `http.response.status_code` | the daemon's events, one line each, in words; warnings are feed and download failures |
|
||||
| `ipx::io` | `ev`, `feeds`, `pending`, `downloaded` on the `status` reply | commands arriving (`-> {...}`) and the healthcheck's answer |
|
||||
| `ipx` | message, sometimes fields | start-up, shutdown, account and config messages |
|
||||
|
||||
Events (`ev`): `feed_start`, `feed_done` (new, downloaded, failed, torrents), `feed_skip` (not due,
|
||||
@@ -37,7 +41,10 @@ routine), `feed_error` (msg), `download_done` (bytes), `download_error` (msg, ur
|
||||
`torrent_deferred`, `reaped`, `scan_done` (feeds checked), `reap_done`, `status` (feeds, pending,
|
||||
downloaded: the healthcheck's, every 30s), `error` (msg).
|
||||
|
||||
Filter on the text before `| json` where you can (`|= "\"target\":\"ipx::io\""`): it is much
|
||||
`error.type` is the HTTP status (`404`, `503`) or one of `dns`, `redirect_loop`, `timeout`, `tls`,
|
||||
`not_a_feed`, `site_message`, `connect`, `parse`, `other`; Loki's `| json` names it `error_type`.
|
||||
|
||||
Filter on the text before `| json` where you can (`|= "\"ev\":\"feed_error\""`): it is much
|
||||
cheaper than parsing every line. Lines before 2026-09-29 14:00 UTC are text, not JSON, and
|
||||
`| json | __error__=""` drops them.
|
||||
|
||||
@@ -53,8 +60,8 @@ a count says something happened, the lines and traces say why.
|
||||
Feed and download failures name the feed in the message; group them in the next check instead.
|
||||
2. **Failing feeds and downloads.**
|
||||
```
|
||||
$Q metric 'sum by (feed, msg) (count_over_time({container="iPX"} |= "\"target\":\"ipx::io\"" | json | __error__="" | ev="feed_error" [$range]))' 7d
|
||||
$Q metric 'sum by (feed, msg) (count_over_time({container="iPX"} |= "\"target\":\"ipx::io\"" | json | __error__="" | ev="download_error" [$range]))' 7d
|
||||
$Q metric 'sum by (feed, error_type) (count_over_time({container="iPX"} |= "\"ev\":\"feed_error\"" | json | __error__="" [$range]))' 7d
|
||||
$Q metric 'sum by (feed, error_type) (count_over_time({container="iPX"} |= "\"ev\":\"download_error\"" | json | __error__="" [$range]))' 7d
|
||||
```
|
||||
Tell the publisher's problems from ipx's. A 404, 410, DNS failure or 503 from the feed's own
|
||||
server is the publisher (worth saying, since the feed may have moved; one issue for a feed
|
||||
@@ -63,7 +70,7 @@ a count says something happened, the lines and traces say why.
|
||||
3. **Server errors and slow requests.**
|
||||
```
|
||||
$Q metric 'sum by (method, route, status) (count_over_time({container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | status >= 500 [$range]))'
|
||||
$Q metric 'topk(10, quantile_over_time(0.95, {container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | route != "" | route != "/api/events" | unwrap ms [$range]) by (method, route))'
|
||||
$Q metric 'topk(10, quantile_over_time(0.95, {container="iPX"} |= "\"target\":\"ipx::http\"" | json | __error__="" | route != "" | route != "/api/events" | unwrap duration_ms [$range]) by (method, route))'
|
||||
```
|
||||
Any 5xx is worth a look. 401s are people signing in, not a problem unless one address is
|
||||
hammering. For a slow route, find its traces (check 5) and see which span holds the time.
|
||||
|
||||
Reference in New Issue
Block a user