The daemon times out fetching some feeds that answer at once from anywhere else #112

Open
opened 2026-10-02 12:36:32 -07:00 by rays · 0 comments
Owner

Since the 30 s feed timeout (#108) went out on 2026-10-02, matthew-garrett (https://mjg59.dreamwidth.org/data/rss) times out on every try in the daemon: feed_error error_type=timeout at 13:48, 16:56, 17:57, 18:59 UTC, each fetch span exactly 30.0 s (trace 4d7448382e236cee5598e921f8bfbad7). In the first scan after the deploy it took 18.6 s and succeeded (d56e4df581f512bbc0115026e484ba54); pluralistic.net took the full 30 s in that scan and 0.4 s in a later one.

The same requests are fast everywhere else, tried at 19:30:

  • curl from Tower: 0.3 s (pluralistic 1.2 s), any of plain, compressed, HTTP/1.1, HTTP/2.
  • wget from a container on content_default, ipx's network: fine.
  • ipx's own client in a fresh process inside the iPX container (ipx --local add, throwaway SQLite): 1-2 s, three times.
  • curl replaying the daemon's conditional request (If-Modified-Since: Mon, 06 Jul 2026 09:03:04 GMT, no ETag): 304 in 0.14 s.

So it is the long-running daemon, not the site, the network or the request. Before the timeout the same feed took 60-67 s and ended in 504s (#108), so it was this all along rather than Dreamwidth being down. Unconfirmed guess: a pooled connection (reqwest keeps them, HTTP/2 multiplexed) that has gone stale to some hosts; the prefetch tasks (#104) share the one client. Next step: log or trace connection reuse for the fetch, or try pool_idle_timeout / http1-only for feed fetches and watch whether the timeouts stop.

Query: {container="iPX"} |= ""ev":"feed_error"" | json | error_type="timeout"

Since the 30 s feed timeout (#108) went out on 2026-10-02, matthew-garrett (https://mjg59.dreamwidth.org/data/rss) times out on every try in the daemon: feed_error error_type=timeout at 13:48, 16:56, 17:57, 18:59 UTC, each fetch span exactly 30.0 s (trace 4d7448382e236cee5598e921f8bfbad7). In the first scan after the deploy it took 18.6 s and succeeded (d56e4df581f512bbc0115026e484ba54); pluralistic.net took the full 30 s in that scan and 0.4 s in a later one. The same requests are fast everywhere else, tried at 19:30: - curl from Tower: 0.3 s (pluralistic 1.2 s), any of plain, compressed, HTTP/1.1, HTTP/2. - wget from a container on content_default, ipx's network: fine. - ipx's own client in a fresh process inside the iPX container (ipx --local add, throwaway SQLite): 1-2 s, three times. - curl replaying the daemon's conditional request (If-Modified-Since: Mon, 06 Jul 2026 09:03:04 GMT, no ETag): 304 in 0.14 s. So it is the long-running daemon, not the site, the network or the request. Before the timeout the same feed took 60-67 s and ended in 504s (#108), so it was this all along rather than Dreamwidth being down. Unconfirmed guess: a pooled connection (reqwest keeps them, HTTP/2 multiplexed) that has gone stale to some hosts; the prefetch tasks (#104) share the one client. Next step: log or trace connection reuse for the fetch, or try pool_idle_timeout / http1-only for feed fetches and watch whether the timeouts stop. Query: {container="iPX"} |= "\"ev\":\"feed_error\"" | json | error_type="timeout"
rays added the bug label 2026-10-02 12:36:32 -07:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: rays/ipx#112