Back off a failing feed exponentially, up to a day (#99)

A feed that failed was tried again on its usual schedule however long it had been failing:
gizmodo's 404, pelgrane's 403, daily-quests' 503 and toddstashwick's redirect loop every hour,
each a request to a site that had said no, a warning and scan time. A failing feed now waits as
long as it has been failing, from error_since to its last check, never less than its usual
interval and never more than a day: 1h, 1h, 2h, 4h, 8h, 16h, then daily on an hourly schedule.
No new column: error_since already marks the run's start and the first success clears it. A
forced refresh skips the due check, so it still tries at once. The feed list's next check
follows the backoff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-29 19:33:57 +00:00
parent 527777efdf
commit db87bc0f42
5 changed files with 50 additions and 7 deletions

View File

@@ -23,6 +23,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Changed
- A feed that keeps failing is checked less and less often, waiting as long as it has been failing, up to once a day; it goes back to its schedule as soon as it works. Refreshing it still checks it at once.
- A scan's trace shows the time a feed spends on its artwork and in the database after the fetch.
- The JSON log carries each line's `trace_id` and `span_id`, logs each scan event once instead of
twice, names a failure's kind in `error.type` (and its HTTP status in

View File

@@ -45,7 +45,9 @@ media_types = ["audio", "video"]
database as described above; `download_dir`, `socket` and `organize` stay in config.toml.
* **`schedule`** — how often feeds are re-checked. A feed's own `<ttl>` still wins when it asks to
be polled *less* often, and a per-feed `schedule` overrides both. Admin-only from the UI.
be polled *less* often, and a per-feed `schedule` overrides both. A feed that keeps failing
waits as long as it has been failing before the next try, up to a day, and is back on schedule
after its first success. Admin-only from the UI.
* **`organize`** — `feed` files downloads under the feed's folder; `date` under `YYYY-MM-DD`.
* **`max_total_gb`** — the reaper deletes to get back under this, oldest first, keeping a 50 MB
pad. Kept items are never deleted, and a file only counts as read once every subscriber has

View File

@@ -415,6 +415,8 @@ pub struct HttpState {
pub last_modified: Option<String>,
pub last_checked: Option<i64>,
pub ttl_mins: Option<u64>,
/// When the feed's current run of failures began, for backing off (`due_after`).
pub error_since: Option<i64>,
}
impl Db {
@@ -427,6 +429,7 @@ impl Db {
last_modified: f.last_modified,
last_checked: f.last_checked,
ttl_mins: f.ttl_mins.map(|t| t.max(0) as u64),
error_since: f.error_since,
})
.unwrap_or_default())
}

View File

@@ -981,7 +981,7 @@ async fn fetch(ctx: &Arc<Ctx>, only: Option<&str>, force: bool, scope: &[String]
}
if !force && let Some(last) = state.last_checked {
let due = last + due_after(&cfg, feed_cfg, state.ttl_mins) as i64;
let due = last + due_after(&cfg, feed_cfg, state.ttl_mins, state.error_since, last) as i64;
if due > db::now() {
ctx.out.emit(Event::FeedSkip {
feed: id.clone(),
@@ -1165,13 +1165,34 @@ async fn retire_stranded(ctx: &Ctx) -> Result<usize> {
///
/// A per-feed schedule is an explicit instruction and wins outright. Without one, the
/// global schedule applies, but the feed's own <ttl> raises it when the publisher asks to
/// be polled less often.
pub fn due_after(cfg: &config::Config, feed: &config::Feed, ttl_mins: Option<u64>) -> u64 {
/// be polled less often. A feed that is failing backs off (`backoff`), from `error_since`, when
/// its run of failures began, to `last_checked`, when it last failed.
pub fn due_after(
cfg: &config::Config,
feed: &config::Feed,
ttl_mins: Option<u64>,
error_since: Option<i64>,
last_checked: i64,
) -> u64 {
let mins = match feed.schedule.as_deref().and_then(config::parse_interval) {
Some(explicit) => explicit,
None => ttl_mins.unwrap_or(0).max(cfg.general.interval()),
};
mins * 60
let failing_for = error_since.map(|since| (last_checked - since).max(0) as u64);
backoff(mins * 60, failing_for)
}
/// A failing feed waits as long as it has been failing, so the wait doubles with each failure
/// (1h, 1h, 2h, 4h, ... on an hourly schedule), never less than its usual interval and, past that,
/// never more than a day. Retried hourly, a feed dead for good cost a request, a warning and scan
/// time every hour (#99). The first success clears error_since, and with it the backoff; a
/// refresh someone asks for is not held back by it.
fn backoff(usual: u64, failing_for: Option<u64>) -> u64 {
const CEILING: u64 = 86_400;
match failing_for {
Some(f) => usual.max(f.min(CEILING)),
None => usual,
}
}
#[derive(Default)]
@@ -1836,6 +1857,22 @@ fn duration(secs: u64) -> String {
mod tests {
use super::*;
#[test]
fn a_failing_feed_backs_off_doubling_up_to_a_day() {
let hour = 3600;
assert_eq!(backoff(hour, None), hour);
// Failing since 0: each check waits as long as the failure has lasted so far.
let (mut at, mut waits) = (0, vec![]);
for _ in 0..8 {
let w = backoff(hour, Some(at));
waits.push(w / hour);
at += w;
}
assert_eq!(waits, [1, 1, 2, 4, 8, 16, 24, 24]);
// A weekly schedule is longer than the ceiling and stays as it is.
assert_eq!(backoff(7 * 86_400, Some(30 * 86_400)), 7 * 86_400);
}
#[tokio::test]
async fn the_first_start_moves_the_configuration_in_and_trims_the_file() {
let dir = std::env::temp_dir().join(format!("ipx-assemble-{}", std::process::id()));

View File

@@ -790,11 +790,11 @@ async fn feeds(
.schedule
.as_deref()
.and_then(crate::config::parse_interval),
every_mins: crate::due_after(&cfg, feed, ttl_mins) / 60,
every_mins: crate::due_after(&cfg, feed, ttl_mins, None, 0) / 60,
last_checked: s.last_checked,
next_check: s
.last_checked
.map(|t| t + crate::due_after(&cfg, feed, ttl_mins) as i64),
.map(|t| t + crate::due_after(&cfg, feed, ttl_mins, s.error_since, t) as i64),
failing: s
.error_since
.filter(|since| crate::db::now() - since >= FLAG_AFTER_SECS)