Back off a failing feed exponentially, up to a day (#99)

A feed that failed was tried again on its usual schedule however long it had been failing:
gizmodo's 404, pelgrane's 403, daily-quests' 503 and toddstashwick's redirect loop every hour,
each a request to a site that had said no, a warning and scan time. A failing feed now waits as
long as it has been failing, from error_since to its last check, never less than its usual
interval and never more than a day: 1h, 1h, 2h, 4h, 8h, 16h, then daily on an hourly schedule.
No new column: error_since already marks the run's start and the first success clears it. A
forced refresh skips the due check, so it still tries at once. The feed list's next check
follows the backoff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-29 19:33:57 +00:00
parent 527777efdf
commit db87bc0f42
5 changed files with 50 additions and 7 deletions

View File

@@ -981,7 +981,7 @@ async fn fetch(ctx: &Arc<Ctx>, only: Option<&str>, force: bool, scope: &[String]
}
if !force && let Some(last) = state.last_checked {
let due = last + due_after(&cfg, feed_cfg, state.ttl_mins) as i64;
let due = last + due_after(&cfg, feed_cfg, state.ttl_mins, state.error_since, last) as i64;
if due > db::now() {
ctx.out.emit(Event::FeedSkip {
feed: id.clone(),
@@ -1165,13 +1165,34 @@ async fn retire_stranded(ctx: &Ctx) -> Result<usize> {
///
/// A per-feed schedule is an explicit instruction and wins outright. Without one, the
/// global schedule applies, but the feed's own <ttl> raises it when the publisher asks to
/// be polled less often.
pub fn due_after(cfg: &config::Config, feed: &config::Feed, ttl_mins: Option<u64>) -> u64 {
/// be polled less often. A feed that is failing backs off (`backoff`), from `error_since`, when
/// its run of failures began, to `last_checked`, when it last failed.
pub fn due_after(
cfg: &config::Config,
feed: &config::Feed,
ttl_mins: Option<u64>,
error_since: Option<i64>,
last_checked: i64,
) -> u64 {
let mins = match feed.schedule.as_deref().and_then(config::parse_interval) {
Some(explicit) => explicit,
None => ttl_mins.unwrap_or(0).max(cfg.general.interval()),
};
mins * 60
let failing_for = error_since.map(|since| (last_checked - since).max(0) as u64);
backoff(mins * 60, failing_for)
}
/// A failing feed waits as long as it has been failing, so the wait doubles with each failure
/// (1h, 1h, 2h, 4h, ... on an hourly schedule), never less than its usual interval and, past that,
/// never more than a day. Retried hourly, a feed dead for good cost a request, a warning and scan
/// time every hour (#99). The first success clears error_since, and with it the backoff; a
/// refresh someone asks for is not held back by it.
fn backoff(usual: u64, failing_for: Option<u64>) -> u64 {
const CEILING: u64 = 86_400;
match failing_for {
Some(f) => usual.max(f.min(CEILING)),
None => usual,
}
}
#[derive(Default)]
@@ -1836,6 +1857,22 @@ fn duration(secs: u64) -> String {
mod tests {
use super::*;
#[test]
fn a_failing_feed_backs_off_doubling_up_to_a_day() {
let hour = 3600;
assert_eq!(backoff(hour, None), hour);
// Failing since 0: each check waits as long as the failure has lasted so far.
let (mut at, mut waits) = (0, vec![]);
for _ in 0..8 {
let w = backoff(hour, Some(at));
waits.push(w / hour);
at += w;
}
assert_eq!(waits, [1, 1, 2, 4, 8, 16, 24, 24]);
// A weekly schedule is longer than the ceiling and stays as it is.
assert_eq!(backoff(7 * 86_400, Some(30 * 86_400)), 7 * 86_400);
}
#[tokio::test]
async fn the_first_start_moves_the_configuration_in_and_trims_the_file() {
let dir = std::env::temp_dir().join(format!("ipx-assemble-{}", std::process::id()));