A download limit means a show's newest episodes, not a pace (#97)

pending() took the newest files still pending, up to the limit, so once a show's latest three
were down, each full read of its feed took the three before them, working back through its
whole history. In production 4420 files (about 310 GB) were queued this way across 12 shows,
all on the default limit of 3, which is meant as "the latest three". It now takes only from the
feed's newest `limit` items with a file. 0, unlimited, still takes the whole back catalogue:
that is how the shows kept as an archive are set, along with limits of 100 and 10000.

The settings' wording followed the old behaviour ("The rest wait for the next scan"); the field
is now "Newest episodes to download", and says what 0 does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-29 18:50:11 +00:00
parent 4a26c73b82
commit 527777efdf
6 changed files with 50 additions and 14 deletions

View File

@@ -48,6 +48,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed ### Fixed
- A limit on new downloads per check means a show's newest episodes: with it set to 3, ipx no longer works back through the show's history three at a time. Set to 0, it still takes the whole back catalogue.
- With no limit on new downloads per check (`max_new_per_check = 0`), downloads work on Postgres; they failed. - With no limit on new downloads per check (`max_new_per_check = 0`), downloads work on Postgres; they failed.
- Checking a feed with a long history is much quicker: items already stored are no longer written again on every check. - Checking a feed with a long history is much quicker: items already stored are no longer written again on every check.
- A scan no longer asks a feed's website for its icon every time; it asks again when the artwork changes or you refresh the feed. - A scan no longer asks a feed's website for its icon every time; it asks again when the artwork changes or you refresh the feed.

View File

@@ -17,8 +17,8 @@ behind iPodderX (2004-2008, Ray Slakinski & August Trometer).
sign in with a password or through a proxy (Cloudflare Zero Trust or Authentik), and admins sign in with a password or through a proxy (Cloudflare Zero Trust or Authentik), and admins
manage accounts and settings. manage accounts and settings.
- **Scanning.** Feeds are checked on a schedule, globally or per feed, and a feed's own TTL is - **Scanning.** Feeds are checked on a schedule, globally or per feed, and a feed's own TTL is
honoured. Keyword, explicit-content and media-type filters decide what is downloaded, with a cap honoured. Keyword, explicit-content and media-type filters decide what is downloaded, with a limit
on new downloads per scan. on how many of a feed's newest episodes are downloaded.
- **Downloads.** Files come over HTTP or BitTorrent and are filed into a folder per feed. - **Downloads.** Files come over HTTP or BitTorrent and are filed into a folder per feed.
Retention deletes the oldest files to stay under a disk quota or an age limit, and never touches Retention deletes the oldest files to stay under a disk quota or an age limit, and never touches
an item someone has pinned. an item someone has pinned.

View File

@@ -52,8 +52,10 @@ database as described above; `download_dir`, `socket` and `organize` stay in con
read it. `0` disables it entirely. read it. `0` disables it entirely.
* **`max_age_days`** — items older than this with no file on disk are pruned from the database. * **`max_age_days`** — items older than this with no file on disk are pruned from the database.
Kept ones stay. `0` disables it. Kept ones stay. `0` disables it.
* **`max_new_per_check`** — the cap that stops a new subscription pulling a whole back catalogue. * **`max_new_per_check`** — how many of a feed's newest episodes are downloaded; older ones stay
`0` means unlimited, which is rarely what you want: subscribing to an OPML of 80 feeds with no cap listed to download by hand. It stops a new subscription pulling a whole back catalogue.
`0` means every episode, for an archive; set it on the feeds you want archived, since on the
global default it applies to every feed: subscribing to an OPML of 80 feeds with no limit
fetched 216 files and 22 GB in one scan. fetched 216 files and 22 GB in one scan.
* **`media_types`** — top-level MIME types taken automatically. Anything else is still listed and * **`media_types`** — top-level MIME types taken automatically. Anything else is still listed and
can be fetched by hand; blog feeds put each article's header image in an `<enclosure>`, and can be fetched by hand; blog feeds put each article's header image in an `<enclosure>`, and

View File

@@ -623,16 +623,26 @@ pub struct Pending {
} }
impl Db { impl Db {
/// The download queue is the table, not the parse result: an enclosure held back by /// The download queue is the table, not the parse result: what a scan could not finish is
/// `max_new_per_check` is simply picked up by the next scan, in feed order. /// picked up by the next. Only from the feed's `limit` newest items with a file, though: a
/// limit of 3 means the three latest episodes. Taking the newest three still pending, each
/// full read of a feed took the three before the ones already downloaded, working back
/// through every show's history: 4420 files queued across 12 shows on the default of 3
/// (#97). Unlimited, 0 in the settings, is the whole back catalogue, for an archive. An item
/// whose files were all skipped by a filter does not hold one of the places.
#[tracing::instrument(skip_all)] #[tracing::instrument(skip_all)]
pub async fn pending(&self, feed_id: &str, limit: usize) -> Result<Vec<Pending>> { pub async fn pending(&self, feed_id: &str, limit: usize) -> Result<Vec<Pending>> {
// Newest first: a cap of 3 should mean the three latest episodes, not the three that
// happen to have been recorded first.
self.rows( self.rows(
"SELECT x.id, x.url, x.mime FROM enclosures x "SELECT x.id, x.url, x.mime FROM enclosures x
JOIN entries e ON e.feed_id = x.feed_id AND e.guid = x.guid JOIN entries e ON e.feed_id = x.feed_id AND e.guid = x.guid
WHERE x.feed_id = $1 AND x.state = 'pending' WHERE x.feed_id = $1 AND x.state = 'pending'
AND e.guid IN (SELECT n.guid FROM entries n
WHERE n.feed_id = $1
AND EXISTS (SELECT 1 FROM enclosures y
WHERE y.feed_id = n.feed_id AND y.guid = n.guid
AND y.state <> 'skipped')
ORDER BY coalesce(n.published, n.first_seen) DESC, n.guid DESC
LIMIT $2)
ORDER BY coalesce(e.published, e.first_seen) DESC, x.id DESC ORDER BY coalesce(e.published, e.first_seen) DESC, x.id DESC
LIMIT $2", LIMIT $2",
// Unlimited is usize::MAX, which `as i64` makes -1: SQLite reads LIMIT -1 as no // Unlimited is usize::MAX, which `as i64` makes -1: SQLite reads LIMIT -1 as no
@@ -1936,6 +1946,29 @@ mod tests {
assert_eq!((list["f"].unread, list["g"].unread), (1, 1)); // a read, c hidden; sam's read is sam's assert_eq!((list["f"].unread, list["g"].unread), (1, 1)); // a read, c hidden; sam's read is sam's
} }
#[tokio::test]
async fn a_limit_takes_the_newest_episodes_not_the_back_catalogue() {
let db = Db::memory().await.unwrap();
db.exec_for_test(
"INSERT INTO entries (feed_id, guid, first_seen, published) VALUES
('f','e1',0,100),('f','e2',0,200),('f','e3',0,300),('f','e4',0,400),('f','e5',0,500);
INSERT INTO enclosures (id, feed_id, guid, url, state, path) VALUES
(1,'f','e1','u1','pending',null),(2,'f','e2','u2','pending',null),
(3,'f','e3','u3','done','/tmp/3'),(4,'f','e4','u4','done','/tmp/4'),
(5,'f','e5','u5','skipped',null);",
).await
.unwrap();
let ids = |p: Vec<Pending>| p.into_iter().map(|p| p.id).collect::<Vec<_>>();
// The newest three with a file worth having are e4, e3 and e2 (e5's was skipped): only
// e2 is waiting. e1 is back catalogue.
assert_eq!(ids(db.pending("f", 3).await.unwrap()), vec![2]);
// Once e2 is down, nothing: before #97 the next scan took e1.
db.exec_for_test("UPDATE enclosures SET state = 'done', path = '/tmp/2' WHERE id = 2").await.unwrap();
assert!(db.pending("f", 3).await.unwrap().is_empty());
// Unlimited is the archive: everything still waiting.
assert_eq!(ids(db.pending("f", usize::MAX).await.unwrap()), vec![1]);
}
#[tokio::test] #[tokio::test]
async fn an_unlimited_queue_takes_everything_waiting() { async fn an_unlimited_queue_takes_everything_waiting() {
let db = Db::memory().await.unwrap(); let db = Db::memory().await.unwrap();

View File

@@ -34,11 +34,11 @@ async function drawServer(){
</div> </div>
<span class="hint">Applies to every feed that does not set its own. A feed's suggested <span class="hint">Applies to every feed that does not set its own. A feed's suggested
interval (its <b>ttl</b>) is still honoured when it asks to be polled less often.</span></div> interval (its <b>ttl</b>) is still honoured when it asks to be polled less often.</span></div>
<div class="field"><label>Max new downloads per scan, per feed</label> <div class="field"><label>Newest episodes to download, per feed</label>
<input type="number" id="gmax" min="0" max="999" value="${g.max_new_per_check}"> <input type="number" id="gmax" min="0" max="999" value="${g.max_new_per_check}">
<span class="hint">Applies to any feed that does not set its own — including every feed <span class="hint">Applies to any feed that does not set its own — including every feed
inside an OPML subscription. <b>0 means unlimited</b>, which will pull a whole back inside an OPML subscription. Older episodes stay listed to download by hand. <b>0 means
catalogue the first time a feed is scanned.</span></div> every episode</b>, the whole back catalogue.</span></div>
<div class="field"><label>Download these media types automatically</label> <div class="field"><label>Download these media types automatically</label>
<input type="text" id="gtypes" value="${esc((g.media_types||[]).join(', '))}" placeholder="audio, video"> <input type="text" id="gtypes" value="${esc((g.media_types||[]).join(', '))}" placeholder="audio, video">
<span class="hint">Anything else is still listed and can be downloaded by hand — blog feeds <span class="hint">Anything else is still listed and can be downloaded by hand — blog feeds

View File

@@ -275,10 +275,10 @@ function settingsModal(f, newUrl?: string){
<span class="hint">Comma separated words or phrases. An item with one in its title or text <span class="hint">Comma separated words or phrases. An item with one in its title or text
is hidden from you and not downloaded for you, as well as those your Settings hide is hidden from you and not downloaded for you, as well as those your Settings hide
everywhere.</span></div>`} everywhere.</span></div>`}
<div class="field"><label>Max new downloads per scan</label> <div class="field"><label>Newest episodes to download</label>
<input type="number" id="smax" min="0" value="${f.max_new_per_check??''}"> <input type="number" id="smax" min="0" value="${f.max_new_per_check??''}">
<span class="hint">Blank follows the global default (${globalMax}). The rest wait for <span class="hint">Blank follows the global default (${globalMax}). Older episodes stay
the next scan.</span></div> listed to download by hand. <b>0 means every episode</b>, the whole back catalogue.</span></div>
<label class="check"><input type="checkbox" id="sauto" ${f.auto_download?'checked':''}> Download new items automatically</label> <label class="check"><input type="checkbox" id="sauto" ${f.auto_download?'checked':''}> Download new items automatically</label>
<label class="check"><input type="checkbox" id="sexp" ${f.allow_explicit?'checked':''}> Allow items marked explicit</label> <label class="check"><input type="checkbox" id="sexp" ${f.allow_explicit?'checked':''}> Allow items marked explicit</label>
<div class="field"><label>Feed URL</label> <div class="field"><label>Feed URL</label>