Step 4: enclosure downloads

Streaming download with progress, content sniffing that replaces the
long-dead detectFileType(), HTML-body rejection, and placement by rename
from an incomplete dir on the same filesystem.

Filters are applied once at discovery and recorded in enclosures.state,
so the download queue is the table rather than the parse result and
max_new_per_check defers work instead of dropping it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPyeapneuXrCdojsaiXGbe
This commit is contained in:
2026-09-09 20:38:30 +00:00
parent ab2583fc64
commit a8e1f474a6
6 changed files with 540 additions and 14 deletions

View File

@@ -10,9 +10,7 @@ The full design and step list live in the plan file at
README, this file.
- [x] **2. `config.rs` + `db.rs`** — TOML config structs + SQLite schema.
- [x] **3. `feed.rs`** — conditional GET, RSS-then-Atom parse, persist entries.
- [ ] **4. `download.rs`** filename derivation + sanitizer, streaming download, `infer` sniff,
HTML rejection, move into place, dedupe, keyword/explicit filters, `max_new_per_check`.
*Done when:* smoke 1, 2, 4 pass.
- [x] **4. `download.rs`** — downloads, filters, dedupe.
- [ ] **5. `retention.rs`** — oldest-first quota + age reaper, `ipx reap [--dry-run]`.
*Done when:* smoke 6 passes.
- [ ] **6. `ipc.rs` + daemon** — broadcast event bus, UDS JSON-lines server, TTL scheduler,
@@ -33,6 +31,38 @@ The full design and step list live in the plan file at
---
## 2026-09-09 — Step 4: download.rs
`src/download.rs`: streaming download to `<download_dir>/.ipx-incomplete/` (same filesystem as the
destination, so filing it is a rename, not the original's copy-then-unlink), content sniffing,
then `place()`. Filename comes from the URL's last path segment, percent-decoded, unless
`Content-Disposition` names one (RFC 5987 `filename*=` preferred). The sanitizer keeps UTF-8 —
`latin1_to_ascii` existed because 2004 filesystems demanded ASCII — strips the same characters
`stringCleaning()` did plus control chars, and adds a real 255-byte cap the Python never had,
preserving the extension across truncation.
Sniffing replaces `detectFileType()`, which called a `typeFile` module that was already missing in
2008 and so always answered `'data'`. Two rules survive: an HTML body is a failed download (login
wall/error page), and a torrent body is a torrent whatever the MIME claimed.
Filters run once at discovery and are recorded in `enclosures.state`; the download queue is then
just "everything still `pending`", so an enclosure held back by `max_new_per_check` is picked up by
the next scan instead of being lost. Keywords are OR'd across keywords and AND'd within one — the
original's nested loop let a later keyword silently undo an earlier miss.
Verified: `cargo test` 16/16. Smoke against a local server, five enclosures, each filter path hit:
`ep1.mp3 -> done`, `ep2.mp3 -> skipped (explicit)`, `ep2.mp3?v=3 -> skipped (no keyword match)`,
`paywall.html -> error (HTML page, not media)`, `ep5.torrent -> torrent (deferred to step 7)`.
Smoke 2 and 4 pass, and because a plain rerun 304s before parsing, dedupe was proved separately by
clearing the stored etag/last-modified and re-parsing all five entries: 0 downloaded, hand-deleted
file not refetched, `.ipx-incomplete` left empty.
Known wart: `ipx list` counts `path IS NOT NULL`, so a hand-deleted file still reads as downloaded.
Reconciling rows against the filesystem belongs in step 5.
Next: step 5 — `retention.rs`.
---
## 2026-09-09 — Step 3: feed.rs
`src/feed.rs`: conditional GET (`If-None-Match` + `If-Modified-Since`, optional basic auth) and a