Step 4: enclosure downloads
Streaming download with progress, content sniffing that replaces the long-dead detectFileType(), HTML-body rejection, and placement by rename from an incomplete dir on the same filesystem. Filters are applied once at discovery and recorded in enclosures.state, so the download queue is the table rather than the parse result and max_new_per_check defers work instead of dropping it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RPyeapneuXrCdojsaiXGbe
This commit is contained in:
36
PROGRESS.md
36
PROGRESS.md
@@ -10,9 +10,7 @@ The full design and step list live in the plan file at
|
||||
README, this file.
|
||||
- [x] **2. `config.rs` + `db.rs`** — TOML config structs + SQLite schema.
|
||||
- [x] **3. `feed.rs`** — conditional GET, RSS-then-Atom parse, persist entries.
|
||||
- [ ] **4. `download.rs`** — filename derivation + sanitizer, streaming download, `infer` sniff,
|
||||
HTML rejection, move into place, dedupe, keyword/explicit filters, `max_new_per_check`.
|
||||
*Done when:* smoke 1, 2, 4 pass.
|
||||
- [x] **4. `download.rs`** — downloads, filters, dedupe.
|
||||
- [ ] **5. `retention.rs`** — oldest-first quota + age reaper, `ipx reap [--dry-run]`.
|
||||
*Done when:* smoke 6 passes.
|
||||
- [ ] **6. `ipc.rs` + daemon** — broadcast event bus, UDS JSON-lines server, TTL scheduler,
|
||||
@@ -33,6 +31,38 @@ The full design and step list live in the plan file at
|
||||
|
||||
---
|
||||
|
||||
## 2026-09-09 — Step 4: download.rs
|
||||
|
||||
`src/download.rs`: streaming download to `<download_dir>/.ipx-incomplete/` (same filesystem as the
|
||||
destination, so filing it is a rename, not the original's copy-then-unlink), content sniffing,
|
||||
then `place()`. Filename comes from the URL's last path segment, percent-decoded, unless
|
||||
`Content-Disposition` names one (RFC 5987 `filename*=` preferred). The sanitizer keeps UTF-8 —
|
||||
`latin1_to_ascii` existed because 2004 filesystems demanded ASCII — strips the same characters
|
||||
`stringCleaning()` did plus control chars, and adds a real 255-byte cap the Python never had,
|
||||
preserving the extension across truncation.
|
||||
|
||||
Sniffing replaces `detectFileType()`, which called a `typeFile` module that was already missing in
|
||||
2008 and so always answered `'data'`. Two rules survive: an HTML body is a failed download (login
|
||||
wall/error page), and a torrent body is a torrent whatever the MIME claimed.
|
||||
|
||||
Filters run once at discovery and are recorded in `enclosures.state`; the download queue is then
|
||||
just "everything still `pending`", so an enclosure held back by `max_new_per_check` is picked up by
|
||||
the next scan instead of being lost. Keywords are OR'd across keywords and AND'd within one — the
|
||||
original's nested loop let a later keyword silently undo an earlier miss.
|
||||
|
||||
Verified: `cargo test` 16/16. Smoke against a local server, five enclosures, each filter path hit:
|
||||
`ep1.mp3 -> done`, `ep2.mp3 -> skipped (explicit)`, `ep2.mp3?v=3 -> skipped (no keyword match)`,
|
||||
`paywall.html -> error (HTML page, not media)`, `ep5.torrent -> torrent (deferred to step 7)`.
|
||||
Smoke 2 and 4 pass, and because a plain rerun 304s before parsing, dedupe was proved separately by
|
||||
clearing the stored etag/last-modified and re-parsing all five entries: 0 downloaded, hand-deleted
|
||||
file not refetched, `.ipx-incomplete` left empty.
|
||||
|
||||
Known wart: `ipx list` counts `path IS NOT NULL`, so a hand-deleted file still reads as downloaded.
|
||||
Reconciling rows against the filesystem belongs in step 5.
|
||||
Next: step 5 — `retention.rs`.
|
||||
|
||||
---
|
||||
|
||||
## 2026-09-09 — Step 3: feed.rs
|
||||
|
||||
`src/feed.rs`: conditional GET (`If-None-Match` + `If-Modified-Since`, optional basic auth) and a
|
||||
|
||||
Reference in New Issue
Block a user