Commit Graph

77 Commits

Author SHA1 Message Date
e7c59489ee Announce a followed move as a feed_moved event (#115)
A feed moved to its new address was only logged by follow_move. It is now an event, feed_moved
with the feed and its old and new addresses, so it goes where every other event goes: the log,
with from and to as fields, the admin page's Scans view, `ipx fetch`, and the page, which gets
the feed's new row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 20:22:28 +00:00
f7f4b466dc Follow a feed that has moved for good to its new address (#115)
A feed whose address answered with a permanent redirect was read through it on every check,
and the catalogue kept the old address: 28 of 147 feeds in production, most http to https, some
to a new path or domain.

Feeds are now fetched with a client of their own that follows no redirects (Ctx::feed_client),
and feed::fetch follows them itself, up to 10 hops, so it sees each one. When every hop was
permanent (301 or 308) it says where the feed ended up, and the scan moves the feed there in the
catalogue (follow_move). A temporary hop (302, 307) anywhere moves nothing. A feed an OPML lists
is left alone, as the OPML would put the old address back, and so is a move onto an address
another feed has. A password goes only to the feed's own host, never to a redirect elsewhere;
reqwest's own following dropped it the same way. Ten hops is a loop, worded as reqwest worded
it so it still reads as redirect_loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 20:05:43 +00:00
ec8fd5dd86 Sleep until the next feed is due instead of scanning every minute (#114)
The daemon ticked every 60 s and ran a scan pass each time: the sweep, then a check-state query
per feed (about 180) to find which were due. In the six hours before, 293 of 362 passes found
nothing due. Now, after each pass, it works out when the earliest feed is due (due_at, shared
with the scan's own check, over Db::http_states, one query) and sleeps until then: at least
30 s, so a feed that never gets a check time cannot spin it, and at most 10 minutes, so what no
command announces, ipx add or a shorter schedule, is picked up. Commands still wake it at once,
and the first pass after starting runs straight away, as the tick's did. The scan reads every
feed's state in one query too.

The prod-check skill says what to expect now: tens of scans in six hours, and pending as the
real queue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 19:56:13 +00:00
a387012c69 Count as queued only what a scan will download on its own (#113)
Every file in state 'pending' counted as waiting to download: 4420 in production, across 12
shows. Since #97 a scan only downloads among a feed's newest max_new_per_check items, so those
were back-catalogue episodes no scan would take; the real queue was 0.

A new state, 'held': listed and downloadable by hand, but outside the feed's newest items, or of
a feed nothing downloads automatically. Db::hold_back moves a feed's waiting files between
'pending' and 'held' each time the feed is due, changed or not, and again after its items are
stored, so a new episode, a raised limit or auto-download turned on or off moves them. A held
file keeps its item's place among the newest, as a downloaded one does. 'pending' now means
queued, so ipx status, /api/status and the dashboard read true without changing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 19:47:52 +00:00
a00516687a Keep artwork on disk and serve every image from iPX (#111)
The page loaded artwork from each publisher's server, or through /api/art, fetched every time,
for http-only hosts. Nothing was kept, every visit asked every publisher, and artwork went when
a publisher's server did.

- The page draws every image a feed or item names from /api/art. The first time, iPX fetches
  it (only an address a feed or item names, only an image, up to 5 MB, within the feed timeout)
  and keeps it in art/ beside the database, under a hash of its address with its type beside
  it (src/art.rs). Later it comes from disk, which marks it as used.
- art_cache_mb, a server setting on the admin page, 500 by default, caps what is kept: the
  sweep before each scan drops the least recently shown until it fits. 0 keeps nothing, and
  artwork is fetched through iPX each time. Settings saved before it get the default.
- Served from iPX's own address, someone else's image must stay an image: nosniff, and a CSP
  with sandbox, so an SVG opened on its own runs no script as iPX.
- The fixture server sends .jpg as image/jpeg, which /api/art requires.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 13:50:46 +00:00
38ebc7b02b Move old items' http artwork to https too, for hosts the feed no longer names (#110)
The first pass only asked the hosts the feed's current body names. IGN's feed kept five 2009
items with artwork on assets1/assets2.ignimgs.com, which its feed no longer mentions, so they
stayed on http. A scan now also asks, once per host, the hosts of the feed's stored http
artwork (Db::http_images).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 13:35:51 +00:00
54d827f655 Time out a hung feed, serve the precomposed touch icon, and store http artwork on https (#108, #109, #110)
#108: the HTTP client had no timeout, and scans handle feeds in order, so a hung server held
every scan. Dreamwidth answered 504 after 60-67 s for a day and each scan took 70-80 s instead of
15. A feed fetch, and a Patreon creator's show list, now gets 30 s from connecting to the last
byte (feed::FEED_TIMEOUT); the client gets a 10 s connect timeout, which bounds a download's
start but not a long download.

#109: iOS asks for /apple-touch-icon-precomposed.png first when the site is added to a home
screen; it was a 404 and the only non-feed warning in the log. It serves the same icon.

#110: the page is https and loads no http. Artwork on http came through /api/art (#90) even
when its host serves https too. A scan now tries each http artwork host on https once per feed
(feed::prefer_https) and stores the https address where the host answers with an image,
rewriting that feed's stored items from the same host (Db::secure_images). 4 of the 5 hosts in
production do; cdn.thesecretcabal.com presents another name's certificate and stays on
/api/art.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 13:31:50 +00:00
d52ecfcbd7 A listed feed nobody subscribes to downloads nothing (#107)
With no subscribers a feed falls back to its own settings, where auto_download is on, so a
listed feed with audio would have downloaded files for no one. The seeded news feeds carry
only images, which media_types already skips, so nothing was downloaded. Files skipped for it
are judged again, by the new subscriber's settings, on the next scan after someone subscribes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 21:20:09 +00:00
142610e8b8 ipx add --list or --category updates a feed the catalogue already has (#107)
It refused one already there ("already subscribed as ..."), so a feed added before listing
existed, such as CBC's, could not be put in the Directory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 21:17:49 +00:00
43acc62259 List feeds in the Directory before anyone subscribes, and clear out dead ones (#107)
The Directory showed a catalogue feed only once someone subscribed, and a feed left the
catalogue with its last subscriber, so nothing could be put there for others to find.

- A feed has a listed flag, set by ipx add --list (with --category for the Directory's chip).
  The web page keeps a listed feed in the catalogue when its last subscriber leaves.
- The Directory lists every catalogue feed; Popular still only what people subscribe to.
  popular() reads titles, artwork and categories through Db::feed_list, not three queries a
  feed. Subscribing from the Directory scans the feed at once.
- A feed nobody subscribes to is checked once a day at most.
- clean_directory, in the sweep before each scan, removes from the catalogue and the database
  a feed nobody subscribes to, with no file on disk and not from an OPML, that has failed for
  30 days or published nothing in a year. Run against production first: it removes nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 21:13:13 +00:00
ef5bdb4244 Say what guards the web UI, and warn only without Cloudflare Access (#106)
Every start logged WARN "web ui is reachable off this machine; the token is all that guards
it". A container has to bind 0.0.0.0 for its port to be published, so it fired on every start
of production, and it was out of date: signing in takes an account's password or the admin
token, and through the tunnel Cloudflare Access. It was the only warning in a healthy log. Now
it names what guards it, at info when Access is configured and a warning otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:46:43 +00:00
7a134b810f Fetch feeds six ahead in a scan; reload the list only after a scan that checked something (#103, #104)
Fetching was 65-90% of a scan, each feed waiting for the one before: 13 s of fetches in a 20 s
refresh of 32 feeds. The scan now works out which feeds are due, fetches their bodies up to six
ahead in tasks of their own, and handles each in order as before, so database writes,
downloads and OPML syncs stay one at a time. A Patreon creator still fetches in scan_one.

The page reloaded /api/feeds, and /api/settings with it, on every scan_done: the scheduler
scans every minute, so each open page reloaded the list once a minute, 169 times an hour. It
now reloads only when the scan checked a feed, and asks for settings once, on first load.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:27:51 +00:00
79bc006a30 Add only what is a feed or links one; refuse the rest (#102)
Adding an address looked for the feed a web page links and, finding none, added the address
as it was: every check then failed, and the sidebar called it a feed that had moved. cnn.com
is one; its page links no feed. find_feed replaces feed_behind_page: the address is added if
it is a feed or an OPML list, the feed its page links if it is a web page that links one (and
that is a feed), and otherwise the add is refused with the reason, from the web page (400, the
dialog stays open) and from `ipx add`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:06:04 +00:00
db87bc0f42 Back off a failing feed exponentially, up to a day (#99)
A feed that failed was tried again on its usual schedule however long it had been failing:
gizmodo's 404, pelgrane's 403, daily-quests' 503 and toddstashwick's redirect loop every hour,
each a request to a site that had said no, a warning and scan time. A failing feed now waits as
long as it has been failing, from error_since to its last check, never less than its usual
interval and never more than a day: 1h, 1h, 2h, 4h, 8h, 16h, then daily on an hourly schedule.
No new column: error_since already marks the run's start and the first success clears it. A
forced refresh skips the due check, so it still tries at once. The feed list's next check
follows the backoff.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 19:33:57 +00:00
57ab419718 Insert only a feed's new items and files on a scan (#96)
The spans added in 2799704 showed it: in a full scan of 134 feeds (trace da9a419b...,
2026-09-29 17:31, 315 s), storing items took 117 s, fetching 44 s and every other database call
about 2 s together. A scan inserted every item and file the feed listed, stored or not, one
round trip of about 10 ms each; Clarkesworld's 1200 items took 13 s. It now reads the feed's
stored guids and file URLs once (Db::stored_items) and inserts only the rest. A file URL not
among the feed's own may still be another feed's, so that one still goes to the insert, which
finds it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 17:40:52 +00:00
2799704d30 Look at a feed's artwork only when it may have changed, and trace a scan's database work (#95, #96)
A feed read in full checked its own artwork and, without one, asked its website for an icon,
every time; a feed without validators is read in full every scan, so looking-for-group spent
2 s of every scan loading lfg.co's home page. Now the check runs when the feed names different
artwork from what is stored, or the scan was asked for, which keeps #80's point: a refresh
still picks up an icon the site changes or fixes.

Feed spans ran seconds past their fetch with nothing to say where (#96). The artwork lookup,
the loop that stores each item, and the per-feed database calls (feed_summary, record_feed,
subscribers, adopt, skipped_by_filter, rehide, pending) now have spans of their own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 17:27:31 +00:00
448e557272 Trace ids, failure kinds and one line per event in the JSON log (#91)
From Dash0's structured logging guide, what applies here:

- Each JSON line inside a traced span ends with its trace_id and span_id, so a line in Loki leads
  to its trace in Tempo; the access log is written inside its request's span so it has one too.
  The JSON formatter takes no extra fields, so WithTrace appends them to the object it writes.
- A feed or download failure carries error.type (the HTTP status, or dns, redirect_loop,
  timeout, ...) and http.response.status_code, from failure_kind beside explain_failure, so
  failures group by kind without a regex over msg.
- Each event was logged twice: words under ipx::scan and fields under ipx::io. It is now one
  line under ipx::scan with both; the wire copy is at debug, for the admin page's Daemon I/O tab,
  and out of production's log. The healthcheck's status reply stays under ipx::io.
- The access log's ms is duration_ms. The dashboard and the prod-check skill follow.
- error fields are Display with the anyhow chain everywhere, not a mix of Debug and Display.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 16:25:53 +00:00
928cbaf8f4 Show each feed's id labelled in ipx list (#82)
The id led the title's line unlabelled, so antirez.com's, "feed" with no title beside it, read
as a heading; 'ipx fetch antirez' was tried instead and failed. The title now heads the entry and
the id has its own row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 16:16:25 +00:00
0945f3a9f4 Remove ipx copy-db (#85)
It was the one-off copy from SQLite to Postgres (#18), run once on 2026-09-18. Production has run
on Postgres since; rolling back needs only the old state.db, which is kept, not this command.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 16:15:51 +00:00
d4e00be085 A feed's own artwork has to be there before it is used (#89)
A feed's itunes:image or <image><url> was stored without being asked
for, so a dead one stood in the way of the site's icon. Ken and Robin
Talk About Stuff names http://kenandrobin.wpengine.com/.../kartas_podcast.png,
a 404, while its site's apple-touch-icon works. The feed's artwork now has
to answer as an image, as the site icon already did, when the feed is
read in full; otherwise the site's icon is looked for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:11:16 +00:00
4f8b3d6a1d Log as JSON when IPX_LOG_FORMAT=json (#91)
The log was text, so the Grafana dashboard picked lines apart with
regular expressions, and a change of wording would have blanked its
panels. With IPX_LOG_FORMAT=json each line is one JSON object: the
access log carries method, path, route, status and ms as fields (the
route passed from the routing layer in the response's extensions), and
each wire event its ev, feed, new, downloaded, failed, bytes, msg and
the rest (log_wire), beside the old message. The two startup lines that
were println! are logged, so no line breaks the JSON. Text stays the
default, for a terminal. The dashboard reads the fields with Loki's json
parser, and groups requests by route rather than path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 13:55:50 +00:00
9e7c93e149 Name request traces by route, and no colour codes off a terminal (#87, #88)
A request's trace was named by its path, so every item's GUID in
POST /api/entries/{feed_id}/{guid}/flags made a trace name of its own and
nothing grouped in Tempo. A route layer now renames it once routing has
matched. It renames the OpenTelemetry span directly: tracing-opentelemetry
drops a recorded otel.name once the span has been entered, and access_log
enters it before routing runs.

tracing-subscriber's fmt layer writes ANSI colour by default, so docker
logs and Loki (through Alloy) carried escape codes on every line, which
each query had to strip. Colour is now for a terminal only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 13:33:11 +00:00
1724346da7 Send OpenTelemetry traces over OTLP (#87)
ipx had no spans, only log lines, so there was no way to see where a
slow scan, download or request spent its time. With
OTEL_EXPORTER_OTLP_ENDPOINT set, the daemon now exports traces over
OTLP/HTTP (Tempo on Tower): a scan, each feed in it, the feed fetch and
site icon lookup, downloads, torrents, reaps, and web requests. Log lines
inside a span ride along as its events.

Only the daemon exports: the healthcheck runs ipx status every 30s and
would bury everything else. The web event stream and the log view's
polling get no span, for the same reason. The exporter shares ipx's
reqwest 0.13, so no second HTTP stack comes in.

The stderr log now prefixes lines inside a span with it, as
tracing-subscriber's fmt layer does (scan{only=None force=false}: ...).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 13:22:45 +00:00
8c5eddd783 Drop the start-up pass that folds files WordPress listed twice (#83)
Before 0.6.0 the parser took WordPress's numbered player URLs (?_=2) for
separate files and downloaded some episodes twice. Since then it drops
the repeats while reading (same_file_key), and merge_repeated_enclosures
cleaned up what was already stored. Production has run it; on every
start since it has only cost a query that finds nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 13:12:10 +00:00
469467bf04 A site icon is looked up again when the feed is read in full (#80)
The icon standing in for a feed's missing artwork was looked up once and
kept, so a site that changed or fixed its icon, or a feed that dropped
its own artwork, kept whatever was found first. A dead icon stored
before #79 would have stayed dead. It is now looked up whenever the
feed is read in full: when it has changed, or on a refresh someone asks
for, which reads in full since #77.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 12:51:19 +00:00
bae22e553e A refresh someone asks for reads the feeds in full (#77)
Check every feed, a feed's refresh, pull to refresh and ipx fetch --force
all send force, which only skipped the not-due wait: the request still
carried the stored ETag and Last-Modified, so an unchanged feed answered
304 and was not read. anil-dash got no site icon from a refresh for this
reason. A forced scan now drops the validators; the scheduled scan keeps
them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 23:47:08 +00:00
2004da3459 Feeds from before site icons get one, without waiting for a post (#73)
The icon lookup ran only when a scan got the feed's body, and most feeds
answer 304 to their stored validators, so anildash.com and 68 others
stayed blank until their next post. A feed whose image was never looked
for is now refetched once without validators, as an empty one already
was. A miss is stored as "" (drawn as no art), so neither the refetch nor
the site lookup repeats on every scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 23:27:00 +00:00
c6980c355d Adding a site's address subscribes to the feed it links (#71, #72)
Both add paths, the CLI's and the web's, look behind the URL first: a web
page that names its feed with <link rel="alternate"> is swapped for that
feed, before the duplicate check so it finds a feed someone already has.
Before, the page itself was added and every scan failed on it.

alternate_feed_link found tags in a to_lowercase() copy and sliced the
original at those offsets; Unicode lowercasing changes some characters'
length, so a page with one before its <link> tags lost the href or
panicked off a char boundary. ASCII lowercasing keeps offsets aligned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 23:21:23 +00:00
ce221cee18 A feed with no artwork takes its site's icon (#70)
When a feed names no image and none is stored, the scan fetches the
channel's site link (RSS <link>, Atom rel=alternate) and uses the
apple-touch-icon or icon it names, falling back to /favicon.ico when that
answers with an image.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 23:18:46 +00:00
5be629427a /api/status, for Homepage's dashboard (#67)
The iPX tile on Homepage was a bare link: nothing in ipx gave a summary a
customapi widget could read. /api/status serves what `ipx status` prints
(feeds, items pending, files downloaded), from the same function the
control socket answers with, plus the version. It sits behind sign-in
like the rest of /api; Homepage sends the shared [web] token as the
ipx_token cookie.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 19:10:05 +00:00
bdda9b3d2e Block lists: words that hide items and keep them from downloading (#47)
Each person has a list for every feed they read and one per feed. An item whose title or text
holds one of the words, matched as whole words so "ai" does not hide everything that "said"
anything, is hidden from them and, since the scanner now keeps each subscriber's filters
separate, is fetched only if someone else still wants it.

Whole-word matching is not something LIKE can do on both SQLite and Postgres, so the matches are
worked out in Rust into a `hidden` table whenever a list changes, someone subscribes, or a scan
brings in new items, and the queries only look that table up. Both new tables are tables rather
than columns because create_missing adds tables but never columns. Hidden counts as read for
the reaper and for "others still want this file", since whoever it is hidden from is as done
with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 13:13:03 +00:00
c1187a7926 Verify Cloudflare Access's signed token before trusting the proxy
The proxy sign-in believed Cf-Access-Authenticated-User-Email from any
address in trusted_proxies. On Tower that address is the Docker gateway,
so any container there could name itself anyone (docs/sso.md said as
much, and CLAUDE.md listed it as a known gap).

With [web] access_team and access_aud set, a proxied request must also
carry a Cf-Access-Jwt-Assertion that verifies against Cloudflare's keys
(RS256 only, this application's audience, the team's issuer, not
expired), and the name comes from its email claim. The keys are fetched
at start and again when a token names an unseen key, at most once a
minute, so made-up key ids cannot make every request a request to
Cloudflare. While the keys cannot be had, proxied sign-in is refused;
password and token sign-in are unaffected. Both settings empty, nothing
changes.

jsonwebtoken does the checking, on the aws-lc-rs backend already in the
tree through rustls. Tests sign with throwaway keys in tests/data: a
valid token, another app's audience, expired, a forged signature, HS256,
alg none, the refetch limit, and keys that cannot be fetched. Checked
live on a scratch daemon: the header alone and a forged token got 401,
the admin token still signed in.

vouched_name takes the peer and headers rather than the request: a
&Request held across the new await made the auth middleware's future
unsendable, as a body is not Sync.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 14:38:25 +00:00
3625cf48fb Keep the web token out of the startup log
The daemon printed http://<bind>/?token=<token> at every start. The token
signs in as the admin, and in the container that line lands in docker
logs, readable by anyone with Docker access on Tower. It now says where
the token is kept instead; config.toml already has it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 14:27:06 +00:00
905dfa0b02 A spinner on the feed being checked, not toasts; check only your own feeds
The scan's events reach everyone, so every browser showed "<feed>: N new" and
"Scanning…" toasts, and refreshed, for everyone's feeds. Now a feed's row, and
its folder's, carries a spinner between feed_start and its done, skip or error;
the list refreshes only for the reader's own feeds; the scan toasts are gone, and
"Downloaded" is said only for a file on screen.

"Check every feed" from the web UI sent a scan of every feed on the server.
Command::Fetch takes an optional `feeds` list -- those feeds and the feeds
inside any OPML among them -- and the web fills it with the asker's
subscriptions. The schedule and the CLI send none, meaning every feed.

Closes #37.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 01:17:40 +00:00
1cafd8d6e3 Keep the feed catalogue and server settings in the database
Phase 3 of #18. Two tables: catalogue (each feed's config::Feed as JSON, so a
new feed setting needs no column) and settings (general: the five server
settings the admin page edits). config.toml keeps what is needed before the
database is reached, or decides who gets in: paths, [torrent], [web].

ipx still runs from one in-memory Config, assembled at start from both
(assemble_config). The eight places that saved config.toml and re-read it now
call Ctx::store_cfg, which writes the database and swaps the copy in memory; the
first-run web token, which is config.toml's, is written there.

The first start on a database with no catalogue imports config.toml's feeds and
settings in one transaction whose first insert is the settings row, so two ipx
starting at once cannot both import; it then trims config.toml, keeping the
original as config.toml.pre-database. After that, feeds written into the file are
ignored with a warning. copy-db skips it, and copies both tables.

Rehearsed on a clone of production's database with production's config: all 130
feeds imported, the file trimmed, and the feed list, settings and directory
identical to the live server's.

Postgres connections now ask for no notices. Every CREATE ... IF NOT EXISTS on an
existing table sends one, eleven per open; sqlx logs them, and
tracing-subscriber 0.3.23's per-layer filters then dropped the next line ipx
logged -- the import's own message went missing that way. Proved by toggling it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 22:28:49 +00:00
bc53e0f730 Postgres: pick the database by URL, copy-db, tests on both
- IPX_DATABASE_URL (postgres://...) picks the database; unset, it is the SQLite
  file as before. Passwords are taken out of anything logged.
- `ipx copy-db <state.db>` copies every table into the empty database the URL
  names, in one transaction, and moves the id counters past the copied ids. A
  copy of production went across in 14s with every count and column
  fingerprint identical.
- With IPX_TEST_DATABASE_URL set, each test gets a Postgres schema of its own;
  all 79 pass on both databases. Fixtures write booleans as true/false.
- Sorts say where an item with no value goes (NULLS FIRST going up, LAST going
  down): SQLite counts NULL as smallest, Postgres as largest, so "largest first"
  on Postgres led with every item that has no file. Tested on both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 20:28:50 +00:00
611716d8b7 SeaORM: feeds and scanning; rusqlite gone
The last nineteen functions move to SeaORM: recording feeds, items and
enclosures, managed OPML feeds, folding WordPress's repeated files, and handing a
Patreon creator's files to its shows. Two SQLite-only forms go: GLOB becomes a
LIKE with the underscore escaped (broader, harmlessly: the fold still keys on
`_=` and digits), and UPDATE OR IGNORE becomes an UPDATE ... WHERE NOT EXISTS.
The two transactions are SeaORM transactions.

With nothing left on it, rusqlite goes, with the SQL schema and migrate(). The
entities are the schema: create_missing makes whatever tables and indexes a
database lacks, from them, with CREATE ... IF NOT EXISTS. Production's schema
already has every column migrate() added and none it dropped.

Not SeaORM's schema sync, used until now: despite its docs it drops a unique
index the entities do not describe, so it dropped users_name_lower on every open.
Every `ipx` command then took a write lock, and against a daemon busy writing,
`ipx status` -- the healthcheck -- failed 7 times in 15 where the old code
failed none. Now 15 in 15, as before. On Postgres it would not have started.

WAL is set only when a file is not already in it: setting it takes a lock that
cannot wait out a busy daemon.

Checked on copies of production: a forced scan of all 162 feeds against the real
feeds with no database errors; the feed list, filters, sorts, search and the
reaper's candidates against the old code on the same data, earlier in the
branch. The column comments from the SQL schema move to the entities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 19:10:59 +00:00
a68bfb179b SeaORM: enclosures, downloads and the reaper
Twelve enclosure functions move to SeaORM: recording, the download queue,
marking done or failed, requeueing, and what the reaper may delete. INSERT OR
IGNORE becomes ON CONFLICT DO NOTHING; the reaper's read verdict is true or
false rather than 1 or 0, which Postgres would type as a 32-bit integer and
refuse to read as an i64; `read = 1` and `flagged = 1` test the booleans
themselves. retention::run and its callers (reap, rm, retire_group,
retire_stranded) become async.

The reaper deletes files, so it was checked on a copy of production against the
old SQL on the same file: all 2,195 candidates, identical and in the same order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:38:07 +00:00
6bf1ad6b31 SeaORM: items and read state
The item list, its counts, filters, sorts and search, positions, pins and
mark-all-read move to SeaORM, as SQL written for both databases:

- Parameters are gathered as the SQL is written (Args), so only what a
  statement uses is bound. rusqlite needed every one mentioned, hence the old
  `?1 IS NULL` and `?2 = ''`; Postgres refuses a parameter it cannot type.
- Yes/no columns are tested as booleans (NOT coalesce(s.read, false)) and
  written as true, not 1; SQLite reads true and false as 1 and 0.
- The last tiebreak of the sort is the guid, not SQLite's rowid, which Postgres
  lacks. Only items with the same date change places.
- set_position names entry_state.duration beside excluded.duration.
- The status callback on the control socket returns a future, as the counts
  are now a query.

Checked on a copy of production against the live server: 42 of 48 lists
identical; the other six differ only in how ties fall, or because the test
daemon cleared paths to files this machine does not have. Run on the same file,
every filter's count matches the old SQL exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:46 +00:00
d7f8f2df1d SeaORM: subscriptions and pins
Twelve subscription functions move to SeaORM. Lookups use the entity API; the
joins, counts and upserts are SQL written to run on both databases: $n
parameters, ON CONFLICT DO NOTHING in place of INSERT OR IGNORE, and
CASE WHEN on the yes/no column itself rather than comparing it to 1, which
Postgres would refuse for a boolean. INSERT ... SELECT ... ON CONFLICT gets a
WHERE true, which SQLite needs to tell the two apart.

Checked with a daemon on a copy of production: the feed list, read through the
new code, comes back with every feed and its settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:23:37 +00:00
4927677e66 SeaORM: accounts, sessions and themes
The fourteen user and session functions move from rusqlite to SeaORM and become
async; their callers await them (auth, admin_user, user_cmd, the account
handlers). Checked against a copy of production, where the yes/no columns are
still INTEGER: the admin flag reads back right.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:17:52 +00:00
484aaa1849 SeaORM beside rusqlite: entities, schema sync, a second connection
The first step of moving to SeaORM (#18, phase 1). Nothing a user sees changes.

- src/entity.rs: the seven tables as SeaORM entities, matching the SQLite schema.
  Strings are Text, as the columns are; yes/no columns are bool, which is BOOLEAN
  on Postgres and stays INTEGER in the existing SQLite file (sync notes the
  difference and leaves it alone).
- Db holds a SeaORM connection to the same SQLite file beside the rusqlite one;
  functions move to it one at a time, and rusqlite goes with the last of them.
- db::sync creates what a database is missing from the entities (SeaORM's
  schema-sync, experimental, so sea-orm is pinned to ~2.0), plus the two indexes
  an entity cannot express. Checked against a copy of production: it added the
  lower(name) index and changed nothing else.
- Test databases are now built from the entities alone, in a temporary file
  (two connections to one ":memory:" are two databases), so every test also
  checks that the entities describe what the queries need. That caught the one
  difference: finding a user by name relied on COLLATE NOCASE, which Postgres
  lacks; it now compares lower() on both sides.
- rusqlite steps back to 0.39: 0.40's libsqlite3-sys is newer than sqlx accepts,
  and only one may link SQLite. It goes away at the end of this phase.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:11:28 +00:00
49dafedbbc Inter, one file where WordPress listed two, and a pin heading on line
Inter (#11): the pages are set in Inter's variable font, served from the
binary at /inter.woff2 as the icon is, with its OFL licence beside it in
web/. Classic keeps Lucida Grande, the 2004 app's face.

Double audio (#12): WordPress numbers each audio player on a page by
adding ?_=N to its file's URL, so a post that embeds the file it encloses
listed it twice, and it was downloaded twice. The parser keeps the first
of an item's enclosures that differ only by that number. At startup the
repeats already stored fold into the first; where only the repeat had
been downloaded its file moves to the first rather than being deleted.

Pin heading (#13): the rows' icon buttons kept the browser's side
padding, which pushed their 16px icon 3px right of centre, and the
heading's icon sat at the left of its column. Both are centred now, and
the heading row takes the pixel of border the rows have, so every
heading sits over its column.

Closes #11, closes #12, closes #13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 17:31:03 +00:00
b83523fc92 Directory: let an admin give a blog its category
Almost no blog names a category the Directory can use, so a feed can
carry one of its own in config.toml, set by an admin in the feed's
settings and used when the feed names none. The feed's own iTunes
category still wins. The field offers the categories the Directory
already shows, so a blog about games joins Games rather than starting a
second chip. Setting it on a feed from an OPML promotes it to config, as
any other shared setting does.

Closes #10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 17:19:59 +00:00
4314b11237 Retire davewiner's 922 rows, without taking the ones people still read
davewiner's OPML left config.toml before retire_group existed, so its
derived rows were skipped by every scan but never cleared. At startup the
daemon now retires every group whose parent is gone from config. And
retire_group unmanages a feed that has its own config entry instead of
dropping it: eleven of davewiner's were promoted without being unmanaged,
and dropping them as derived would have deleted their entries. That also
covers removing an OPML or Patreon subscription from the page.

Closes #3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 17:01:27 +00:00
954173cacf Directory: chips by kind and category over a grid of cover art
Feeds take their channel's first <itunes:category> into a new feeds.category
column; the migration drops ETag and Last-Modified once so every feed re-reads
on its normal schedule and picks one up. /api/popular and /api/directory carry
category and podcast (any audio or video enclosure). Directory becomes a grid of
cover-art tiles under a chip rail: All, Podcasts, Blogs, and a podcast's
categories once Podcasts is picked. Popular and Add a feed keep their rows.

Closes #4, closes #5, closes #6, closes #7.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 16:09:47 +00:00
be3820bbbd Release 0.5.3: OPML orphan scan, feed error UI, small UI fixes
- Stop scanning an OPML/Patreon feed's derived rows once nobody subscribes
  to it; retire them (drop or orphan) the way sync_group already does when
  the list itself drops one. This is what let 922 defunct davewiner feeds
  keep scanning hourly after the OPML left config.
- Repair feed XML with a bare `&`, and give a plain reason (moved web page
  with its new address when linked, or nothing yet for an empty body)
  instead of a raw parser error.
- Show a failing feed's plain-English reason and next step (Unsubscribe /
  Use the new address) in the sidebar and on its own page, once it has
  been down a day.
- Fix four small UI bugs: show-note links open in a new tab, video files
  play as video, an opened item no longer disappears from the Unread tab,
  and Subscribe/Unsubscribe get their own icons.
- Fix Settings disappearing for non-admin accounts: it was hiding the
  whole modal instead of just the admin-only parts (Users, the editable
  schedule/quota, Save), which are the only parts the server actually
  refuses them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DmQfE1eFPApnXWyPHBWqUA
2026-09-14 14:53:45 +00:00
2ff2074755 Answer status to the client that asked, not everyone
status is a terminal event. Broadcast, the healthcheck's answer ended any
ipx fetch that was watching a scan, which stopped reading at the next probe
while the scan carried on. It could not happen while status waited behind
the scan; answering it at once made it happen every 30 seconds. Each
connection's writer now takes private replies beside the broadcast, and
the test checks another client hears nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TAC7sLVqfKmY6rsTLXzNgk
2026-09-12 14:27:42 +00:00
1698cf8d1e Answer status on the socket instead of queuing it behind the worker
The worker runs one job at a time, and status was one of its jobs, so the
Docker healthcheck waited behind the startup scan (54 seconds of it after
the last deploy) and timed out at 5. Any scan or download longer than
three probes would have had a working daemon marked unhealthy. The socket
now answers status straight away; everything else still queues. A test
fills the queue and checks status comes back anyway.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TAC7sLVqfKmY6rsTLXzNgk
2026-09-12 14:19:00 +00:00
586d2c07a1 ipx user rename: give an account the name the proxy signs it in as
An account made by hand before the proxy was set up is called what it was
given ('rays'), while Cloudflare Access vouches for an email address. With
auto_create_users on, the first visit through the tunnel would make a
second, empty account. Renaming keeps the id, so feeds, read state and
admin rights go with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TAC7sLVqfKmY6rsTLXzNgk
2026-09-12 13:54:29 +00:00