Looking for a page's feed link can panic on non-ASCII text #72

Closed
opened 2026-09-28 16:19:09 -07:00 by rays · 1 comment
Owner

feed.rs alternate_feed_link finds tags in text.to_lowercase() and slices the original text at those byte offsets. Unicode lowercasing can change a character's length (e.g. U+0130 becomes 3 bytes from 2), so a page with such a character before its tags is sliced at the wrong place, off a char boundary, which panics. Found while reading the code for #71.

feed.rs alternate_feed_link finds <link> tags in text.to_lowercase() and slices the original text at those byte offsets. Unicode lowercasing can change a character's length (e.g. U+0130 becomes 3 bytes from 2), so a page with such a character before its <link> tags is sliced at the wrong place, off a char boundary, which panics. Found while reading the code for #71.
rays added the bug label 2026-09-28 16:19:09 -07:00
Author
Owner

Fixed in c6980c3: alternate_feed_link lowercases ASCII only, so offsets match the original text; test a_feed_link_after_non_ascii_text_is_found. Deployed.

Fixed in c6980c3: alternate_feed_link lowercases ASCII only, so offsets match the original text; test a_feed_link_after_non_ascii_text_is_found. Deployed.
rays closed this issue 2026-09-28 16:22:51 -07:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: rays/ipx#72