Two content-destruction bugs in extract_clean_markdown:
- A cleanup regex stripped every %XX sequence from the converted markdown,
corrupting all percent-encoded URLs (%20, %2F, ...) — precisely when
extract_links=True was requested.
- The JSON-blob line filter dropped any line over 100 chars starting with
'{' OR '[' — silently deleting long markdown links [text](long-url),
clickable images, and citation-style lines.
Delete the %XX regex, and only drop long lines that actually parse as JSON
(json.loads) so SPA state blobs are still filtered while markdown links
survive. Also extract the HTML->markdown conversion into a pure
convert_html_to_markdown() helper so this stage is unit-testable.