Commit Graph

2 Commits

Author SHA1 Message Date
Saurav Panda ab08dea62c Fix markdown extraction destroying URLs and dropping long link lines
Two content-destruction bugs in extract_clean_markdown:

- A cleanup regex stripped every %XX sequence from the converted markdown,
  corrupting all percent-encoded URLs (%20, %2F, ...) — precisely when
  extract_links=True was requested.
- The JSON-blob line filter dropped any line over 100 chars starting with
  '{' OR '[' — silently deleting long markdown links [text](long-url),
  clickable images, and citation-style lines.

Delete the %XX regex, and only drop long lines that actually parse as JSON
(json.loads) so SPA state blobs are still filtered while markdown links
survive. Also extract the HTML->markdown conversion into a pure
convert_html_to_markdown() helper so this stage is unit-testable.
2026-07-06 14:45:25 -07:00
Cursor Agent 24e276d5f4 Refactor markdown preprocessing and add tests
Co-authored-by: mailmertunsal <mailmertunsal@gmail.com>
2025-12-11 00:00:21 +00:00