mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-08-13 20:26:51 +08:00
### Summary `format_document_soup` tracks "am I inside a table" and "am I inside a link" with sticky flags that are meant to be reset by `elif e.name == "/table"` and `elif e.name == "/a"`. BeautifulSoup's `.descendants` only yields opening tags — a `Tag` named `/table` or `/a` never exists — so both branches are dead code and neither flag is ever cleared. Everything after the first `<table>` on a page is therefore formatted as if it were still table content: paragraphs lose their newline, list items lose their `- ` marker, headings lose their break, and the text is glued onto the last table cell. Under `HTML_BASED_CONNECTOR_TRANSFORM_LINKS_STRATEGY=markdown` the same bug leaks a link's `href` into everything that follows it, including whole subsequent paragraphs. The Confluence connector (`confluence_connector.py:948`) goes through this path. Real output for a Confluence-shaped page (heading, intro, spec table, then the body) via the public `parse_html_page_basic`: **Before** ``` prod us-east-1 Rollback procedure If the canary fails, run the rollback script immediately. Drain the load balancer Revert the deployment Escalate to the on-call rota if the rollback stalls. Do not skip the post-mortem. ``` **After** ``` prod us-east-1 Rollback procedure If the canary fails, run the rollback script immediately. - Drain the load balancer - Revert the deployment Escalate to [the on-call rota](http://oncall.example.com) if the rollback stalls. Do not skip the post-mortem. ``` Every heading, paragraph and list marker after the table is lost, and the whole body is indexed as one run-on line hanging off a table cell. ### Fix Derive both scopes from each element's **ancestors** instead of from flags that nothing can clear, and drop the two dead branches plus the two that become redundant. The scopes are resolved in one up-front pass into `id`-keyed maps (`table_scope`, `href_scope`) and looked up in O(1) per element. Probing per element with `find_parent` instead is O(depth) each, which measured 12–13× slower on table-heavy pages and up to 103× on deeply nested markup; the map version costs a depth-independent 1.13–1.35× over `main`. Numbers and method are in the round-2 comment below. This also changes one adjacent behaviour worth calling out explicitly: a link **inside** a table cell now renders as markdown, where before it rendered as plain text. That previous behaviour was not by design — it only held when no link preceded the table. With a link before the table, `main` stamps the stale href onto every cell: ``` main: '[pre](http://STALE.com)\n\t[cellA](http://STALE.com)\t[cellB](http://STALE.com)' branch: '[pre](http://STALE.com)\n\tcellA\tcellB' ``` Those cells are not links. Both symptoms are the same sticky-state bug, so they are fixed together rather than left half-done. ### Testing `test/unit_test/data_source/test_html_utils.py` is new — `format_document_soup` had no test coverage. 11 tests: 8 fail on `main` and pass on this branch, 3 are controls that pass on both (the table itself still separates rows and cells, anchor text is still linkified, the default `strip` strategy still strips). Representative failures on `main`: ``` assert '\nAfter' in 'Before\n\tA\tB After' assert '\n- item1' in 'Before\n\tA\tB item1 item2' assert 'see [link](http://x.com) [ after](http://x.com)' == 'see [link](http://x.com) after' assert '[next paragraph]' not in '[link](http://x.com)\n[next paragraph](http://x.com)' ``` Reverting each clause of the fix independently keeps the anchors honest: reverting only the table clause fails exactly the 4 table tests and leaves the link tests green; reverting only the link clause fails exactly the 3 link tests and leaves the table tests green. (`test_link_inside_a_table_cell_is_linkified` needs both clauses broken to fail, so it appears in neither single-clause revert — it is covered by the 8-fail run against `main`.) Full `test/unit_test/data_source/` suite: **3 failed, 199 passed**, and the failure set is byte-identical to clean `main` (**3 failed, 188 passed**) — the 3 are `TestSSRFValidation::*`, which resolve `api.example.com` against real DNS and are unrelated to this change. `ruff check` and `ruff format --check` are clean on both touched files. --- This PR was drafted with AI assistance (Claude). I reviewed the change, independently reproduced both symptoms against `main`, and take responsibility for it.
(1). Deploy RAGFlow services and images
https://ragflow.io/docs/build_docker_image
(2). Configure the required environment for testing
Install Python dependencies (including test dependencies):
uv sync --python 3.13 --only-group test --no-default-groups --frozen
Activate the environment:
source .venv/bin/activate
Install SDK:
uv pip install sdk/python
Modify the .env file: Add the following code:
COMPOSE_PROFILES=${COMPOSE_PROFILES},tei-cpu
TEI_MODEL=BAAI/bge-small-en-v1.5
RAGFLOW_IMAGE=infiniflow/ragflow:v0.26.4 #Replace with the image you are using
Start the container(wait two minutes):
docker compose -f docker/docker-compose.yml up -d
(3). Test Elasticsearch
a) Run sdk tests against Elasticsearch:
export HTTP_API_TEST_LEVEL=p2
export HOST_ADDRESS=http://127.0.0.1:9380 # Ensure that this port is the API port mapped to your localhost
pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_sdk_api
b) Run http api tests against Elasticsearch:
pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_http_api
(4). Test Infinity
Modify the .env file:
DOC_ENGINE=${DOC_ENGINE:-infinity}
Start the container:
docker compose -f docker/docker-compose.yml down -v
docker compose -f docker/docker-compose.yml up -d
a) Run sdk tests against Infinity:
DOC_ENGINE=infinity pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_sdk_api
b) Run http api tests against Infinity:
DOC_ENGINE=infinity pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_http_api