mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-07-20 22:51:06 +08:00
### Problem Parsing a Q&A `.csv` can splice unrelated text into the wrong answers (reported in #16791). ### Root cause The `.csv` branch of `rag/app/qa.py`'s `chunk()` builds records with `csv.reader(lines, delimiter=delimiter)` (default `quotechar='"'`), but then indexes `lines[i]` by the reader's *record* index in `answer += "\n" + lines[i]`. When a line's field opens with a `"`, `csv.reader` treats it as an unclosed quoted field and merges several physical lines into one record. From there the record index permanently desyncs from the physical line numbers, so `lines[i]` returns the wrong line and unrelated Q&A content gets appended to the wrong answer. ### Reproduction (stdlib only) ```python import csv lines = 'Q1,A1\n"quoted answer start\ncontinues here,extra\nQ2,A2\n'.split("\n") list(csv.reader(lines, delimiter=",")) # record 1 swallows 3 physical lines: ['quoted answer startcontinues here,extraQ2,A2'] # -> the reader index no longer matches lines[i] list(csv.reader(lines, delimiter=",", quoting=csv.QUOTE_NONE)) # one physical line per record; index stays aligned ``` ### Fix Pass `quoting=csv.QUOTE_NONE` so one physical line maps to one record, keeping the reader index aligned with `lines[i]` (the surrounding code already relies on that 1:1 mapping). Fixes #16791. --------- Signed-off-by: Yash Raj Pandey <yashpn62@gmail.com> Co-authored-by: Yingfeng <yingfeng.zhang@gmail.com>
(1). Deploy RAGFlow services and images
https://ragflow.io/docs/build_docker_image
(2). Configure the required environment for testing
Install Python dependencies (including test dependencies):
uv sync --python 3.13 --only-group test --no-default-groups --frozen
Activate the environment:
source .venv/bin/activate
Install SDK:
uv pip install sdk/python
Modify the .env file: Add the following code:
COMPOSE_PROFILES=${COMPOSE_PROFILES},tei-cpu
TEI_MODEL=BAAI/bge-small-en-v1.5
RAGFLOW_IMAGE=infiniflow/ragflow:v0.26.4 #Replace with the image you are using
Start the container(wait two minutes):
docker compose -f docker/docker-compose.yml up -d
(3). Test Elasticsearch
a) Run sdk tests against Elasticsearch:
export HTTP_API_TEST_LEVEL=p2
export HOST_ADDRESS=http://127.0.0.1:9380 # Ensure that this port is the API port mapped to your localhost
pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_sdk_api
b) Run http api tests against Elasticsearch:
pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_http_api
(4). Test Infinity
Modify the .env file:
DOC_ENGINE=${DOC_ENGINE:-infinity}
Start the container:
docker compose -f docker/docker-compose.yml down -v
docker compose -f docker/docker-compose.yml up -d
a) Run sdk tests against Infinity:
DOC_ENGINE=infinity pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_sdk_api
b) Run http api tests against Infinity:
DOC_ENGINE=infinity pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_http_api