Fix: Q&A CSV parser splices wrong text when a field opens with a quote (#16881)

### Problem

Parsing a Q&A `.csv` can splice unrelated text into the wrong answers
(reported in #16791).

### Root cause

The `.csv` branch of `rag/app/qa.py`'s `chunk()` builds records with
`csv.reader(lines, delimiter=delimiter)` (default `quotechar='"'`), but
then indexes `lines[i]` by the reader's *record* index in `answer +=
"\n" + lines[i]`. When a line's field opens with a `"`, `csv.reader`
treats it as an unclosed quoted field and merges several physical lines
into one record. From there the record index permanently desyncs from
the physical line numbers, so `lines[i]` returns the wrong line and
unrelated Q&A content gets appended to the wrong answer.

### Reproduction (stdlib only)

```python
import csv
lines = 'Q1,A1\n"quoted answer start\ncontinues here,extra\nQ2,A2\n'.split("\n")
list(csv.reader(lines, delimiter=","))
# record 1 swallows 3 physical lines: ['quoted answer startcontinues here,extraQ2,A2']
# -> the reader index no longer matches lines[i]

list(csv.reader(lines, delimiter=",", quoting=csv.QUOTE_NONE))
# one physical line per record; index stays aligned
```

### Fix

Pass `quoting=csv.QUOTE_NONE` so one physical line maps to one record,
keeping the reader index aligned with `lines[i]` (the surrounding code
already relies on that 1:1 mapping).

Fixes #16791.

---------

Signed-off-by: Yash Raj Pandey <yashpn62@gmail.com>
Co-authored-by: Yingfeng <yingfeng.zhang@gmail.com>
This commit is contained in:
Yash Raj Pandey
2026-07-18 06:31:43 -04:00
committed by GitHub
parent a7b193d77b
commit 20a7c7d17a
2 changed files with 45 additions and 2 deletions

View File

@@ -353,12 +353,16 @@ def chunk(filename, binary=None, from_page=0, to_page=MAXIMUM_PAGE_NUMBER, lang=
fails = []
question, answer = "", ""
res = []
reader = csv.reader(lines, delimiter=delimiter)
reader = csv.reader((line + "\n" for line in lines), delimiter=delimiter)
prev_line_num = 0
# line_num tracks the physical span when quoted fields cross lines.
for i, row in enumerate(reader):
raw = "\n".join(lines[prev_line_num : reader.line_num])
prev_line_num = reader.line_num
if len(row) != 2:
if question:
answer += "\n" + lines[i]
answer += "\n" + raw
else:
fails.append(str(i + 1))
elif len(row) == 2: