Files
ragflow/rag/flow/tests/test_token_chunker_delimiter.py
Jack 9b05e5c67e Fix: delimiter is chunk boundary, drop token_size atom-split (OVER_CAP default) (#17808)
## Summary

Fixes a regression introduced by #17203 (strict-cap atom-split) and a
secondary delimiter-handling bug from #17723.

**Root cause:**
- #17203 added `_split_oversized_unit` / `_compute_chunk_update`, which
split oversize units into ≤ token_size pieces. This collapsed
`token_size=1` into 1-token chunks and set the cap at 512, mismatching
the model-layer truncation boundary (embedding ~8191 / rerank
500/4096/8192/2048). Atom-split is unnecessary: oversize units stay
whole and the model layer truncates.
- #17723's delimiter handling dropped consecutive delimiters (`A####B`
-> `A##B`), glued JSON items with `"".join`, ignored
`children_delimiters`, and stripped whitespace delimiters.

## Changes

- New pure helper `merge_paragraphs(paragraphs, token_size, strategy)`
with a `MergeStrategy` enum (`UNDER_CAP` / `OVER_CAP`); **default
`OVER_CAP`**. `UNDER_CAP` is a strict cap (never overflows
`token_size`); `OVER_CAP` greedily accumulates adjacent paragraphs while
the projected total stays within `token_size`, merging one
boundary-overflow paragraph before closing. Oversize paragraphs stand
alone.
- `naive_merge` / `naive_merge_with_images` /
`RAGFlowTxtParser.parser_txt` now use `merge_paragraphs`; atom-split
removed. `naive_merge` / `naive_merge_with_images` always split a
section on the delimiter whenever one is present (even when the section
already fits `token_size`), so delimiter text never leaks into a chunk.
Only the empty-delimiter (size-only) mode skips splitting.
- `token_chunker`: delimiter text is dropped (not stripped); JSON flush
joins buffered items with `"\n"`; `children_delimiters` and
`PDF_POSITIONS_KEY` are preserved on the delimiter path. PDF positions
are now attributed **per segment** — each split chunk carries only the
positions of the item(s) that contributed to it — fixing a leak where
page-N coordinates were attached to page-M chunks and all segments
shared one preview image.
- `test_txt_parser.py` rewritten to assert the new contract (not the old
strict cap); `naive_merge` and delimiter-case-sensitive matrices
updated.

## Contract (refs #17799)

- user specified delimiter = chunk boundary; user specified delimiter
text never enters a chunk.
- `token_size` = soft target + merge strategy; no atom-split.
- Default strategy = `OVER_CAP`; migration can switch to `UNDER_CAP`
(strict cap).
- `OVER_CAP` has no hard cap; the model layer truncates oversize units.
`UNDER_CAP` enforces a strict cap.

## Notes

- Closes the wrong-object revert in #17774 (revert #17723 would
re-introduce delimiter-in-chunk and the strict cap).
- Go-side alignment (`internal/ingestion/component/chunker/token.go`) is
a follow-up PR.

---------

Co-authored-by: CodeBuddy <noreply@tencent.com>
2026-08-05 11:50:07 +08:00

117 lines
5.3 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Regression test for token_size-mode delimiter handling in TokenChunker.
The default ``delimiters`` a user supplies (e.g. ``["\\n"]``) are only honoured
when they are backtick-wrapped; a bare delimiter makes
``_compile_delimiter_pattern`` return ``""``, so ``TokenChunker._invoke`` falls
through to ``naive_merge(payload, chunk_token_size, "")`` (token_chunker.py:323).
Passing ``""`` discards the sentence-boundary preference that ``naive_merge``
carries by default (``"\\n。"``). Without a forced ``\\n`` break, the token
budget flush can land *inside* a sentence, so a line like
``"Sentence number 7. alpha beta gamma delta epsilon zeta eta theta iota kappa"``
gets split across two chunks. The Go port (``sentenceDelimiter = (\\n|[!?。;!?])``
in token.go:340) always breaks on ``\\n``, so it never cuts a sentence mid-stream.
This test exposes that divergence: with ``delimiters=["\\n"]`` in token_size mode,
every original sentence must survive intact in a single chunk. It currently FAILS
because of the ``""`` passed at token_chunker.py:326; forwarding
``"".join(self._param.delimiters)`` to ``naive_merge`` (or at least
``"\\n。"``) makes it pass and aligns Python with the Go side.
Note: the test needs a live tiktoken encoder. ``common.token_utils`` returns 0 on
any encoder failure, which would make every token budget "never exceeded" and
hide the bug behind a single-chunk baseline. We skip rather than pass under a
dead tokenizer, mirroring ``capture_golden.py::_assert_tokenizer_alive``.
"""
import asyncio
import types
from rag.flow.chunker.token_chunker import TokenChunker, TokenChunkerParam
from rag.nlp import naive_merge
from common.token_utils import num_tokens_from_string
def _build_token_chunker(param: dict) -> TokenChunker:
"""Construct a TokenChunker without the real Graph the normal __init__ wants.
``ComponentBase.__init__`` asserts ``canvas`` is a real ``Graph``; a unit test
has none. Bypass it and attach the three attributes ``_invoke`` actually reads
(``_canvas``, ``_param``, ``callback``) so the genuine chunking path runs with
the real ``naive_merge`` and the real ``num_tokens_from_string``. Mirrors
``capture_golden.py::_build_component``.
"""
p = TokenChunkerParam()
for key, value in param.items():
setattr(p, key, value)
p.check() # __new__ bypassed the constructor's validation; reproduce it.
comp = TokenChunker.__new__(TokenChunker)
comp._canvas = types.SimpleNamespace(_doc_id=None, _tenant_id="t")
comp._param = p
comp.callback = lambda *_a, **_kw: None
return comp
def _invoke_text(comp: TokenChunker, text: str) -> list[dict]:
asyncio.run(comp._invoke(name="t", output_format="text", text=text))
return comp._param.outputs["chunks"]["value"]
def _make_sentences(n: int) -> list[str]:
# Each sentence is ~16 tokens of plain English; well under the 128 budget,
# so several fit per chunk and the token-budget flush lands mid-sentence.
return [f"Sentence number {i}. alpha beta gamma delta epsilon zeta eta theta iota kappa" for i in range(n)]
def test_token_chunker_token_size_mode_does_not_split_sentences():
# Guard: a dead tokenizer would collapse everything to one chunk and hide
# the bug. Skip instead of recording a poisoned pass.
if num_tokens_from_string("alive tokenizer probe sentence") <= 0:
import pytest
pytest.skip("tiktoken encoder unavailable; num_tokens_from_string returned 0")
sentences = _make_sentences(30)
payload = "\n".join(sentences)
comp = _build_token_chunker({"chunk_token_size": 128, "delimiters": ["\n"]})
chunks = _invoke_text(comp, payload)
texts = [c["text"] for c in chunks]
# An original sentence is "intact" when its full text appears in exactly one
# chunk. If the budget flush fell inside it, the sentence is split across two
# chunks and this assertion fails.
split_sentences = [s for s in sentences if not any(s in t for t in texts)]
assert not split_sentences, (
f"{len(split_sentences)} sentence(s) were cut mid-stream by "
f"token_size mode with delimiters=['\\n'] (got {len(chunks)} chunks). "
"naive_merge was called with an empty delimiter (token_chunker.py:326), "
"dropping the '\\n' sentence-boundary preference. Example split: "
f"{split_sentences[0]!r}"
)
def test_naive_merge_empty_delimiter_keeps_unit_whole():
"""Empty delimiter -> no split; the whole payload is one chunk (no
atom-split). A '\\n' delimiter still honours the newline boundary.
Documents the new contract: without a delimiter there is nothing to split
on, so the unit is kept whole and the model layer truncates it.
"""
if num_tokens_from_string("alive tokenizer probe sentence") <= 0:
import pytest
pytest.skip("tiktoken encoder unavailable; num_tokens_from_string returned 0")
sentences = _make_sentences(30)
payload = "\n".join(sentences)
chunks_empty = naive_merge(payload, 128, "")
chunks_nl = naive_merge(payload, 128, "\n")
# Empty delimiter no longer cuts (no atom-split): everything stays in one chunk.
assert len(chunks_empty) == 1
split_nl = [s for s in sentences if not any(s in t for t in chunks_nl)]
assert not split_nl, "naive_merge('\\n') should preserve every sentence boundary"