- text_parser.go now uses the shared DefaultTextCodeDelimiter from
delimiter.go instead of a private duplicate that could drift silently.
- Fix the golden-regeneration comment: no generator script is committed;
the baseline is reproducible from the golden's meta block.
- Add TestTextParser_AdjacentDelimiters pinning the documented adjacent-
delimiter behavior (Go drops the standalone second delimiter) so a
future silent change in splitCapturingDelims is caught.
Make Go TextParser match Python deepdoc TxtParser._code:
- Split on the flow default delimiter set "\n!?;。;!?" (was blank-line only).
- keep_delimiters=True: each trailing delimiter is retained on its segment.
- Normalize CRLF/CR -> LF like Python.
- Do NOT perform the OVER_CAP token merge; chunking stays with the downstream
Chunker (PARSER_ALIGNMENT_HANDOFF.md §2.3).
- The 8192-byte per-item cap is dropped: the parser is a pure delimiter
splitter, matching Python's parser_txt (no size slicing). Sizing is delegated
to the Chunker + embedding truncation per contract #17799.
Tests:
- TestTextParser_ParseWithResult_DefaultDelimiter pins the new split rule
(single-newline + sentence-delimiter splitting with delimiter retained).
- TestTextParser_ParseWithResult_NoSizeCap pins no per-item byte slicing.
- TestTextParser_AlignmentGolden verifies content-equivalence vs the Python
golden on a shared sample.
Depends on #18014 and the foundation helpers PR (align_test.go).
Files:
internal/parser/parser/text_parser.go
internal/parser/parser/parse_with_result_test.go
internal/parser/parser/testdata/textcode.sample.txt (new)
internal/parser/parser/testdata/textcode.python.golden.json (new)
DefaultTextCodeDelimiter and DefaultMarkdownDelimiter are shared by both
the production parsers and the alignment tests. They previously lived in
align_test.go (a _test.go file), so production text_parser.go had to keep
a private duplicate that could drift silently. Move them to a non-test
delimiter.go as the single source of truth.
Extract the alignment scaffolding so the format-specific PRs (text&code,
markdown golden, HTML) don't conflict on align_test.go.
Adds (test-only, no production behavior change):
- GoldenDoc{Meta,Items} and parseGolden: tolerant of both the legacy bare
array and the new {meta, items} golden format.
- LoadGoldenDoc, AcceptedDivergences(meta), FilterOutDocTypes(items, drop):
meta-driven divergence handling, no hardcoded divergence lists.
- TextCodeAlignOptions / DefaultTextCodeDelimiter, HTMLAlignOptions /
StripHTMLHeadingMarker: normalizer presets reused by the alignment tests.
- LoadGolden now tolerates the {meta, items} format.
Stacks on #18014 (the align_test.go framework is already in main).
Port dataset-level wiki incremental compile and refactor splitByTokens
token budgeting. Includes replace-only wiki merge, KNN dedup routing,
and template/config wiring.
## Summary
This PR improves the RAGFlow agentic-search path in three areas: it
stops the outer agent from re-looping over the same rag call, lets the
medium thinking mode discover and follow new sub-claims mid-loop, and
strengthens retrieval by having the LLM emit synonym-rich queries with
time/date/number terms boosted.
1. Avoid the outer re-loop — keep all multi-hop cycles inside agentic
RAG
2. Dynamic claims in medium mode — keep querying newly discovered
sub-questions
medium now enables allows_dynamic_claims. During orchestration, when
claim analysis discovers a new required sub-question
(discovered_claims), the loop spawns it as a new ClaimTarget and
continues searching it in subsequent cycles (bounded by the
dynamic-claim budget) instead of stopping. Also added:
3. Stronger query strategy — synonym-rich queries + time/date/number
weighting
LLM-generated synonyms: the claim-analysis prompt now instructs the
model to write each next_queries entry as a retrieval-boosted query that
actively folds in entity aliases, DATE/TIME synonyms (e.g. 1994 → 1994,
66th Academy Awards), and number/unit variants (e.g. 1.95 m → 6 ft 5
in).
Time/date/number boosting: query.py boosts numeric/date tokens to a high
weight (_NUM_DATE_TOKEN_RE).