Files
ragflow/internal/ingestion/component/chunker/qa_batch2_test.go
Jack 554925b583 Fix(go): align ingestion pipeline with Python (parser/media dispatch + PDF coordinate chain + Chunker) (#17349)
## Summary
Aligns the Go ingestion pipeline with the Python implementation, closing
several behavioral gaps found during the Python→Go migration (tracked in
`docs/migration_python_go_diff.md`). Covers parser/media dispatch
alignment, the PDF coordinate-chain (preview images, outline→title,
chunk coordinate finalization), and the Chunker Token/QA batches below.

Commits are grouped as follows.

### 1. Fix parser params (c524f450e)
Fixes parser/media wiring and several dispatch gaps:
- **docx/pdf vision dispatch**: correct parameter handling and VLM
invocation.
- **markdown vision (diff 2.5)**: also enhance items whose
`doc_type_kwd` is `table`, not only `image` (parser/utils.py:181).
- **media audio (diff 2.11)**: when `output_format` is `json`, carry the
ASR transcription as a JSON item instead of only the `Text` field (the
Invoke switch had no `json` branch and dropped it).
- **email (diff 2.2)**: default `output_format` is `json`
(parser.py:212), not `text`.
- **tokenizer**: handle empty/whitespace-only names; trim before
embedding.
- **extractor**: tag-matching parameter wiring.
- **split**: keyword-split regex now covers CJK/English separators.
- **parser.go**: parser-param plumbing.

### 2. fix parser gap (373537da1)
Image dispatch now mirrors `rag/app/picture.py:chunk()`:
- Always OCR the image (PaddleOCR or local ONNX).
- When OCR text is short, also call VLM (`describe`) and combine `OCR +
VLM` text.
- Emits a **structured JSON item** carrying the image data-URI and
`doc_type_kwd:"image"`, instead of a bare `Text` string. This fixes the
payload being rejected downstream by OneChunker/TokenChunker (JSON=nil).

### 3. PDF coordinate-chain fixes (55367a820, 727f8167c)
Closes three items from the migration tracker in the
chunker/tokenizer/task layer:

- **(Chunker-1.3) `restore_pdf_text_previews`** — `needsCrop` now also
returns true for `text` chunks that carry PDF positions
(`pdfcrop_cgo.go`), so text blocks get a rendered preview image uploaded
to storage via `imageUploadDecorator`/`ChunkImageUploader`, matching
Python `restore_pdf_text_previews` + `image2id`.

- **(Chunker-1.5) PDF outline → title levels** — `title.go` adds
`outlineSimilarity` (rune-bigram Jaccard, mirroring
`common.py:_outline_similarity`), `resolveOutlineLevels` (matches text
lines to outline entries at similarity > 0.8, with a sparse guard
`len(outline)/len(records) <= 0.03`), and `outlineFromInputs` (reads
`file.outline`). Wired into `newLevelContext` in both `group.go` and
`hierarchy.go`; falls back to the title-shape heuristic when no outline
is present.

- **(Tokenizer-(T)1) `finalize_pdf_chunk`** — the coordinate →
`position_int`/`page_num_int`/`top_int` conversion is owned by the task
layer (`processChunkPositions`→`AddPositions`), which runs *after* the
tokenizer and consumes the tokenizer-owned fields. The tokenizer only
preserves the raw `positions`/`_pdf_positions` (no duplicate
conversion), pinned by `TestChunkDocsToMaps_PreservesPDFPositions`.

### 4. Integration test made environment-free
(`internal/ingestion/task/pipeline_real_integration_test.go`)
- Removed the `//go:build integration` tag so the contract tests run
under the default `build.sh --test` (which does not pass `-tags
integration`).
- External dependencies replaced with in-memory substitutes so no
MySQL/MinIO/ES is required:
- MySQL → on-disk sqlite (`glebarez/sqlite`) with the needed tables
auto-migrated.
  - MinIO → `storage.NewMemoryStorage()`.
- Elasticsearch → chunks captured via `WithInsertFunc` instead of
`engine.InsertChunks`/`Search`.
- `requireTokenizerPool` still skips gracefully when the native
tokenizer pool is unavailable; `WithLogCreateFunc(noop)` avoids
depending on the operation-log table.
- Added `taskChunkFieldEqualsStr` to tolerate `kb_id` being a
`[]string`/`[]any` in the raw chunk payload (the search engine flattens
it to a string on read).

### 5. TokenChunker alignment — Batch 1
(`internal/ingestion/component/chunker/token.go`)
Closes four Chunker items from the migration tracker:
- **(Chunker-2.1) sentence delimiter** — the boundary regex now also
breaks on ASCII `!`/`?`. Extracted to a package-level `var
sentenceDelimiter` and used in `mergeByTokenSize`, matching Python's
full delimiter set.
- **(Chunker-2.2) overlap tag leakage** — when a new chunk starts, its
overlap prefix is taken from the previous chunk *after* `removeTag`, in
both the text path (`mergeByTokenSize`) and the JSON path
(`mergeByTokenSizeFromJSON`). Parser tags (`@@…##`) no longer leak into
the overlap region (mirrors `nlp/__init__.py:1181`).
- **(Chunker-2.11) empty-text merge** — merging a non-empty chunk into
an empty previous chunk now assigns the text directly instead of being
skipped (`mergeByTokenSizeFromJSON`), mirroring
`token_chunker.py:236-239`.
- **(Chunker-2.4) overlap token counting** —
`takeFromEnd`/`takeFromStart` now count tokens exactly via `tokenizeStr`
instead of the 4-bytes/token heuristic, fixing over-counting for CJK
text.

### 6. QA Chunker alignment — Batch 2
(`internal/ingestion/component/chunker/qa.go` + `schema`)
Closes three Chunker items from the migration tracker:
- **(Chunker-2.13) default language** — an empty `lang` now defaults to
Chinese prefixes (`问题:`/`回答:`) instead of English, matching `qa.py:299`.
- **(Chunker-2.12) `rmQAPrefix` regex** — the separator is changed to
`[\t:: ]+` (one-or-more), matching `qa.py:241`, so multiple separators
(e.g. `Q:: answer`) are fully stripped.
- **(Chunker-1.8 QA) missing chunk fields** — QA chunks now preserve:
- `top_int` — the source row/record index, threaded through the
tab/csv/markdown extractors (mirrors `qa.py` `beAdoc(..., row_num=i)`);
  - `image` + `doc_type_kwd:"image"`;
  - `_pdf_positions` / `positions` carried from the upstream JSON item.
`schema.ChunkDoc` gains a `TopInt []int` field (serialized as `top_int`,
registered in `UnmarshalJSON`). Note: the Tag/Table/Presentation/One
chunker field gaps under 1.8 remain pending.

## Test plan
- Added/updated unit tests: `pdfcrop_cgo_test.go` (`TestNeedsCrop`,
`TestRestorePDFTextPreview`), `title_test.go`
(`TestResolveOutlineLevels`, `TestResolveOutlineLevels_SparseGuard`,
`TestNewLevelContext_OutlineBranch`, `TestOutlineFromInputs`),
`tokenizer_unit_test.go` (`TestChunkDocsToMaps_PreservesPDFPositions`),
`token_pdfpos_test.go`.
- **Batch 1** — `token_batch1_test.go`:
`TestSentenceDelimiterMatchesBangAndQuestion`,
`TestMergeByTokenSizeFromJSON_OverlapStripsTags`,
`TestMergeByTokenSizeFromJSON_EmptyPrevKeepsChunk`,
`TestTakeFromEndRespectsTokenCount`,
`TestTakeFromStartRespectsTokenCount`.
- **Batch 2** — `qa_batch2_test.go`:
`TestQAChunker_DefaultLangIsChinese`,
`TestRmQAPrefixStripsMultipleSeparators`, `TestQAChunker_SetsTopInt`,
`TestQAChunker_CarriesImageAndPositions`. Existing `qa_test.go`
expectations were updated to the corrected language default / separator
behavior.
- `pipeline_real_integration_test.go`
(`TestPipelineExecutor_Run_RealCanvasDSL_UsesGeneralPipeline`,
`TestPipelineExecutor_Run_RealPDF_ProducesIndexedChunks`,
`TestRunPipeline_RealPipelineOutput_ProducesIndexFields`) now runs
without any external service.
- `bash build.sh --test ./internal/ingestion/...` passes.
- No files deleted.
2026-07-24 21:06:38 +08:00

148 lines
4.7 KiB
Go
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//
package chunker
import (
"context"
"strings"
"testing"
)
// qaInvoke is a small helper that runs the QA chunker on a upstream-style
// input map and returns the produced chunks as generic maps.
func qaInvoke(t *testing.T, inputs map[string]any) []map[string]any {
t.Helper()
c, err := NewQAChunker(map[string]any{})
if err != nil {
t.Fatalf("NewQAChunker: %v", err)
}
out, err := c.Invoke(context.Background(), inputs)
if err != nil {
t.Fatalf("Invoke: %v", err)
}
chunks, ok := out["chunks"].([]map[string]any)
if !ok {
t.Fatalf("Invoke did not return []map chunks: %#v", out["chunks"])
}
return chunks
}
// TestQAChunker_DefaultLangIsChinese exercises migration diff Chunker-2.13:
// when no language is supplied, Python defaults to Chinese prefixes
// ("问题:"/"回答:"); the legacy Go code defaulted to English.
func TestQAChunker_DefaultLangIsChinese(t *testing.T) {
inputs := map[string]any{
"name": "test.txt",
"output_format": "text",
"text": "What is RAG?\tRAG retrieves then generates.",
}
chunks := qaInvoke(t, inputs)
if len(chunks) == 0 {
t.Fatalf("no chunks produced")
}
cw := chunks[0]["content_with_weight"].(string)
if !contains(cw, "问题:") || !contains(cw, "回答:") {
t.Errorf("empty lang should default to Chinese prefixes, got %q", cw)
}
if contains(cw, "Question:") || contains(cw, "Answer:") {
t.Errorf("empty lang must not use English prefixes, got %q", cw)
}
}
// TestRmQAPrefixStripsMultipleSeparators exercises migration diff
// Chunker-2.12: the prefix regex must allow one-or-more separator chars
// (Python uses `[\t: ]+`), so "Q:: answer" is fully stripped. The legacy
// Go pattern only matched a single separator, leaving ": answer".
func TestRmQAPrefixStripsMultipleSeparators(t *testing.T) {
if got := rmQAPrefix("Q:: answer"); got != "answer" {
t.Errorf("multi-separator prefix not fully stripped: got %q", got)
}
if got := rmQAPrefix("Question: foo"); got != "foo" {
t.Errorf("single-separator prefix regression: got %q", got)
}
if got := rmQAPrefix("问:答案在此"); got != "答案在此" {
t.Errorf("CJK prefix regression: got %q", got)
}
}
// TestQAChunker_SetsTopInt exercises migration diff Chunker-1.8 (top_int):
// each QA chunk must carry the source row index in `top_int`, matching
// Python beAdoc(..., row_num=i).
func TestQAChunker_SetsTopInt(t *testing.T) {
inputs := map[string]any{
"name": "test.txt",
"output_format": "text",
"text": "Q1\tA1\nQ2\tA2",
}
chunks := qaInvoke(t, inputs)
if len(chunks) != 2 {
t.Fatalf("want 2 QA chunks, got %d", len(chunks))
}
for i, c := range chunks {
raw, ok := c["top_int"]
if !ok {
t.Fatalf("chunk %d missing top_int field", i)
}
arr, ok := raw.([]any)
if !ok || len(arr) != 1 {
t.Fatalf("chunk %d top_int wrong shape: %#v", i, raw)
}
if int(arr[0].(float64)) != i {
t.Errorf("chunk %d top_int = %v, want %d", i, arr[0], i)
}
}
}
// TestQAChunker_CarriesImageAndPositions exercises migration diff
// Chunker-1.8 (image / positions): when the upstream JSON item already
// carries an image id and pdf positions, the QA chunk must preserve them
// (Python beAdocPdf sets d["image"] and add_positions).
func TestQAChunker_CarriesImageAndPositions(t *testing.T) {
inputs := map[string]any{
"name": "test.pdf",
"output_format": "json",
"json": []map[string]any{
{
"text": "Q\tA",
"image": "img-42",
"_pdf_positions": [][]int{{1, 2, 3, 4, 5}},
},
},
}
chunks := qaInvoke(t, inputs)
if len(chunks) != 1 {
t.Fatalf("want 1 QA chunk, got %d", len(chunks))
}
c := chunks[0]
if c["image"] != "img-42" {
t.Errorf("QA chunk lost upstream image: %#v", c["image"])
}
if _, ok := c["_pdf_positions"]; !ok {
t.Errorf("QA chunk lost upstream _pdf_positions")
}
// The prefix-stripped content must still be present.
cw, _ := c["content_with_weight"].(string)
if !contains(cw, "A") {
t.Errorf("QA content missing answer: %q", cw)
}
}
// contains is a tiny helper to avoid importing strings in every test.
func contains(s, sub string) bool {
return strings.Contains(s, sub)
}