Fix(go): align ingestion pipeline with Python (parser/media dispatch + PDF coordinate chain + Chunker) (#17349)
## Summary
Aligns the Go ingestion pipeline with the Python implementation, closing
several behavioral gaps found during the Python→Go migration (tracked in
`docs/migration_python_go_diff.md`). Covers parser/media dispatch
alignment, the PDF coordinate-chain (preview images, outline→title,
chunk coordinate finalization), and the Chunker Token/QA batches below.
Commits are grouped as follows.
### 1. Fix parser params (c524f450e)
Fixes parser/media wiring and several dispatch gaps:
- **docx/pdf vision dispatch**: correct parameter handling and VLM
invocation.
- **markdown vision (diff 2.5)**: also enhance items whose
`doc_type_kwd` is `table`, not only `image` (parser/utils.py:181).
- **media audio (diff 2.11)**: when `output_format` is `json`, carry the
ASR transcription as a JSON item instead of only the `Text` field (the
Invoke switch had no `json` branch and dropped it).
- **email (diff 2.2)**: default `output_format` is `json`
(parser.py:212), not `text`.
- **tokenizer**: handle empty/whitespace-only names; trim before
embedding.
- **extractor**: tag-matching parameter wiring.
- **split**: keyword-split regex now covers CJK/English separators.
- **parser.go**: parser-param plumbing.
### 2. fix parser gap (373537da1)
Image dispatch now mirrors `rag/app/picture.py:chunk()`:
- Always OCR the image (PaddleOCR or local ONNX).
- When OCR text is short, also call VLM (`describe`) and combine `OCR +
VLM` text.
- Emits a **structured JSON item** carrying the image data-URI and
`doc_type_kwd:"image"`, instead of a bare `Text` string. This fixes the
payload being rejected downstream by OneChunker/TokenChunker (JSON=nil).
### 3. PDF coordinate-chain fixes (55367a820, 727f8167c)
Closes three items from the migration tracker in the
chunker/tokenizer/task layer:
- **(Chunker-1.3) `restore_pdf_text_previews`** — `needsCrop` now also
returns true for `text` chunks that carry PDF positions
(`pdfcrop_cgo.go`), so text blocks get a rendered preview image uploaded
to storage via `imageUploadDecorator`/`ChunkImageUploader`, matching
Python `restore_pdf_text_previews` + `image2id`.
- **(Chunker-1.5) PDF outline → title levels** — `title.go` adds
`outlineSimilarity` (rune-bigram Jaccard, mirroring
`common.py:_outline_similarity`), `resolveOutlineLevels` (matches text
lines to outline entries at similarity > 0.8, with a sparse guard
`len(outline)/len(records) <= 0.03`), and `outlineFromInputs` (reads
`file.outline`). Wired into `newLevelContext` in both `group.go` and
`hierarchy.go`; falls back to the title-shape heuristic when no outline
is present.
- **(Tokenizer-(T)1) `finalize_pdf_chunk`** — the coordinate →
`position_int`/`page_num_int`/`top_int` conversion is owned by the task
layer (`processChunkPositions`→`AddPositions`), which runs *after* the
tokenizer and consumes the tokenizer-owned fields. The tokenizer only
preserves the raw `positions`/`_pdf_positions` (no duplicate
conversion), pinned by `TestChunkDocsToMaps_PreservesPDFPositions`.
### 4. Integration test made environment-free
(`internal/ingestion/task/pipeline_real_integration_test.go`)
- Removed the `//go:build integration` tag so the contract tests run
under the default `build.sh --test` (which does not pass `-tags
integration`).
- External dependencies replaced with in-memory substitutes so no
MySQL/MinIO/ES is required:
- MySQL → on-disk sqlite (`glebarez/sqlite`) with the needed tables
auto-migrated.
- MinIO → `storage.NewMemoryStorage()`.
- Elasticsearch → chunks captured via `WithInsertFunc` instead of
`engine.InsertChunks`/`Search`.
- `requireTokenizerPool` still skips gracefully when the native
tokenizer pool is unavailable; `WithLogCreateFunc(noop)` avoids
depending on the operation-log table.
- Added `taskChunkFieldEqualsStr` to tolerate `kb_id` being a
`[]string`/`[]any` in the raw chunk payload (the search engine flattens
it to a string on read).
### 5. TokenChunker alignment — Batch 1
(`internal/ingestion/component/chunker/token.go`)
Closes four Chunker items from the migration tracker:
- **(Chunker-2.1) sentence delimiter** — the boundary regex now also
breaks on ASCII `!`/`?`. Extracted to a package-level `var
sentenceDelimiter` and used in `mergeByTokenSize`, matching Python's
full delimiter set.
- **(Chunker-2.2) overlap tag leakage** — when a new chunk starts, its
overlap prefix is taken from the previous chunk *after* `removeTag`, in
both the text path (`mergeByTokenSize`) and the JSON path
(`mergeByTokenSizeFromJSON`). Parser tags (`@@…##`) no longer leak into
the overlap region (mirrors `nlp/__init__.py:1181`).
- **(Chunker-2.11) empty-text merge** — merging a non-empty chunk into
an empty previous chunk now assigns the text directly instead of being
skipped (`mergeByTokenSizeFromJSON`), mirroring
`token_chunker.py:236-239`.
- **(Chunker-2.4) overlap token counting** —
`takeFromEnd`/`takeFromStart` now count tokens exactly via `tokenizeStr`
instead of the 4-bytes/token heuristic, fixing over-counting for CJK
text.
### 6. QA Chunker alignment — Batch 2
(`internal/ingestion/component/chunker/qa.go` + `schema`)
Closes three Chunker items from the migration tracker:
- **(Chunker-2.13) default language** — an empty `lang` now defaults to
Chinese prefixes (`问题:`/`回答:`) instead of English, matching `qa.py:299`.
- **(Chunker-2.12) `rmQAPrefix` regex** — the separator is changed to
`[\t:: ]+` (one-or-more), matching `qa.py:241`, so multiple separators
(e.g. `Q:: answer`) are fully stripped.
- **(Chunker-1.8 QA) missing chunk fields** — QA chunks now preserve:
- `top_int` — the source row/record index, threaded through the
tab/csv/markdown extractors (mirrors `qa.py` `beAdoc(..., row_num=i)`);
- `image` + `doc_type_kwd:"image"`;
- `_pdf_positions` / `positions` carried from the upstream JSON item.
`schema.ChunkDoc` gains a `TopInt []int` field (serialized as `top_int`,
registered in `UnmarshalJSON`). Note: the Tag/Table/Presentation/One
chunker field gaps under 1.8 remain pending.
## Test plan
- Added/updated unit tests: `pdfcrop_cgo_test.go` (`TestNeedsCrop`,
`TestRestorePDFTextPreview`), `title_test.go`
(`TestResolveOutlineLevels`, `TestResolveOutlineLevels_SparseGuard`,
`TestNewLevelContext_OutlineBranch`, `TestOutlineFromInputs`),
`tokenizer_unit_test.go` (`TestChunkDocsToMaps_PreservesPDFPositions`),
`token_pdfpos_test.go`.
- **Batch 1** — `token_batch1_test.go`:
`TestSentenceDelimiterMatchesBangAndQuestion`,
`TestMergeByTokenSizeFromJSON_OverlapStripsTags`,
`TestMergeByTokenSizeFromJSON_EmptyPrevKeepsChunk`,
`TestTakeFromEndRespectsTokenCount`,
`TestTakeFromStartRespectsTokenCount`.
- **Batch 2** — `qa_batch2_test.go`:
`TestQAChunker_DefaultLangIsChinese`,
`TestRmQAPrefixStripsMultipleSeparators`, `TestQAChunker_SetsTopInt`,
`TestQAChunker_CarriesImageAndPositions`. Existing `qa_test.go`
expectations were updated to the corrected language default / separator
behavior.
- `pipeline_real_integration_test.go`
(`TestPipelineExecutor_Run_RealCanvasDSL_UsesGeneralPipeline`,
`TestPipelineExecutor_Run_RealPDF_ProducesIndexedChunks`,
`TestRunPipeline_RealPipelineOutput_ProducesIndexFields`) now runs
without any external service.
- `bash build.sh --test ./internal/ingestion/...` passes.
- No files deleted.
2026-07-24 21:06:38 +08:00
|
|
|
|
//
|
|
|
|
|
|
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
|
|
|
|
|
|
//
|
|
|
|
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
|
|
|
|
// you may not use this file except in compliance with the License.
|
|
|
|
|
|
// You may obtain a copy of the License at
|
|
|
|
|
|
//
|
|
|
|
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
|
|
|
|
//
|
|
|
|
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
|
|
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
|
|
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
|
|
|
|
// See the License for the specific language governing permissions and
|
|
|
|
|
|
// limitations under the License.
|
|
|
|
|
|
|
|
|
|
|
|
package component
|
|
|
|
|
|
|
2026-08-11 23:18:04 +09:00
|
|
|
|
import (
|
|
|
|
|
|
"context"
|
|
|
|
|
|
"testing"
|
|
|
|
|
|
|
2026-08-25 13:17:01 +08:00
|
|
|
|
"ragflow/internal/common"
|
2026-08-11 23:18:04 +09:00
|
|
|
|
"ragflow/internal/dao"
|
|
|
|
|
|
"ragflow/internal/entity"
|
|
|
|
|
|
modelModule "ragflow/internal/entity/models"
|
2026-08-25 13:17:01 +08:00
|
|
|
|
"ragflow/internal/ingestion/component/schema"
|
2026-08-11 23:18:04 +09:00
|
|
|
|
"ragflow/internal/utility"
|
|
|
|
|
|
|
|
|
|
|
|
"gorm.io/gorm"
|
|
|
|
|
|
)
|
|
|
|
|
|
|
2026-08-25 13:17:01 +08:00
|
|
|
|
// paddleOCRFakeDriver embeds the ModelDriver interface and only implements
|
|
|
|
|
|
// the OCRFile method needed by dispatchPaddleOCRPdf.
|
|
|
|
|
|
type paddleOCRFakeDriver struct {
|
|
|
|
|
|
modelModule.ModelDriver
|
|
|
|
|
|
text string
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
func (d *paddleOCRFakeDriver) Name() string { return "PaddleOCR" }
|
|
|
|
|
|
|
|
|
|
|
|
func (d *paddleOCRFakeDriver) OCRFile(_ context.Context, _ *string, _ []byte, _ *string, _ *modelModule.APIConfig, _ *modelModule.OCRConfig, _ *common.ModelUsage) (*modelModule.OCRFileResponse, error) {
|
|
|
|
|
|
return &modelModule.OCRFileResponse{Text: &d.text}, nil
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// TestDispatchPaddleOCRPdfLabelsPayloadAsMarkdown guards against the
|
|
|
|
|
|
// format-mismatch bug: PaddleOCR backends always return markdown text via
|
|
|
|
|
|
// OCRFile.Text, so the dispatch result MUST be labelled OutputFormat
|
|
|
|
|
|
// "markdown" regardless of what setup["output_format"] says (the pdf default
|
|
|
|
|
|
// is "json"). If the setup value leaked into OutputFormat, buildParserOutputs
|
|
|
|
|
|
// would emit a nil "json" payload and the downstream TokenChunker would
|
|
|
|
|
|
// consume an empty JSONResult -> "completed with 0 chunks".
|
|
|
|
|
|
func TestDispatchPaddleOCRPdfLabelsPayloadAsMarkdown(t *testing.T) {
|
|
|
|
|
|
orig := resolvePaddleOCRModelForDispatch
|
|
|
|
|
|
t.Cleanup(func() { resolvePaddleOCRModelForDispatch = orig })
|
|
|
|
|
|
|
|
|
|
|
|
md := "## 《道德经》全文及翻译\n\n道可道,非常道。"
|
|
|
|
|
|
resolvePaddleOCRModelForDispatch = func(context.Context, *gorm.DB, string, string) (modelModule.ModelDriver, string, *modelModule.APIConfig, error) {
|
|
|
|
|
|
baseURL := "http://localhost:9380"
|
|
|
|
|
|
return &paddleOCRFakeDriver{text: md}, "ocr-model", &modelModule.APIConfig{BaseURL: &baseURL}, nil
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// The real run had output_format=json in the pdf setup; the payload must
|
|
|
|
|
|
// still be labelled markdown because that is what the backend produced.
|
|
|
|
|
|
res, err := dispatchPaddleOCRPdf(t.Context(), dao.DB, "test.pdf", []byte("%PDF-1.4"), "tenant", schema.ParserSetup{"output_format": "json"}, "some-uuid")
|
|
|
|
|
|
if err != nil {
|
|
|
|
|
|
t.Fatalf("dispatchPaddleOCRPdf: %v", err)
|
|
|
|
|
|
}
|
|
|
|
|
|
if res.OutputFormat != "markdown" {
|
|
|
|
|
|
t.Errorf("OutputFormat = %q, want markdown", res.OutputFormat)
|
|
|
|
|
|
}
|
|
|
|
|
|
if res.Markdown != md {
|
|
|
|
|
|
t.Errorf("Markdown = %q, want %q", res.Markdown, md)
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// Default setup (no output_format key) must behave identically.
|
|
|
|
|
|
res, err = dispatchPaddleOCRPdf(t.Context(), dao.DB, "test.pdf", []byte("%PDF-1.4"), "tenant", nil, "some-uuid")
|
|
|
|
|
|
if err != nil {
|
|
|
|
|
|
t.Fatalf("dispatchPaddleOCRPdf (default setup): %v", err)
|
|
|
|
|
|
}
|
|
|
|
|
|
if res.OutputFormat != "markdown" {
|
|
|
|
|
|
t.Errorf("OutputFormat = %q, want markdown", res.OutputFormat)
|
|
|
|
|
|
}
|
|
|
|
|
|
if res.Markdown != md {
|
|
|
|
|
|
t.Errorf("Markdown = %q, want %q", res.Markdown, md)
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// TestDispatchPaddleOCRPdfEmptyTextFails guards against the silent-empty
|
|
|
|
|
|
// result: OCRFileResponse.Text is a *string that stays non-nil even when the
|
|
|
|
|
|
// backend produced zero text, so the old nil-only guard let an empty payload
|
|
|
|
|
|
// through and the pipeline emitted a "completed with 0 chunks" document.
|
|
|
|
|
|
// Empty text must surface as an explicit error instead.
|
|
|
|
|
|
func TestDispatchPaddleOCRPdfEmptyTextFails(t *testing.T) {
|
|
|
|
|
|
orig := resolvePaddleOCRModelForDispatch
|
|
|
|
|
|
t.Cleanup(func() { resolvePaddleOCRModelForDispatch = orig })
|
|
|
|
|
|
|
|
|
|
|
|
resolvePaddleOCRModelForDispatch = func(context.Context, *gorm.DB, string, string) (modelModule.ModelDriver, string, *modelModule.APIConfig, error) {
|
|
|
|
|
|
baseURL := "http://localhost:9380"
|
|
|
|
|
|
return &paddleOCRFakeDriver{text: ""}, "ocr-model", &modelModule.APIConfig{BaseURL: &baseURL}, nil
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
res, err := dispatchPaddleOCRPdf(t.Context(), dao.DB, "test.pdf", []byte("%PDF-1.4"), "tenant", nil, "some-uuid")
|
|
|
|
|
|
if err == nil {
|
|
|
|
|
|
t.Fatalf("dispatchPaddleOCRPdf with empty text: expected error, got result %+v", res)
|
|
|
|
|
|
}
|
|
|
|
|
|
if res.OutputFormat != "" || res.Markdown != "" {
|
|
|
|
|
|
t.Errorf("expected zero-value result on error, got %+v", res)
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
|
|
|
|
|
|
2026-08-11 23:18:04 +09:00
|
|
|
|
func TestMaybeDispatchPDFVisionEnhancementForwardsDatasetLanguage(t *testing.T) {
|
|
|
|
|
|
origResolver := resolveTenantModelByType
|
|
|
|
|
|
origInvoker := visionChatInvoker
|
|
|
|
|
|
origPrompt := figureVisionPromptBuilder
|
|
|
|
|
|
t.Cleanup(func() {
|
|
|
|
|
|
resolveTenantModelByType = origResolver
|
|
|
|
|
|
visionChatInvoker = origInvoker
|
|
|
|
|
|
figureVisionPromptBuilder = origPrompt
|
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
|
|
resolveTenantModelByType = func(context.Context, *gorm.DB, string, entity.ModelType) (modelModule.ModelDriver, string, *modelModule.APIConfig, int, error) {
|
|
|
|
|
|
return &docxVisionFakeDriver{}, "pdf-vision-model", &modelModule.APIConfig{}, 0, nil
|
|
|
|
|
|
}
|
|
|
|
|
|
invoker := &docxVisionCaptureInvoker{}
|
|
|
|
|
|
visionChatInvoker = invoker.invoke
|
|
|
|
|
|
capturedLanguage := ""
|
|
|
|
|
|
figureVisionPromptBuilder = func(_, _, language string) (string, error) {
|
|
|
|
|
|
capturedLanguage = language
|
|
|
|
|
|
return "describe the figure", nil
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
dispatched := parserDispatchResult{
|
|
|
|
|
|
OutputFormat: "json",
|
|
|
|
|
|
DocType: "pdf",
|
|
|
|
|
|
JSON: []map[string]any{
|
|
|
|
|
|
{"text": "caption", "image": "data:image/png;base64,aW1hZ2U=", "doc_type_kwd": "image"},
|
|
|
|
|
|
},
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
res, modified, err := maybeDispatchPDFVisionEnhancement(
|
|
|
|
|
|
t.Context(),
|
|
|
|
|
|
dao.DB,
|
|
|
|
|
|
utility.FileTypePDF,
|
|
|
|
|
|
dispatched,
|
|
|
|
|
|
map[string]any{"tenant_id": "t1", "lang": "Dutch"},
|
|
|
|
|
|
)
|
|
|
|
|
|
if err != nil {
|
|
|
|
|
|
t.Fatalf("maybeDispatchPDFVisionEnhancement: %v", err)
|
|
|
|
|
|
}
|
|
|
|
|
|
if !modified {
|
|
|
|
|
|
t.Fatal("modified = false, want true")
|
|
|
|
|
|
}
|
|
|
|
|
|
if capturedLanguage != "Dutch" {
|
|
|
|
|
|
t.Fatalf("figure prompt language = %q, want Dutch", capturedLanguage)
|
|
|
|
|
|
}
|
|
|
|
|
|
if got := res.JSON[0]["text"]; got != "caption\na diagram of a pipeline" {
|
|
|
|
|
|
t.Fatalf("enhanced text = %q", got)
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
Fix(go): align ingestion pipeline with Python (parser/media dispatch + PDF coordinate chain + Chunker) (#17349)
## Summary
Aligns the Go ingestion pipeline with the Python implementation, closing
several behavioral gaps found during the Python→Go migration (tracked in
`docs/migration_python_go_diff.md`). Covers parser/media dispatch
alignment, the PDF coordinate-chain (preview images, outline→title,
chunk coordinate finalization), and the Chunker Token/QA batches below.
Commits are grouped as follows.
### 1. Fix parser params (c524f450e)
Fixes parser/media wiring and several dispatch gaps:
- **docx/pdf vision dispatch**: correct parameter handling and VLM
invocation.
- **markdown vision (diff 2.5)**: also enhance items whose
`doc_type_kwd` is `table`, not only `image` (parser/utils.py:181).
- **media audio (diff 2.11)**: when `output_format` is `json`, carry the
ASR transcription as a JSON item instead of only the `Text` field (the
Invoke switch had no `json` branch and dropped it).
- **email (diff 2.2)**: default `output_format` is `json`
(parser.py:212), not `text`.
- **tokenizer**: handle empty/whitespace-only names; trim before
embedding.
- **extractor**: tag-matching parameter wiring.
- **split**: keyword-split regex now covers CJK/English separators.
- **parser.go**: parser-param plumbing.
### 2. fix parser gap (373537da1)
Image dispatch now mirrors `rag/app/picture.py:chunk()`:
- Always OCR the image (PaddleOCR or local ONNX).
- When OCR text is short, also call VLM (`describe`) and combine `OCR +
VLM` text.
- Emits a **structured JSON item** carrying the image data-URI and
`doc_type_kwd:"image"`, instead of a bare `Text` string. This fixes the
payload being rejected downstream by OneChunker/TokenChunker (JSON=nil).
### 3. PDF coordinate-chain fixes (55367a820, 727f8167c)
Closes three items from the migration tracker in the
chunker/tokenizer/task layer:
- **(Chunker-1.3) `restore_pdf_text_previews`** — `needsCrop` now also
returns true for `text` chunks that carry PDF positions
(`pdfcrop_cgo.go`), so text blocks get a rendered preview image uploaded
to storage via `imageUploadDecorator`/`ChunkImageUploader`, matching
Python `restore_pdf_text_previews` + `image2id`.
- **(Chunker-1.5) PDF outline → title levels** — `title.go` adds
`outlineSimilarity` (rune-bigram Jaccard, mirroring
`common.py:_outline_similarity`), `resolveOutlineLevels` (matches text
lines to outline entries at similarity > 0.8, with a sparse guard
`len(outline)/len(records) <= 0.03`), and `outlineFromInputs` (reads
`file.outline`). Wired into `newLevelContext` in both `group.go` and
`hierarchy.go`; falls back to the title-shape heuristic when no outline
is present.
- **(Tokenizer-(T)1) `finalize_pdf_chunk`** — the coordinate →
`position_int`/`page_num_int`/`top_int` conversion is owned by the task
layer (`processChunkPositions`→`AddPositions`), which runs *after* the
tokenizer and consumes the tokenizer-owned fields. The tokenizer only
preserves the raw `positions`/`_pdf_positions` (no duplicate
conversion), pinned by `TestChunkDocsToMaps_PreservesPDFPositions`.
### 4. Integration test made environment-free
(`internal/ingestion/task/pipeline_real_integration_test.go`)
- Removed the `//go:build integration` tag so the contract tests run
under the default `build.sh --test` (which does not pass `-tags
integration`).
- External dependencies replaced with in-memory substitutes so no
MySQL/MinIO/ES is required:
- MySQL → on-disk sqlite (`glebarez/sqlite`) with the needed tables
auto-migrated.
- MinIO → `storage.NewMemoryStorage()`.
- Elasticsearch → chunks captured via `WithInsertFunc` instead of
`engine.InsertChunks`/`Search`.
- `requireTokenizerPool` still skips gracefully when the native
tokenizer pool is unavailable; `WithLogCreateFunc(noop)` avoids
depending on the operation-log table.
- Added `taskChunkFieldEqualsStr` to tolerate `kb_id` being a
`[]string`/`[]any` in the raw chunk payload (the search engine flattens
it to a string on read).
### 5. TokenChunker alignment — Batch 1
(`internal/ingestion/component/chunker/token.go`)
Closes four Chunker items from the migration tracker:
- **(Chunker-2.1) sentence delimiter** — the boundary regex now also
breaks on ASCII `!`/`?`. Extracted to a package-level `var
sentenceDelimiter` and used in `mergeByTokenSize`, matching Python's
full delimiter set.
- **(Chunker-2.2) overlap tag leakage** — when a new chunk starts, its
overlap prefix is taken from the previous chunk *after* `removeTag`, in
both the text path (`mergeByTokenSize`) and the JSON path
(`mergeByTokenSizeFromJSON`). Parser tags (`@@…##`) no longer leak into
the overlap region (mirrors `nlp/__init__.py:1181`).
- **(Chunker-2.11) empty-text merge** — merging a non-empty chunk into
an empty previous chunk now assigns the text directly instead of being
skipped (`mergeByTokenSizeFromJSON`), mirroring
`token_chunker.py:236-239`.
- **(Chunker-2.4) overlap token counting** —
`takeFromEnd`/`takeFromStart` now count tokens exactly via `tokenizeStr`
instead of the 4-bytes/token heuristic, fixing over-counting for CJK
text.
### 6. QA Chunker alignment — Batch 2
(`internal/ingestion/component/chunker/qa.go` + `schema`)
Closes three Chunker items from the migration tracker:
- **(Chunker-2.13) default language** — an empty `lang` now defaults to
Chinese prefixes (`问题:`/`回答:`) instead of English, matching `qa.py:299`.
- **(Chunker-2.12) `rmQAPrefix` regex** — the separator is changed to
`[\t:: ]+` (one-or-more), matching `qa.py:241`, so multiple separators
(e.g. `Q:: answer`) are fully stripped.
- **(Chunker-1.8 QA) missing chunk fields** — QA chunks now preserve:
- `top_int` — the source row/record index, threaded through the
tab/csv/markdown extractors (mirrors `qa.py` `beAdoc(..., row_num=i)`);
- `image` + `doc_type_kwd:"image"`;
- `_pdf_positions` / `positions` carried from the upstream JSON item.
`schema.ChunkDoc` gains a `TopInt []int` field (serialized as `top_int`,
registered in `UnmarshalJSON`). Note: the Tag/Table/Presentation/One
chunker field gaps under 1.8 remain pending.
## Test plan
- Added/updated unit tests: `pdfcrop_cgo_test.go` (`TestNeedsCrop`,
`TestRestorePDFTextPreview`), `title_test.go`
(`TestResolveOutlineLevels`, `TestResolveOutlineLevels_SparseGuard`,
`TestNewLevelContext_OutlineBranch`, `TestOutlineFromInputs`),
`tokenizer_unit_test.go` (`TestChunkDocsToMaps_PreservesPDFPositions`),
`token_pdfpos_test.go`.
- **Batch 1** — `token_batch1_test.go`:
`TestSentenceDelimiterMatchesBangAndQuestion`,
`TestMergeByTokenSizeFromJSON_OverlapStripsTags`,
`TestMergeByTokenSizeFromJSON_EmptyPrevKeepsChunk`,
`TestTakeFromEndRespectsTokenCount`,
`TestTakeFromStartRespectsTokenCount`.
- **Batch 2** — `qa_batch2_test.go`:
`TestQAChunker_DefaultLangIsChinese`,
`TestRmQAPrefixStripsMultipleSeparators`, `TestQAChunker_SetsTopInt`,
`TestQAChunker_CarriesImageAndPositions`. Existing `qa_test.go`
expectations were updated to the corrected language default / separator
behavior.
- `pipeline_real_integration_test.go`
(`TestPipelineExecutor_Run_RealCanvasDSL_UsesGeneralPipeline`,
`TestPipelineExecutor_Run_RealPDF_ProducesIndexedChunks`,
`TestRunPipeline_RealPipelineOutput_ProducesIndexFields`) now runs
without any external service.
- `bash build.sh --test ./internal/ingestion/...` passes.
- No files deleted.
2026-07-24 21:06:38 +08:00
|
|
|
|
|
|
|
|
|
|
// TestIsNamedPDFParseMethodWhitelistAligned verifies that the runtime
|
|
|
|
|
|
// "named parse_method" classifier agrees with (*ParserComponent).Check()'s
|
|
|
|
|
|
// PDF whitelist (parser.go:200-203):
|
|
|
|
|
|
//
|
|
|
|
|
|
// deepdoc, plain_text, mineru, docling,
|
|
|
|
|
|
// opendataloader, tcadp parser, paddleocr, somark
|
|
|
|
|
|
//
|
|
|
|
|
|
// Diff 2.10: a parse_method that Check() rejects must NOT be treated as a
|
|
|
|
|
|
// recognized named method by isNamedPDFParseMethod — otherwise it silently
|
|
|
|
|
|
// falls through to the CustomVLM vision path instead of failing fast at
|
|
|
|
|
|
// construction (and Python would have rejected it outright).
|
|
|
|
|
|
func TestIsNamedPDFParseMethodWhitelistAligned(t *testing.T) {
|
|
|
|
|
|
// Values that MUST be recognized (subset of the Check() whitelist,
|
|
|
|
|
|
// case-insensitive).
|
|
|
|
|
|
named := []string{
|
|
|
|
|
|
"deepdoc", "plain_text", "mineru", "docling",
|
|
|
|
|
|
"opendataloader", "tcadp parser", "paddleocr", "somark",
|
|
|
|
|
|
"DeepDoc", "PLAIN_TEXT", "MinerU", "DocLing",
|
|
|
|
|
|
"OpenDataLoader", "TCADP Parser", "PaddleOCR", "SoMark",
|
|
|
|
|
|
}
|
|
|
|
|
|
for _, v := range named {
|
|
|
|
|
|
if !isNamedPDFParseMethod(v) {
|
|
|
|
|
|
t.Errorf("isNamedPDFParseMethod(%q) = false, want true (in Check() whitelist)", v)
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// Values that MUST NOT be recognized. These either duplicate the
|
|
|
|
|
|
// whitelist with non-canonical spelling ("plain text"/"plaintext")
|
|
|
|
|
|
// or are bare-family abbreviations ("tcadp") that Check() does not
|
|
|
|
|
|
// accept, so they should be funneled to the CustomVLM path (or fail
|
|
|
|
|
|
// construction) rather than masquerading as a named method.
|
|
|
|
|
|
notNamed := []string{
|
|
|
|
|
|
"plain text", "plaintext", "tcadp",
|
|
|
|
|
|
"CustomVLM", "some_vlm", "gpt-4o",
|
|
|
|
|
|
"", " ",
|
|
|
|
|
|
}
|
|
|
|
|
|
for _, v := range notNamed {
|
|
|
|
|
|
if isNamedPDFParseMethod(v) {
|
|
|
|
|
|
t.Errorf("isNamedPDFParseMethod(%q) = true, want false (not in Check() whitelist)", v)
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// TestIsNamedPDFParseMethodLayoutSuffixes verifies that "@"-suffixed
|
|
|
|
|
|
// layout_recognizer spellings are NOT treated as named parse methods. They
|
|
|
|
|
|
// are layout_recognizer selectors (resolved separately at
|
|
|
|
|
|
// pdf_vision_dispatch.go:62-68), and Check() rejects them as parse_method,
|
|
|
|
|
|
// so they must fall through to the CustomVLM/VLM path — consistent with the
|
|
|
|
|
|
// (*ParserComponent).Check() whitelist (parser.go:200-203).
|
|
|
|
|
|
func TestIsNamedPDFParseMethodLayoutSuffixes(t *testing.T) {
|
|
|
|
|
|
suffixed := []string{
|
|
|
|
|
|
"foo@mineru", "@mineru",
|
|
|
|
|
|
"foo@paddleocr", "@paddleocr",
|
|
|
|
|
|
"foo@somark", "@somark",
|
|
|
|
|
|
"foo@opendataloader", "@opendataloader",
|
|
|
|
|
|
}
|
|
|
|
|
|
for _, v := range suffixed {
|
|
|
|
|
|
if isNamedPDFParseMethod(v) {
|
|
|
|
|
|
t.Errorf("isNamedPDFParseMethod(%q) = true, want false (layout_recognizer selector, not a named parse_method)", v)
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
// An unknown suffix is also not a named method.
|
|
|
|
|
|
if isNamedPDFParseMethod("foo@unknown") {
|
|
|
|
|
|
t.Errorf("isNamedPDFParseMethod(%q) = true, want false", "foo@unknown")
|
|
|
|
|
|
}
|
|
|
|
|
|
}
|