Files
ragflow/internal/utility/file_test.go

129 lines
3.8 KiB
Go
Raw Normal View History

fix(ingestion): align laws DSL with Python — heading fallback, colon-title, short-line filter, remove_toc, and image extension mapping (#17200) ## Summary This PR aligns the Go ingestion pipeline's **Laws** DSL template with the Python implementation by fixing heading-detection gaps, adds image-extension support, refactors the **Extractor** component's LLM resolution, hardens heading detection for CJK text, and makes the Extractor accept the Python DSL prompt key names (`sys_prompt`/`prompts`) alongside the Go names. ## Changes ### 1. Picture file-type detection (`internal/utility/file.go`) Adds explicit mapping for common image extensions (png, jpg, jpeg, gif, bmp, tiff, tif, webp, svg, ico, avif, heic, apng) → `FileTypeVISUAL`, with regression tests. ### 2. Laws DSL heading-detection alignment (`internal/ingestion/component/chunker/`) Four fixes to `resolveTitleLevels`: | Fix | What changed | Why | |-----|-------------|-----| | **DOCX `ck_type` fallback** | `ckType` field on `lineRecord`, propagated from `ChunkDoc.CKType` in `recordsFromStructured`. When `ck_type=="heading"`, assign `fallbackLevel`. | office_oxide extracts DOCX heading metadata, but the info was lost before reaching the heading detector. Word headings whose text doesn't match any regex (e.g. "Introduction") were treated as body. | | **`make_colon_as_title` promotion** | `isColonTitle()`: promotes lines ending with `:`/`:` that have sentence-ending punctuation before the colon and ≥32 runes between them. | Mirrors Python's `make_colon_as_title` in `rag/nlp/__init__.py`. Triple guard prevents false positives. | | **Short/numeric line filter** | Lines with ≤1 rune or purely numeric are pinned to body level. | Mirrors Python `tree_merge`'s filter of `sections` where `len(...) <= 1` or `re.match(r"[0-9]+$", ...)`. | | **PDF `remove_toc`** | `"remove_toc": true` added to the PDF parser setup in `ingestion_pipeline_laws.json`. | The Go PDF parser already supports TOC removal; the Book template already enables it. | ### 3. Extractor llm_id resolution (`internal/ingestion/component/extractor.go`) Refactored to handle both **bare tenant_model UUIDs** and **composite model@provider** strings via the shared `resolveModelConfig` (`dispatch_model.go`): - **`resolveExtractorChatConfig`** — UUID path calls `resolveModelConfigByID` directly (one DB hit); composite path goes through `resolveModelConfig`. Added `isBareTenantModelID` pre-check for clear errors when a UUID doesn't exist. - **`resolveExtractorChatTarget`** — propagates resolution errors instead of silently returning empty driver. - **`Chat()`** — removed `driver = "dummy"` fallback. Missing driver is now an explicit error. - **Removed dead code**: `splitExtractorLLID`, `findExtractorSoleActiveInstance`. ### 4. `InjectExtractorLLMID` — fallback when no user config (`internal/common/parser_config.go`) Injects the tenant's global default LLM into extractor components **only when their `llm_id` is empty**. Preserves user-selected UUID or model@provider values. Priority: user-configured llm_id > tenant global default > error (no silent dummy fallback). ### 5. `ResponseHeaderTimeout` increase (`internal/entity/models/base_model.go`) `ResponseHeaderTimeout` 60s → 120s in `NewDriverHTTPClient`. Reasoning models with large extraction prompts can take longer than 60s to produce the first response token. ### 6. CJK rune-aware heading detection (`internal/ingestion/component/chunker/title.go`) Two byte-vs-rune bugs that only manifest on CJK text: | Fix | What changed | Why | |-----|-------------|-----| | **`isColonTitle` byte offset** | `body[lastPunct+1:]` → `body[lastPunct+runeLen:]` via `utf8.DecodeRuneInString` | `strings.LastIndexAny` returns a byte index; `+1` skips only 1 byte, corrupting multi-byte CJK punctuation (e.g. `。` = 3 bytes) and inflating the rune count past the 32-rune threshold → false-positive heading promotion. | | **Short-line filter byte count** | `len(text) <= 1` → `utf8.RuneCountInString(text) <= 1` | Go `len` is UTF-8 bytes; a single CJK char (3 bytes) passed the filter, but Python's `len` returns 1 → mismatch. | ### 7. `extractor_tag.go` — log error when llm fails When `resolveExtractorChatTarget` returned an error, `runAutoTags` will log error. ### 8. Python DSL prompt-key compatibility (`internal/ingestion/component/extractor.go`) The Resume DSL template uses Python-side key names (`sys_prompt`, `prompts`). `NewExtractorComponent` now accepts them as fallbacks alongside the Go names: - `system_prompt` (Go) ← `sys_prompt` (Python) as fallback - `prompt` (Go string) ← `prompts` (Python array `[{"role","content"}]`, takes `[0].content`) as fallback Mirrors the alias pattern already in `internal/agent/component/llm.go`. `resolveInputs` accepts per-call `sys_prompt` override too. ## Remaining gaps vs Python | Gap | Scope | Impact | |-----|-------|--------| | **TOC removal for TXT/MD/HTML** | Python's `remove_contents_table` works on all text formats; Go's `remove_toc` is PDF-only. | Low — plain-text documents rarely contain structured TOCs. | | **Regex pattern details** | Minor differences in quantifiers, missing H5/H6 markdown patterns, missing 4-level numbering pattern. | Low — Go's variants are stricter; DOCX headings are covered by `ck_type` fallback. | ## Testing - `TestHierarchyTitleChunker_CKTypeHeadingFallback` — DOCX `ck_type` heading promotion - `TestHierarchyTitleChunker_ColonTitlePromotion` / `_ColonTitleShortLine_Negative` — colon-title promotion + guard - `TestIsColonTitle_CJKEdgeCase` / `TestIsColonTitle_ASCII_NoRegression` — CJK byte-offset fix + ASCII regression - `TestHierarchyTitleChunker_ColonTitlePromotion_CJK_EdgeCase` — CJK colon edge case through full pipeline - `TestHierarchyTitleChunker_ShortSingleCJKLineFilter` — single CJK char filtered to body - `TestHierarchyTitleChunker_ShortNumericLineFilter` — purely numeric lines filtered - `TestGetFileType_ImageExtensions` / `_ExistingFormats_NoRegression` — image extension mapping - `TestInjectExtractorLLMID_SkipWhenUUID` / `_SkipWhenComposite` / `_InjectWhenEmpty` — llm_id injection guard - `TestIsBareTenantModelID` — UUID detection - `TestResolveExtractorChatTarget_AtSplitFallback` / `_NoDriver` — @ split fallback without DB - `TestNewExtractorComponent_SysPromptAlias` / `_PromptsArray` / `_PromptsArray_PromptWins` / `_SystemPromptWinsOverSysPrompt` — Python key compatibility - `TestBuildDOCXJSONSections_List` / `_TextBox` / `_MixedWithList` — DOCX list/text_box parsing - Full ingestion test suite passes (chunker, pipeline, task, service, component packages)
2026-07-22 19:14:32 +08:00
//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//
package utility
import (
"testing"
)
func TestGetFileType_ImageExtensions(t *testing.T) {
cases := []struct {
name string
filename string
want FileType
}{
// Image formats that should resolve to VISUAL.
{"png", "image.png", FileTypeVISUAL},
{"jpg", "photo.jpg", FileTypeVISUAL},
{"jpeg", "photo.jpeg", FileTypeVISUAL},
{"gif", "animation.gif", FileTypeVISUAL},
{"bmp", "bitmap.bmp", FileTypeVISUAL},
{"tiff", "image.tiff", FileTypeVISUAL},
{"tif", "image.tif", FileTypeVISUAL},
{"webp", "image.webp", FileTypeVISUAL},
{"svg", "vector.svg", FileTypeVISUAL},
{"ico", "icon.ico", FileTypeVISUAL},
{"avif", "image.avif", FileTypeVISUAL},
{"heic", "photo.heic", FileTypeVISUAL},
{"apng", "animation.apng", FileTypeVISUAL},
// Capitalised extension should still match.
{"png uppercase", "image.PNG", FileTypeVISUAL},
{"jpg mixed case", "photo.JpG", FileTypeVISUAL},
// Path with directory should still resolve.
{"png with path", "/path/to/image.png", FileTypeVISUAL},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := GetFileType(tc.filename)
if got != tc.want {
t.Errorf("GetFileType(%q) = %q, want %q", tc.filename, got, tc.want)
}
})
}
}
func TestGetFileType_ExistingFormats_NoRegression(t *testing.T) {
cases := []struct {
name string
filename string
want FileType
}{
{"pdf", "doc.pdf", FileTypePDF},
{"doc", "old.doc", FileTypeDOC},
{"docx", "report.docx", FileTypeDOCX},
{"xls", "spreadsheet.xls", FileTypeXLS},
{"xlsx", "spreadsheet.xlsx", FileTypeXLSX},
{"csv", "data.csv", FileTypeCSV},
{"ppt", "slides.ppt", FileTypePPT},
{"pptx", "slides.pptx", FileTypePPTX},
{"html", "page.html", FileTypeHTML},
{"htm", "page.htm", FileTypeHTML},
{"md", "readme.md", FileTypeMarkdown},
{"markdown", "readme.markdown", FileTypeMarkdown},
{"txt", "notes.txt", FileTypeTXT},
{"py", "script.py", FileTypeTXT},
{"js", "script.js", FileTypeTXT},
{"go", "main.go", FileTypeTXT},
{"java", "Main.java", FileTypeTXT},
{"epub", "book.epub", FileTypeEPUB},
{"json", "data.json", FileTypeJSON},
{"jsonl", "data.jsonl", FileTypeJSON},
{"eml", "email.eml", FileTypeEMAIL},
{"msg", "email.msg", FileTypeEMAIL},
{"mp3", "audio.mp3", FileTypeAURAL},
{"wav", "audio.wav", FileTypeAURAL},
{"flac", "audio.flac", FileTypeAURAL},
{"mp4", "video.mp4", FileTypeVIDEO},
{"avi", "video.avi", FileTypeVIDEO},
{"mkv", "video.mkv", FileTypeVIDEO},
{"mov", "video.mov", FileTypeVIDEO},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := GetFileType(tc.filename)
if got != tc.want {
t.Errorf("GetFileType(%q) = %q, want %q (regression)", tc.filename, got, tc.want)
}
})
}
}
func TestGetFileType_UnknownExtension_ReturnsOTHER(t *testing.T) {
cases := []struct {
name string
filename string
}{
{"no extension", "Makefile"},
{"unknown extension", "data.xyz"},
{"dotfile", ".hidden"},
{"empty string", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := GetFileType(tc.filename)
if got != FileTypeOTHER {
t.Errorf("GetFileType(%q) = %q, want %q (FileTypeOTHER)", tc.filename, got, FileTypeOTHER)
}
})
}
}