mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-07-24 17:36:47 +08:00
fix(ingestion): align laws DSL with Python — heading fallback, colon-title, short-line filter, remove_toc, and image extension mapping (#17200)
## Summary
This PR aligns the Go ingestion pipeline's **Laws** DSL template with
the Python implementation by fixing heading-detection gaps, adds
image-extension support, refactors the **Extractor** component's LLM
resolution, hardens heading detection for CJK text, and makes the
Extractor accept the Python DSL prompt key names
(`sys_prompt`/`prompts`) alongside the Go names.
## Changes
### 1. Picture file-type detection (`internal/utility/file.go`)
Adds explicit mapping for common image extensions (png, jpg, jpeg, gif,
bmp, tiff, tif, webp, svg, ico, avif, heic, apng) → `FileTypeVISUAL`,
with regression tests.
### 2. Laws DSL heading-detection alignment
(`internal/ingestion/component/chunker/`)
Four fixes to `resolveTitleLevels`:
| Fix | What changed | Why |
|-----|-------------|-----|
| **DOCX `ck_type` fallback** | `ckType` field on `lineRecord`,
propagated from `ChunkDoc.CKType` in `recordsFromStructured`. When
`ck_type=="heading"`, assign `fallbackLevel`. | office_oxide extracts
DOCX heading metadata, but the info was lost before reaching the heading
detector. Word headings whose text doesn't match any regex (e.g.
"Introduction") were treated as body. |
| **`make_colon_as_title` promotion** | `isColonTitle()`: promotes lines
ending with `:`/`:` that have sentence-ending punctuation before the
colon and ≥32 runes between them. | Mirrors Python's
`make_colon_as_title` in `rag/nlp/__init__.py`. Triple guard prevents
false positives. |
| **Short/numeric line filter** | Lines with ≤1 rune or purely numeric
are pinned to body level. | Mirrors Python `tree_merge`'s filter of
`sections` where `len(...) <= 1` or `re.match(r"[0-9]+$", ...)`. |
| **PDF `remove_toc`** | `"remove_toc": true` added to the PDF parser
setup in `ingestion_pipeline_laws.json`. | The Go PDF parser already
supports TOC removal; the Book template already enables it. |
### 3. Extractor llm_id resolution
(`internal/ingestion/component/extractor.go`)
Refactored to handle both **bare tenant_model UUIDs** and **composite
model@provider** strings via the shared `resolveModelConfig`
(`dispatch_model.go`):
- **`resolveExtractorChatConfig`** — UUID path calls
`resolveModelConfigByID` directly (one DB hit); composite path goes
through `resolveModelConfig`. Added `isBareTenantModelID` pre-check for
clear errors when a UUID doesn't exist.
- **`resolveExtractorChatTarget`** — propagates resolution errors
instead of silently returning empty driver.
- **`Chat()`** — removed `driver = "dummy"` fallback. Missing driver is
now an explicit error.
- **Removed dead code**: `splitExtractorLLID`,
`findExtractorSoleActiveInstance`.
### 4. `InjectExtractorLLMID` — fallback when no user config
(`internal/common/parser_config.go`)
Injects the tenant's global default LLM into extractor components **only
when their `llm_id` is empty**. Preserves user-selected UUID or
model@provider values.
Priority: user-configured llm_id > tenant global default > error (no
silent dummy fallback).
### 5. `ResponseHeaderTimeout` increase
(`internal/entity/models/base_model.go`)
`ResponseHeaderTimeout` 60s → 120s in `NewDriverHTTPClient`. Reasoning
models with large extraction prompts can take longer than 60s to produce
the first response token.
### 6. CJK rune-aware heading detection
(`internal/ingestion/component/chunker/title.go`)
Two byte-vs-rune bugs that only manifest on CJK text:
| Fix | What changed | Why |
|-----|-------------|-----|
| **`isColonTitle` byte offset** | `body[lastPunct+1:]` →
`body[lastPunct+runeLen:]` via `utf8.DecodeRuneInString` |
`strings.LastIndexAny` returns a byte index; `+1` skips only 1 byte,
corrupting multi-byte CJK punctuation (e.g. `。` = 3 bytes) and inflating
the rune count past the 32-rune threshold → false-positive heading
promotion. |
| **Short-line filter byte count** | `len(text) <= 1` →
`utf8.RuneCountInString(text) <= 1` | Go `len` is UTF-8 bytes; a single
CJK char (3 bytes) passed the filter, but Python's `len` returns 1 →
mismatch. |
### 7. `extractor_tag.go` — log error when llm fails
When `resolveExtractorChatTarget` returned an error, `runAutoTags` will
log error.
### 8. Python DSL prompt-key compatibility
(`internal/ingestion/component/extractor.go`)
The Resume DSL template uses Python-side key names (`sys_prompt`,
`prompts`). `NewExtractorComponent` now accepts them as fallbacks
alongside the Go names:
- `system_prompt` (Go) ← `sys_prompt` (Python) as fallback
- `prompt` (Go string) ← `prompts` (Python array `[{"role","content"}]`,
takes `[0].content`) as fallback
Mirrors the alias pattern already in `internal/agent/component/llm.go`.
`resolveInputs` accepts per-call `sys_prompt` override too.
## Remaining gaps vs Python
| Gap | Scope | Impact |
|-----|-------|--------|
| **TOC removal for TXT/MD/HTML** | Python's `remove_contents_table`
works on all text formats; Go's `remove_toc` is PDF-only. | Low —
plain-text documents rarely contain structured TOCs. |
| **Regex pattern details** | Minor differences in quantifiers, missing
H5/H6 markdown patterns, missing 4-level numbering pattern. | Low — Go's
variants are stricter; DOCX headings are covered by `ck_type` fallback.
|
## Testing
- `TestHierarchyTitleChunker_CKTypeHeadingFallback` — DOCX `ck_type`
heading promotion
- `TestHierarchyTitleChunker_ColonTitlePromotion` /
`_ColonTitleShortLine_Negative` — colon-title promotion + guard
- `TestIsColonTitle_CJKEdgeCase` / `TestIsColonTitle_ASCII_NoRegression`
— CJK byte-offset fix + ASCII regression
- `TestHierarchyTitleChunker_ColonTitlePromotion_CJK_EdgeCase` — CJK
colon edge case through full pipeline
- `TestHierarchyTitleChunker_ShortSingleCJKLineFilter` — single CJK char
filtered to body
- `TestHierarchyTitleChunker_ShortNumericLineFilter` — purely numeric
lines filtered
- `TestGetFileType_ImageExtensions` / `_ExistingFormats_NoRegression` —
image extension mapping
- `TestInjectExtractorLLMID_SkipWhenUUID` / `_SkipWhenComposite` /
`_InjectWhenEmpty` — llm_id injection guard
- `TestIsBareTenantModelID` — UUID detection
- `TestResolveExtractorChatTarget_AtSplitFallback` / `_NoDriver` — @
split fallback without DB
- `TestNewExtractorComponent_SysPromptAlias` / `_PromptsArray` /
`_PromptsArray_PromptWins` / `_SystemPromptWinsOverSysPrompt` — Python
key compatibility
- `TestBuildDOCXJSONSections_List` / `_TextBox` / `_MixedWithList` —
DOCX list/text_box parsing
- Full ingestion test suite passes (chunker, pipeline, task, service,
component packages)
This commit is contained in:
@@ -139,7 +139,7 @@ func extractDOCXFiguresFromIR(irJSON string) []DOCXFigure {
|
||||
flat = append(flat, flatBlock{image: b64})
|
||||
continue
|
||||
}
|
||||
text := joinDOCXIRRuns(el.Content)
|
||||
text := joinDOCXIRRuns(el.contentRuns())
|
||||
flat = append(flat, flatBlock{text: text})
|
||||
}
|
||||
}
|
||||
@@ -237,6 +237,44 @@ func joinDOCXIRRuns(runs []docxIRRun) string {
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// extractTextFromListItem extracts the plain text content from a list item.
|
||||
// Each list item contains block-level elements (typically a Paragraph),
|
||||
// whose text runs are concatenated.
|
||||
func extractTextFromListItem(item docxIRListItem) string {
|
||||
var parts []string
|
||||
for _, el := range item.Content {
|
||||
if el.Type == "paragraph" || el.Type == "heading" {
|
||||
t := joinDOCXIRRuns(el.contentRuns())
|
||||
if t != "" {
|
||||
parts = append(parts, t)
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(parts) == 0 {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSpace(strings.Join(parts, "\n"))
|
||||
}
|
||||
|
||||
// extractTextFromBlockElements extracts text from a slice of block-level
|
||||
// elements (paragraphs/headings), used by text_box and other compound
|
||||
// element types.
|
||||
func extractTextFromBlockElements(blocks []docxIRElement) string {
|
||||
var parts []string
|
||||
for _, el := range blocks {
|
||||
if el.Type == "paragraph" || el.Type == "heading" {
|
||||
t := joinDOCXIRRuns(el.contentRuns())
|
||||
if t != "" {
|
||||
parts = append(parts, t)
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(parts) == 0 {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSpace(strings.Join(parts, "\n"))
|
||||
}
|
||||
|
||||
// buildFiguresMap converts the internal DOCXFigure slice to the
|
||||
// map form attached to fileMeta["figures"].
|
||||
func buildFiguresMap(figures []DOCXFigure) []map[string]any {
|
||||
@@ -257,7 +295,7 @@ func buildFiguresMap(figures []DOCXFigure) []map[string]any {
|
||||
func joinCellText(cell docxIRCell) string {
|
||||
var parts []string
|
||||
for _, el := range cell.Content {
|
||||
if text := joinDOCXIRRuns(el.Content); text != "" {
|
||||
if text := joinDOCXIRRuns(el.contentRuns()); text != "" {
|
||||
parts = append(parts, text)
|
||||
}
|
||||
}
|
||||
@@ -294,7 +332,7 @@ func buildDOCXJSONSections(irJSON string) []map[string]any {
|
||||
for _, el := range sec.Elements {
|
||||
switch el.Type {
|
||||
case "paragraph", "heading":
|
||||
text := joinDOCXIRRuns(el.Content)
|
||||
text := joinDOCXIRRuns(el.contentRuns())
|
||||
if strings.TrimSpace(text) == "" {
|
||||
continue
|
||||
}
|
||||
@@ -326,6 +364,30 @@ func buildDOCXJSONSections(irJSON string) []map[string]any {
|
||||
"image": nil,
|
||||
"doc_type_kwd": "table",
|
||||
})
|
||||
|
||||
case "list":
|
||||
for _, item := range el.Items {
|
||||
text := extractTextFromListItem(item)
|
||||
if text == "" {
|
||||
continue
|
||||
}
|
||||
sections = append(sections, map[string]any{
|
||||
"text": text,
|
||||
"image": nil,
|
||||
"doc_type_kwd": "text",
|
||||
})
|
||||
}
|
||||
|
||||
case "text_box":
|
||||
text := extractTextFromBlockElements(el.contentBlocks())
|
||||
if text == "" {
|
||||
continue
|
||||
}
|
||||
sections = append(sections, map[string]any{
|
||||
"text": text,
|
||||
"image": nil,
|
||||
"doc_type_kwd": "text",
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -344,12 +406,38 @@ type docxIRSection struct {
|
||||
}
|
||||
|
||||
type docxIRElement struct {
|
||||
Type string `json:"type"` // "paragraph", "heading", "table", "image"
|
||||
Level int `json:"level"` // heading level (1-6)
|
||||
Style string `json:"style"` // Word style name (e.g. "Normal", "Heading 1")
|
||||
Content []docxIRRun `json:"content"` // rich text runs
|
||||
Data []byte `json:"data"` // raw image bytes (for "image" type)
|
||||
Rows []docxIRRow `json:"rows"` // table rows
|
||||
Type string `json:"type"` // "paragraph", "heading", "table", "image", "list", "text_box", ...
|
||||
Level int `json:"level"` // heading level (1-6) or list nesting level
|
||||
Style string `json:"style"` // Word style name (e.g. "Normal", "Heading 1")
|
||||
Content json.RawMessage `json:"content"` // rich text runs or block-level content; decoded per type
|
||||
Data []byte `json:"data"` // raw image bytes (for "image" type)
|
||||
Rows []docxIRRow `json:"rows"` // table rows
|
||||
Ordered bool `json:"ordered"` // true=numbered list, false=bullet list (for "list" type)
|
||||
Items []docxIRListItem `json:"items"` // list items (for "list" type)
|
||||
}
|
||||
|
||||
// contentRuns decodes Content as flat text runs (paragraph/heading type).
|
||||
func (e docxIRElement) contentRuns() []docxIRRun {
|
||||
var runs []docxIRRun
|
||||
if len(e.Content) > 0 {
|
||||
_ = json.Unmarshal(e.Content, &runs)
|
||||
}
|
||||
return runs
|
||||
}
|
||||
|
||||
// contentBlocks decodes Content as block-level elements (text_box type).
|
||||
func (e docxIRElement) contentBlocks() []docxIRElement {
|
||||
var blocks []docxIRElement
|
||||
if len(e.Content) > 0 {
|
||||
_ = json.Unmarshal(e.Content, &blocks)
|
||||
}
|
||||
return blocks
|
||||
}
|
||||
|
||||
// docxIRListItem represents one item in an ordered/unordered list.
|
||||
type docxIRListItem struct {
|
||||
Content []docxIRElement `json:"content"` // block-level content (typically a single Paragraph)
|
||||
Nested json.RawMessage `json:"nested,omitempty"` // nested sub-list (stored as raw JSON for now)
|
||||
}
|
||||
|
||||
type docxIRRun struct {
|
||||
|
||||
@@ -20,10 +20,16 @@ package parser
|
||||
|
||||
import (
|
||||
"encoding/base64"
|
||||
"encoding/json"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func rawJSON(v any) json.RawMessage {
|
||||
data, _ := json.Marshal(v)
|
||||
return json.RawMessage(data)
|
||||
}
|
||||
|
||||
func TestBuildDOCXJSONSections_Paragraphs(t *testing.T) {
|
||||
ir := `{"sections":[{"title":"","elements":[
|
||||
{"type":"paragraph","content":[{"type":"text","text":"Hello world"}],"style":"Normal"}
|
||||
@@ -144,6 +150,65 @@ func TestBuildDOCXJSONSections_MixedContent(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildDOCXJSONSections_List(t *testing.T) {
|
||||
ir := `{"sections":[{"title":"","elements":[
|
||||
{"type":"list","ordered":false,"items":[
|
||||
{"content":[{"type":"paragraph","content":[{"type":"text","text":"Item 1"}]}]},
|
||||
{"content":[{"type":"paragraph","content":[{"type":"text","text":"Item 2"}]}]},
|
||||
{"content":[{"type":"paragraph","content":[{"type":"text","text":"Item 3"}]}]}
|
||||
],"level":0}
|
||||
]}]}`
|
||||
sections := buildDOCXJSONSections(ir)
|
||||
if len(sections) != 3 {
|
||||
t.Fatalf("got %d sections, want 3 (each list item should be a section)", len(sections))
|
||||
}
|
||||
for i, want := range []string{"Item 1", "Item 2", "Item 3"} {
|
||||
if got, _ := sections[i]["text"].(string); got != want {
|
||||
t.Errorf("section[%d].text = %q, want %q", i, got, want)
|
||||
}
|
||||
if got, ok := sections[i]["doc_type_kwd"].(string); !ok || got != "text" {
|
||||
t.Errorf("section[%d].doc_type_kwd = %q, want %q", i, got, "text")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildDOCXJSONSections_TextBox(t *testing.T) {
|
||||
ir := `{"sections":[{"title":"","elements":[
|
||||
{"type":"text_box","content":[{"type":"paragraph","content":[{"type":"text","text":"Boxed paragraph"}]}],"width_emu":null}
|
||||
]}]}`
|
||||
sections := buildDOCXJSONSections(ir)
|
||||
if len(sections) != 1 {
|
||||
t.Fatalf("got %d sections, want 1 (text_box content should become a section)", len(sections))
|
||||
}
|
||||
if got, _ := sections[0]["text"].(string); got != "Boxed paragraph" {
|
||||
t.Errorf("text = %q, want %q", got, "Boxed paragraph")
|
||||
}
|
||||
if got, ok := sections[0]["doc_type_kwd"].(string); !ok || got != "text" {
|
||||
t.Errorf("doc_type_kwd = %q, want %q", got, "text")
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildDOCXJSONSections_MixedWithList(t *testing.T) {
|
||||
ir := `{"sections":[{"title":"","elements":[
|
||||
{"type":"paragraph","content":[{"type":"text","text":"Preamble"}],"style":"Normal"},
|
||||
{"type":"list","ordered":false,"items":[
|
||||
{"content":[{"type":"paragraph","content":[{"type":"text","text":"Bullet A"}]}]},
|
||||
{"content":[{"type":"paragraph","content":[{"type":"text","text":"Bullet B"}]}]}
|
||||
],"level":0},
|
||||
{"type":"paragraph","content":[{"type":"text","text":"Trailer"}],"style":"Normal"}
|
||||
]}]}`
|
||||
sections := buildDOCXJSONSections(ir)
|
||||
if len(sections) != 4 {
|
||||
t.Fatalf("got %d sections, want 4 (para + 2 list items + para)", len(sections))
|
||||
}
|
||||
wants := []string{"Preamble", "Bullet A", "Bullet B", "Trailer"}
|
||||
for i, want := range wants {
|
||||
if got, _ := sections[i]["text"].(string); got != want {
|
||||
t.Errorf("section[%d].text = %q, want %q", i, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildDOCXJSONSections_EmptySkipped(t *testing.T) {
|
||||
ir := `{"sections":[{"title":"","elements":[
|
||||
{"type":"paragraph","content":[{"type":"text","text":""}]},
|
||||
@@ -174,7 +239,7 @@ func TestDocxIRTableToHTML_Single(t *testing.T) {
|
||||
Cells: []docxIRCell{{
|
||||
Content: []docxIRElement{{
|
||||
Type: "paragraph",
|
||||
Content: []docxIRRun{{Type: "text", Text: "hello"}},
|
||||
Content: json.RawMessage(`[{"type":"text","text":"hello"}]`),
|
||||
}},
|
||||
}},
|
||||
}},
|
||||
@@ -191,7 +256,7 @@ func TestDocxIRTableToHTML_MultiRowCol(t *testing.T) {
|
||||
return docxIRCell{
|
||||
Content: []docxIRElement{{
|
||||
Type: "paragraph",
|
||||
Content: []docxIRRun{{Type: "text", Text: text}},
|
||||
Content: rawJSON([]docxIRRun{{Type: "text", Text: text}}),
|
||||
}},
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user