mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-07-30 20:49:21 +08:00
989529c00a2f1cf4f3f797aefbc7b87ec437e19d
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9b0719fa94 |
fix: Go ingestion migration batch 5 (Parser 1.1/1.7/2.11, Chunker 1.7/1.8/2.6/2.7, Tokenizer 6x fixes) (#17419)
## Summary
Continuation of the Python→Go ingestion pipeline migration (File →
Parser → Chunker → Extractor → Tokenizer). Fixes cover Parser, Chunker,
and Tokenizer gaps identified. Fix page number (0-indexed and 1-index
mixed before fix; use 1-indexed after fix) and chunk order issues.
### Parser
- **Slides TCADP (1.7):** `pptx_tcadp.go` + TCADP branch in
`pptx_parser.go`/`ppt_parser.go` — PowerPoint files now support
`parse_method="tcadp"` via the TCADP cloud service, matching the
spreadsheet-family TCADP pattern. PPT containers pass `"PPT"` as
fileType (not hardcoded `"PPTX"`).
- **Audio default output_format (2.11):** `defaultSetups()` audio
default changed from `"text"` to `"json"`, aligning with Python
`parser.py:232` and `AllowedOutputFormat["audio"]={"json"}`.
- **PDF VLM enhancement (1.1):** `maybeDispatchPDFVisionEnhancement` in
`pdf_vision_dispatch.go` enriches image/table items with IMAGE2TEXT
model descriptions after PDF parsing, mirroring Python
`enhance_media_sections_with_vision`. Semaphore fix: acquire before
goroutine start to prevent unbounded goroutine creation.
- **json family (2.3):** reclassified as Keep Go — `json_parser.go` is a
functional enhancement, not a parity gap.
- **page number:** changed from "mixed use of 1-indexed & 0-indexed" to
"1-indexed"
### Chunker
- **BULLET_PATTERN fallback (1.7):** 4th-level fallback in
`resolveTitleLevels` (`title.go`) detects bullet/numbered-list patterns
(Chinese legal, numbering, English) when outline + regex levels produce
only bodyLevel. Guarded by `allBodyLevel` to never override existing
structure.
- **Tag/One chunker fields (1.8):** `tag.go` sets `TopInt` from source
row index; `one.go` preserves `Positions`/`PDFPositions` from source
items. TSV multi-line RowNum fix: tracks `contentStart` for correct row
attribution.
- **Overlapped_percent normalization (2.6):**
`NormalizeOverlappedPercent` in `schema/chunker.go` mirrors Python
`common/float_utils.py:50-58` — accepts `[0,1)` fraction or `[0,90]`
percent, normalizes to canonical `[0,90]`.
- **Paragraph splitting (2.7):** aligned to Python flow `naive_merge` —
`CRLF` normalization, `splitKeepingDelimiter` preserves sentence
delimiters, single-section merge with token-budget-governed chunking.
- **chunk order:** sort by reading order
### Tokenizer
- **Phantom chunk filtering (Omission 2):** `isPhantomChunk` + filter
loop in `chunksFromTokenizerUpstream` skips zero-value ChunkDocs.
- **Batch size env var (Omission 3):** `embeddingBatchSize()` reads
`TOKENIZER_EMBEDDING_BATCH_SIZE`, defaults to 16.
- **Summary empty check (Diff 5):** `TrimSpace(s) != ""` → `s != ""`,
matching Python truthy check.
- **chunk_order_int all paths (Diff 8):** set unconditionally before
full_text/embedding branching.
- **Timeout default (Diff 10):** `600s` → `60s`, matching Python
`@timeout(60)`.
- **Small maxTokens truncation (Diff 14):** `truncateForEmbedding`
returns `""` when `maxTokens <= 10`, matching Python.
### Code review fixes
- Semaphore acquire moved before goroutine in `pdf_vision_dispatch.go`
(concurrency control)
- Context propagation in `pptx_tcadp.go` (cancellation support)
- Test resolver leak fix in `media_dispatch_test.go` (defer restore)
- Migration history comments removed per AGENTS.md
## Test plan
```
bash build.sh --test ./internal/parser/parser/... ./internal/ingestion/component/...
```
## Notes
- Migration diff tracking: `docs/migration_python_go_diff.md`
- Remaining gaps: Extractor component only (21 items)
|
||
|
|
76acd499d4 |
Fix: use independent fixtures in ClampsOverlappedPct test (#17416)
## Problem `TestMergeByTokenSizeFromJSON_ClampsOverlappedPct` reused a single `items` fixture across four calls to `mergeByTokenSizeFromJSON`. `mergeByTokenSizeFromJSON` mutates its `perItem` argument in place and returns the same backing array (token.go: `perItem[idx] = merged`). As a result: - The 2nd and later calls merged already-merged chunks instead of the original input. - Because the returned slice aliases the input, later calls silently overwrote the earlier results (`at100`, `at0`, ...). - `reflect.DeepEqual(at100, at150)` was therefore vacuously true — the test was a **false positive** that never actually exercised the clamp. It would still pass even if the clamp were broken. This is exactly the defect flagged in the review comment on PR #17396. ## Fix Add a `clampOverlapFixture()` factory and build a fresh fixture for every call, so that 150 / 1e300 / -5 / -1e300 are applied to the original input. ## Verification - The test passes after the fix. - When the clamp logic was temporarily disabled, the test **failed** (panic at token.go:768 — the negative index produced by an out-of-range pct), proving the fixed test is no longer a false positive and can catch a clamp regression. ## Scope Test file only. No production code change (verified `token.go` is identical to `upstream/main`). Refs: review comment on PR #17396. |
||
|
|
a6a67c5ece |
Chunker: port Python overlapped_percent normalization and BULLET_PATTERN title fallback (#17396)
## Summary Closes two Chunker migration gaps documented in `docs/migration_python_go_diff.md` (diffs **2.6** and **1.7**), improving parity with the Python ingestion pipeline. This is the code portion of commit `261e1fd0b` on branch `fix/batch-3-4`; the migration document itself is tracked separately (untracked in this commit). ### Diff 2.6 — `overlapped_percent` missing Python normalization Python's `normalize_overlapped_percent` (`common/float_utils.py:50-58`) accepts a `[0,1)` fraction (the flow-canvas UI validates `[0,1)`), multiplies it by 100, `int()`-truncates, and clamps to `[0,90]`. Go previously only accepted a raw `[0,90]` percentage and **rejected** out-of-range input, so a Python config passing `0.1` (meaning 10%) silently produced ~0% overlap. Added `normalizeOverlappedPercent` (`internal/ingestion/component/chunker/common.go`) mirroring the Python helper: - parses numbers and numeric strings (mirrors Python `float()`; bad / `NaN` / `Inf` → `0`), - `0 < v < 1 → v *= 100`, - `int()` truncation, - clamps to `[0,90]`. Wired into `tokenChunkerParam.Update` (`token.go`); the merge math (`token.go:705`, `(100-x)/100`) already matched Python. `TokenChunkerParam.Validate` (`schema/chunker.go`) now only guards direct struct construction. Tests: `TestNormalizeOverlappedPercent`, extended `TestTokenChunker_NewAcceptsPythonOverlappedRange` (adds fraction/clamp inputs), new `TestTokenChunker_NormalizesOverlappedPercent`. The existing reject-test cases for `<0` / `>90` were removed because they are now normalized/clamped (Python parity). ### Diff 1.7 — missing `BULLET_PATTERN` title-level fallback `resolveTitleLevels` (`title.go`) now applies a 4th-level fallback: when outline + regex + layout all yield body level, `bulletsCategory` selects the best-matching bullet-pattern group (Chinese legal / numbering / Chinese numbering / English legal — mirroring `rag/nlp/__init__.py:258-320`) and assigns structural levels. Guarded by `allBodyLevel` so it never overrides an existing outline/regex level. Tests: `TestResolveTitleLevels_BulletFallback` (4 subtests). ### Incidental test adjustments included in the commit - `token_batch1_test.go`: overlap input changed `0.3` → `30.0` to reflect the post-normalization 30% semantics. - `real_consumer_test.go`: updated `LoadFromIngestionTask(task)` → `LoadFromIngestionTask(ctx, task)` for the new context-first signature. ## Verification `bash build.sh --test ./internal/ingestion/component/...` — chunker + schema suites pass, no regression (CGO build). ## Migration doc reference `docs/migration_python_go_diff.md` §Chunker 1.7 and 2.6 are marked **Fixed** for these changes. |
||
|
|
554925b583 |
Fix(go): align ingestion pipeline with Python (parser/media dispatch + PDF coordinate chain + Chunker) (#17349)
## Summary Aligns the Go ingestion pipeline with the Python implementation, closing several behavioral gaps found during the Python→Go migration (tracked in `docs/migration_python_go_diff.md`). Covers parser/media dispatch alignment, the PDF coordinate-chain (preview images, outline→title, chunk coordinate finalization), and the Chunker Token/QA batches below. Commits are grouped as follows. ### 1. Fix parser params ( |