mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-07-31 04:59:24 +08:00
## Summary Fixes embedding inputs being emptied and documents chunker heading-detection parity. - `component/tokenizer.go`: `truncateForEmbedding` now mirrors Python `common/token_utils.py:183-185` `truncate(string, max_len)` (keep the first `max_len` tokens). For an unconfigured embedder reporting `maxTokens <= 0`, instead of mirroring Python's `truncate` (which returns `""` for `max_len <= 0` and would make the embeddings API reject the batch with "inputs cannot be empty"), Go **clamps the limit to a safe default** (`defaultEmbeddingTokenLimit = 8192`) so truncation stays active and never produces empty inputs. The previous code returned `""` for `maxTokens <= 10`, which emptied non-empty chunks into the batch and triggered the API error. - A **10-token safety margin** is reserved (mirroring Python's embedding path `rag/svr/task_executor.py` which uses `mdl.max_length - 10`) so Go does not over-send tokens to hard-capped embedding APIs; it is only applied when the limit `> 10` so small limits remain non-empty. - The clamp is applied **centrally inside `truncateForEmbedding`**, covering both Builtin and generic callers — no separate `model_service.go` change is required. - `component/chunker/title.go`: comment-only — documents the already-ported PDF-outline detection (omission 1.5 / Gap C) and the hardcoded `BULLET_PATTERN` fallback (omission 1.7 / Gap C), porting `common.py:_outline_similarity` and `rag/nlp` `BULLET_PATTERN`. - `tokenizer_unit_test.go`: update truncation expectations to the new positive-keeps-tokens / non-positive-clamps-to-default behaviour. ## Test plan - `tokenizer_unit_test.go` updated for the new truncation behaviour (`TestTruncateForEmbedding_SmallMaxTokens`, `TestTruncateForEmbedding_UnconfiguredClampsToDefault`). - `./build.sh --test ./internal/ingestion/component/` passes. 🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy)