Commit Graph

174 Commits

Author SHA1 Message Date
Haruko386
a33ee7d4a0 fix: dataset ingestion log has no other type instead of success (#18605) 2026-08-21 15:41:40 +08:00
Jack
99e1c2eaca fix(parser): allow text output_format for audio family (#18575) 2026-08-20 22:41:01 +08:00
euvre
c87aa3b168 fix: store source key only in pipeline_operation_log.source_from (#18239) 2026-08-20 20:17:52 +08:00
buua436
64d928e45f fix: align wiki graph and dataset logs (#18566) 2026-08-20 19:38:40 +08:00
euvre
cc28035881 fix(ingestion): show data pipeline column for builtin pipeline ingestion logs (#18484) 2026-08-20 19:04:12 +08:00
buua436
a3df588951 feat: improve incremental Wiki compilation (#18557) 2026-08-20 16:45:34 +08:00
Jack
4eec5e4f35 feat(chunker): enforce a hard token cap in the TokenChunker merge (#18525) 2026-08-20 16:28:20 +08:00
Jack
d76c15e2a5 feat(title-chunker): enforce chunk_token_cap on the Go title-family chunks (#18455) (#18551) 2026-08-20 15:41:48 +08:00
jay77721
2022efa9fc refactor(ingestion): unify auto-metadata on modular metadata and hard-delete legacy flats (BuiltInMetadata carried, not LLM) (#18511) 2026-08-20 15:35:51 +08:00
Jack
84bed4dec5 refactor(chunker): replace QAChunker/PresentationChunker with PairChunker/PageChunker, drop TagChunker (#18523) 2026-08-20 11:53:00 +08:00
jay77721
2a1a93ed3a refactor(extractor): unify 5-in-1 modular schema across dataset and pipeline, drop legacy field_name and custom prompt (#18432)
Refactor the Extractor component into a pure, unified **5-in-1 modular extraction engine** across both Dataset (`knowledgebase.parser_config`) and Pipeline (Canvas DSL).
2026-08-19 15:02:46 +08:00
Wang Qi
5f92323ced Fix ingestion pipeline child delimiter extra newline (#18490) 2026-08-19 13:47:09 +08:00
buua436
31812e1de5 feat: improve Go wiki compilation (#18479)
Improve Go Wiki compilation with chunk-level MAP caching, topic-aware planning, entity/topic page merging, and correct source chunk counting.
2026-08-19 11:41:20 +08:00
ming1523
6756b48f10 test: add text&code parser-to-chunker parity golden (#18311) 2026-08-18 22:45:57 +08:00
jay77721
c71991bb7a feat(ingestion,web): modularize extractor configuration with sub-tabs, independent prompts, and metadata integration (#18383)
This PR modularizes the **Extractor** component configuration with dedicated feature subtabs, adds independent system prompt configuration, fixes multi-node execution determinism and parameter persistence across save and page refresh, and ensures backward compatibility with legacy flat fields.
2026-08-18 11:38:52 +08:00
Jack
37b945c94d fix(ingestor): stop dropping tasks under burst parse backpressure (#18369)
Fixes the defect where selecting many files (e.g. 20+) in one dataset and starting parsing at once leaves most of them stuck in `RUNNING` forever: a few parse, the rest never do.
2026-08-17 19:09:58 +08:00
questfever
52519e2fd7 refactor: replace Split in loops with more efficient SplitSeq (#18248) 2026-08-17 19:04:04 +08:00
buua436
1ef4ddddec fix: align Go ingestion progress and pipeline selection (#18366)
Fixes Go ingestion progress reporting and pipeline selection:
- Add timestamps to document progress logs.
- Keep document duration and status updated during parsing.
- Start frontend polling immediately after parsing begins.
- Prevent documents explicitly using General from inheriting an old
dataset pipeline.
- Populate missing pipeline operation log fields.
- Remove stale component progress logs between retries.
- Prevent progress values greater than `1`.
2026-08-17 16:58:44 +08:00
jay77721
3e38b6cacb fix(ingestion): unify extractor prompt placeholder rendering and per-chunk substitution (#18355) 2026-08-17 15:19:47 +08:00
jay77721
452720a62d feat(ingestion): end-to-end Auto metadata in Go (extractor merge, builtin UI, built-in fields) (#18273) 2026-08-17 14:17:48 +08:00
euvre
9f358cbeaa fix(ingestion): accumulate component progress lines into document.progress_msg (#18356) 2026-08-17 13:27:40 +08:00
euvre
fe963e69ba fix: route TokenChunker delimiter_mode "one" to OneChunker in Go ingestion (#18121) 2026-08-14 16:55:32 +08:00
Jack
ae256bcf59 feat(chunker): add ManualChunker for manual-layout PDFs (#18272)
Add `ManualChunker`, the Go port of Python's `manual` doc-type chunk method (`rag/app/manual.py`). Like `GroupTitleChunker` it merges adjacent text records into heading-bounded groups, but it first re-sorts the records into physical reading order before grouping.
2026-08-14 16:55:08 +08:00
Jack
423c8489b5 fix(chunker): carry overlap-head PDF positions into new chunk (#18148) (#18227)
When `TokenChunker` starts a fresh chunk with an overlap prefix (Go `computeOverlapPrefix` / Python visible-text cut), the previous chunk's **tail PDF coordinates were dropped**. As a result, the overlap head of a PDF chunk is displayed but **not highlighted** — the highlight box is shifted/truncated relative to the displayed span (infiniflow/ragflow#18148).
2026-08-13 22:18:33 +08:00
Jin Hai
811f9dd0df Go: refactor embed and rerank interface (#18240)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-13 22:14:23 +08:00
Jin Hai
e2acf3aebb Go: add error logs when fail to init infinity (#18235)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-13 20:17:53 +08:00
Zhichang Yu
0784bef5b0 Port dataset-level structure merge for timeline/graph/mindmap (#18201)
Ports Python dataset-level structure aggregation (timeline, graph, mindmap) to Go. Mindmap emits entity/relation rows and merges like graph. Adds dataset_merge guard, engine gate, resolveDatasetStructureKind, kind-required structure graph GET/DELETE API, per-index task-id fields.
2026-08-13 18:37:47 +08:00
jay77721
0a4cc65db0 refactor(common): unify </think>-strip into shared StripThinkTrailing helper (#18179)
Collapse the nine identical **"cut through the last `</think>`"**
implementations — mirroring Python's `re.sub(r"^.*</think>", "", s,
re.DOTALL)` — into one shared helper `common.StripThinkTrailing`,
preventing future behavior drift between copies.
2026-08-13 13:56:56 +08:00
jay77721
e1b8366ecc refactor: remove resume (CV) workflow DSL template and Go handling code (#18177) 2026-08-13 10:47:37 +08:00
Zhichang Yu
6677f14bdf Port dataset nav and structure graph fixes to Go, merge agents list (#18183)
Fix compilation template config validation for JSONMap; merge template groups into agents list ordered by category/name; install nav service in ingestor; write readable nav cluster/doc names and emit nav_doc leaves;
port tree-to-graph projection and full document structure graph endpoint parity.
2026-08-12 22:46:24 +08:00
Jack
b505437db5 fix(chunker): honor bare (non-backtick) delimiters in TokenChunker (#17723) (#18182)
Go's `TokenChunker` previously ignored **bare (non-backtick)
delimiters** such as `::`, `.`, or `;`. `compileDelimPattern` compiled
them with `keepBare=false`, so a bare delimiter produced a `nil` pattern
and the payload was routed to the single-section merge that never split
on it — diverging from Python's `naive_merge`, which splits on bare
delimiters and then merges by token size.
2026-08-12 22:36:11 +08:00
jay77721
207e2eaf3b fix(ingestion): strip mid-text think blocks in extractor metadata path (#18175) 2026-08-12 19:39:49 +08:00
jay77721
91fd114783 fix(dao): honor tenant-configured context window override (D22/D23) (#18171)
Make `ResolveModelContentLength` honor the per-model custom **context window length** (`content_length`) — stored in the Python-legacy `tenant_model.extra["max_tokens"]` field, whose semantic meaning is the context window, NOT the generation cap — **before** any provider-catalog read, and remove the parallel service-layer implementation so every consumer shares one resolution path.
2026-08-12 19:32:37 +08:00
euvre
3d64f8d044 fix(ingestion): remove broken built-in Resume pipeline (#18173) 2026-08-12 19:11:49 +08:00
Jack
a4dbd898bb docs(ingestion): correct stale parser-component scope comment (#18147) 2026-08-12 17:58:05 +08:00
Zhichang Yu
c677e9af36 Port dataset-level knowledge compile to Go with variant dispatch (#18161)
Ports dataset-level knowledge compilation (tree/structure/wiki) to Go:
add compile-type variants to backlog events, route per-variant
dataset-level paths, move dataset-nav to the consumer, add structure
merge and per-variant clean, plus rebuild variant recovery.
2026-08-12 17:24:12 +08:00
jay77721
86b6f64abb fix(ingestion): fit extractor and tagger prompts to model context (#18095)
Trim Extractor call prompts and the automatic tagger prompt to the chat model's context window (`content_length`) before sending, so oversized chunks or tag files are trimmed instead of rejected by the provider with a context-length error.
2026-08-12 15:45:01 +08:00
taek105
492d6d81a9 fix: honor dataset language in Go vision dispatch (#17892)
### Summary

- Propagate the dataset language through Go DOCX, Markdown, PDF
figure-enhancement, and standalone-image vision paths.
- Explicitly render the shared figure prompt's `{{ language }}`
placeholder in Go.
- Use English when the dataset language is empty.
- Make the default standalone-image prompt request the dataset language
while preserving visible text in its original language.
- Add focused tests for caller propagation, language fallback, prompt
rendering, and prompt-cache isolation.
2026-08-11 22:18:04 +08:00
Zhichang Yu
d9ed14ce9c feat: wiki incremental Mode A/B with durable rewrite barrier (#18122)
Port the wiki_incremental dataset-level merge and make its rewrite
barrier durable and concurrency-safe. Wiki pages merge replace-only; the
barrier persists a monotonic numeric generation, and a scheduler-backed
per-dataset lock closes the cross-process TOCTOU window. Adds the
Compiler Plan toggle (frontend) with Mode A grouping.
2026-08-11 22:10:49 +08:00
Jack
a399b93143 Test(chunker): add golden parity harness, fixtures, and live Go<->Python tool (#17735)
Golden-parity test infrastructure for the **Go `TokenChunker` ↔ Python alignment**.
It runs the Go chunker over a committed case set (`testdata/parity/cases/`) and diffs each output against a captured Python golden (`testdata/parity/golden/`), honoring a `known_diffs.json` ratchet (`extra_fields` / `chunk_count` / `chunk_text`) so accepted divergences are tracked rather than silently widening.
2026-08-11 17:46:34 +08:00
Jin Hai
d7661b676d Go: fix context and warnings (#18097)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 16:18:49 +08:00
Zhichang Yu
64533e5b5e Refactor splitByTokens and wire wiki incremental compile (#18083)
Port dataset-level wiki incremental compile and refactor splitByTokens
token budgeting. Includes replace-only wiki merge, KNN dedup routing,
and template/config wiring.
2026-08-11 13:46:55 +08:00
jay77721
9f0663d4d0 fix(ingestion): avoid duplicate chunk text injection in Extractor prompts (#18034) 2026-08-11 11:53:34 +08:00
Jack
dc73163908 Fix(tokenizer): pin alignment guards (#18011) 2026-08-11 10:09:07 +08:00
Jin Hai
c697fcff41 Go: fix go context (#18052)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-10 18:04:17 +08:00
jay77721
ca9e62ee59 fix(ingestion): split Extractor call() into callRaw/callText/callStructured (#18038) 2026-08-10 17:18:24 +08:00
Jin Hai
c0bc146fcb Go: fix env variables (#18032)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-10 14:37:05 +08:00
deadtrickster
2d63ad654d fix(ingestion): make document/KB counter application idempotent per run (#17995) 2026-08-07 23:33:18 +08:00
Jack
4b4a6e72f0 fix(chunker): unify TokenChunker merge and strip coord tags in Python JSON path (#18002)
Unifies the Go TokenChunker merge path on a single `mergeUnits` core and
fixes coordinate-tag drift in the Python JSON merge at `overlap > 0`.
Rebased on top of #17979 (delimiter_mode convergence).
2026-08-07 21:55:07 +08:00
Zhichang Yu
f12c0ec08a feat(knowledge_compile): materialize wiki page graph (wiki_entity/wiki_relation) (#17976)
Re-materialize wiki page graph from merged wiki_page rows after each
batch merge. Adds ProjectWikiGraph/DropWikiGraph, full page_type/slug
identity, delete-then-insert, tests.
2026-08-07 17:47:59 +08:00