Files
ragflow/internal/entity/knowledge_compile_doc.go

51 lines
2.8 KiB
Go
Raw Normal View History

//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package entity
import "time"
refactor(knowledge_compile): global compile pool, token-budget batching, and DocEngine-only deletion (#17679) ## Summary This PR refactors the Go knowledge-compilation ingestion pipeline (`internal/ingestion/knowledge_compile` + `internal/ingestion/component/knowledge_compiler`) with three related changes: - **Token-budget batching for LLM merge decisions.** `LLMMergeDecider.DecideBatch` previously stuffed every `(existing, candidate)` pair into a single LLM call, risking `max_token` overflow. It now splits pairs into token-bounded sub-batches (budget = `llmMaxTokens * 0.85`) via `tokenizer.NumTokensFromString`, runs them concurrently while preserving the global pair index, and never reindexes. - **Process-level global compile pool.** Introduces a single vCPU-sized goroutine pool (`pool.go`, env `KC_COMPILE_CONCURRENCY`) dedicated to *all* knowledge-compilation stages. KNN search loop, `DecideBatch` sub-batches, `WriteMerged`/`DeleteMerged` internals, and the component-level (structure/mindmap) per-call pools are all unified into it via an injected submitter. No more per-job short-lived goroutines in `runCompilerJobs` (futures are collected then awaited on the caller). Fan-out stays bounded by the pool worker count; these stages are docengine-bounded / LLM-bounded, not CPU-bounded. - **DocEngine-only deletion.** `Consumer.processBatch` deletion no longer loads the deleted docs' products into memory. Two sequential DocEngine calls replace the old in-memory surgery: - `DeleteDocLevelForDocs` — one `DeleteChunks` over `doc_id IN deletedDocIDs` (merged rows carry `doc_id == kb`, so only per-doc products match). - `StripMergedSources` — one `Search` of `kc_merged=1` rows filtered by `source_doc_ids IN deletedDocIDs` (intersection pushed down to the engine), `UpdateChunks` the source array of survivors, and `DeleteChunks` the rows whose array became empty. ## Changes - `internal/ingestion/knowledge_compile/pool.go` (new): global `compilerPool` + `runCompilerJobs`/`SubmitCompilerJob`/`SubmitCompilerJobs`. - `internal/ingestion/knowledge_compile/consumer.go`: deletion rewritten to the two DocEngine calls; `mergedBase`/`toDelete`/`stripDeletedSources` removed. - `internal/ingestion/knowledge_compile/writer.go`: `DeleteDocLevelForDocs` + `StripMergedSources` replace `DeleteMergedForDoc`/`DeleteMerged`. - `internal/ingestion/knowledge_compile/reader.go`: drop `LoadMergedBySourceDoc` + `containsString` (keep `LoadDocProducts` for the completion branch). - `internal/ingestion/knowledge_compile/dedup.go`: `NewLLMDeduper` takes `llmMaxTokens`; wires `SetMaxBatchTokens`/`SetSubmitter`. - `internal/ingestion/knowledge_compiler/{structure,merge}.go`, `mindmap/mindmap.go`, `pool_wiring.go`: token-budget split + submitter injection. - Tests: `structure_test.go` (token-budget split), `dedup_test.go`, `consumer_test.go` (tombstone + DocEngine deletion assertions) updated. ## Validation `bash build.sh --test -race ./internal/ingestion/knowledge_compile/... ./internal/ingestion/component/knowledge_compiler/...` passes (unit tier, no external services). 🤖 Generated with [CodeBuddy](https://www.codebuddy.ai) --------- Co-authored-by: yuzhichang <yuzhichang@infiniflow.ai>
2026-08-02 17:06:29 +08:00
// KnowledgeCompileDataset is the MySQL scheduling row for the dataset-level
// post-processing consumer (knowledge_compile_design.md §11.4, Option E). It is
// the scheduling system of record: backlog_doc_ids holds the not-yet-processed
// doc entries for the KB, inflight_doc_ids the ones a worker has claimed (the
// closed batch), and the claim_* fields the owner/lease used for crash
// recovery. NATS notify is only a wake-up; same-KB serialization comes from
// these rows, not from the broker.
//
// The *_doc_ids columns store a JSON array of knowledge_compile.BacklogEntry
// (doc_id + event_type + seq) as TEXT so the consumer can re-apply the same
// out-of-order / tombstone guards as the broker-based design without re-reading
// the queue.
refactor(knowledge_compile): global compile pool, token-budget batching, and DocEngine-only deletion (#17679) ## Summary This PR refactors the Go knowledge-compilation ingestion pipeline (`internal/ingestion/knowledge_compile` + `internal/ingestion/component/knowledge_compiler`) with three related changes: - **Token-budget batching for LLM merge decisions.** `LLMMergeDecider.DecideBatch` previously stuffed every `(existing, candidate)` pair into a single LLM call, risking `max_token` overflow. It now splits pairs into token-bounded sub-batches (budget = `llmMaxTokens * 0.85`) via `tokenizer.NumTokensFromString`, runs them concurrently while preserving the global pair index, and never reindexes. - **Process-level global compile pool.** Introduces a single vCPU-sized goroutine pool (`pool.go`, env `KC_COMPILE_CONCURRENCY`) dedicated to *all* knowledge-compilation stages. KNN search loop, `DecideBatch` sub-batches, `WriteMerged`/`DeleteMerged` internals, and the component-level (structure/mindmap) per-call pools are all unified into it via an injected submitter. No more per-job short-lived goroutines in `runCompilerJobs` (futures are collected then awaited on the caller). Fan-out stays bounded by the pool worker count; these stages are docengine-bounded / LLM-bounded, not CPU-bounded. - **DocEngine-only deletion.** `Consumer.processBatch` deletion no longer loads the deleted docs' products into memory. Two sequential DocEngine calls replace the old in-memory surgery: - `DeleteDocLevelForDocs` — one `DeleteChunks` over `doc_id IN deletedDocIDs` (merged rows carry `doc_id == kb`, so only per-doc products match). - `StripMergedSources` — one `Search` of `kc_merged=1` rows filtered by `source_doc_ids IN deletedDocIDs` (intersection pushed down to the engine), `UpdateChunks` the source array of survivors, and `DeleteChunks` the rows whose array became empty. ## Changes - `internal/ingestion/knowledge_compile/pool.go` (new): global `compilerPool` + `runCompilerJobs`/`SubmitCompilerJob`/`SubmitCompilerJobs`. - `internal/ingestion/knowledge_compile/consumer.go`: deletion rewritten to the two DocEngine calls; `mergedBase`/`toDelete`/`stripDeletedSources` removed. - `internal/ingestion/knowledge_compile/writer.go`: `DeleteDocLevelForDocs` + `StripMergedSources` replace `DeleteMergedForDoc`/`DeleteMerged`. - `internal/ingestion/knowledge_compile/reader.go`: drop `LoadMergedBySourceDoc` + `containsString` (keep `LoadDocProducts` for the completion branch). - `internal/ingestion/knowledge_compile/dedup.go`: `NewLLMDeduper` takes `llmMaxTokens`; wires `SetMaxBatchTokens`/`SetSubmitter`. - `internal/ingestion/knowledge_compiler/{structure,merge}.go`, `mindmap/mindmap.go`, `pool_wiring.go`: token-budget split + submitter injection. - Tests: `structure_test.go` (token-budget split), `dedup_test.go`, `consumer_test.go` (tombstone + DocEngine deletion assertions) updated. ## Validation `bash build.sh --test -race ./internal/ingestion/knowledge_compile/... ./internal/ingestion/component/knowledge_compiler/...` passes (unit tier, no external services). 🤖 Generated with [CodeBuddy](https://www.codebuddy.ai) --------- Co-authored-by: yuzhichang <yuzhichang@infiniflow.ai>
2026-08-02 17:06:29 +08:00
type KnowledgeCompileDataset struct {
DatasetID string `gorm:"primaryKey;column:dataset_id;size:64" json:"dataset_id"`
TenantID string `gorm:"column:tenant_id;size:64;not null;default:''" json:"tenant_id"`
// The *_doc_ids columns store a JSON array as TEXT. No DDL default is set:
// MySQL (8.0.13+) rejects a literal DEFAULT on TEXT/BLOB columns (Error
// 1101), and the application always writes "[]" explicitly on insert/update
// (scheduler.go FirstOrCreate / release paths), so the default is redundant.
BacklogDocIDs string `gorm:"column:backlog_doc_ids;type:text;not null" json:"backlog_doc_ids"`
InflightDocIDs string `gorm:"column:inflight_doc_ids;type:text;not null" json:"inflight_doc_ids"`
ClaimOwner string `gorm:"column:claim_owner;size:64;not null;default:''" json:"claim_owner"`
ClaimToken string `gorm:"column:claim_token;size:64;not null;default:''" json:"claim_token"`
ClaimExpiresAt *time.Time `gorm:"column:claim_expires_at;default:null" json:"claim_expires_at"`
Priority int `gorm:"column:priority;not null;default:0" json:"priority"`
CreatedAt time.Time `gorm:"column:created_at;autoCreateTime" json:"created_at"`
UpdatedAt time.Time `gorm:"column:updated_at;autoUpdateTime" json:"updated_at"`
}
// TableName pins the scheduling table name.
refactor(knowledge_compile): global compile pool, token-budget batching, and DocEngine-only deletion (#17679) ## Summary This PR refactors the Go knowledge-compilation ingestion pipeline (`internal/ingestion/knowledge_compile` + `internal/ingestion/component/knowledge_compiler`) with three related changes: - **Token-budget batching for LLM merge decisions.** `LLMMergeDecider.DecideBatch` previously stuffed every `(existing, candidate)` pair into a single LLM call, risking `max_token` overflow. It now splits pairs into token-bounded sub-batches (budget = `llmMaxTokens * 0.85`) via `tokenizer.NumTokensFromString`, runs them concurrently while preserving the global pair index, and never reindexes. - **Process-level global compile pool.** Introduces a single vCPU-sized goroutine pool (`pool.go`, env `KC_COMPILE_CONCURRENCY`) dedicated to *all* knowledge-compilation stages. KNN search loop, `DecideBatch` sub-batches, `WriteMerged`/`DeleteMerged` internals, and the component-level (structure/mindmap) per-call pools are all unified into it via an injected submitter. No more per-job short-lived goroutines in `runCompilerJobs` (futures are collected then awaited on the caller). Fan-out stays bounded by the pool worker count; these stages are docengine-bounded / LLM-bounded, not CPU-bounded. - **DocEngine-only deletion.** `Consumer.processBatch` deletion no longer loads the deleted docs' products into memory. Two sequential DocEngine calls replace the old in-memory surgery: - `DeleteDocLevelForDocs` — one `DeleteChunks` over `doc_id IN deletedDocIDs` (merged rows carry `doc_id == kb`, so only per-doc products match). - `StripMergedSources` — one `Search` of `kc_merged=1` rows filtered by `source_doc_ids IN deletedDocIDs` (intersection pushed down to the engine), `UpdateChunks` the source array of survivors, and `DeleteChunks` the rows whose array became empty. ## Changes - `internal/ingestion/knowledge_compile/pool.go` (new): global `compilerPool` + `runCompilerJobs`/`SubmitCompilerJob`/`SubmitCompilerJobs`. - `internal/ingestion/knowledge_compile/consumer.go`: deletion rewritten to the two DocEngine calls; `mergedBase`/`toDelete`/`stripDeletedSources` removed. - `internal/ingestion/knowledge_compile/writer.go`: `DeleteDocLevelForDocs` + `StripMergedSources` replace `DeleteMergedForDoc`/`DeleteMerged`. - `internal/ingestion/knowledge_compile/reader.go`: drop `LoadMergedBySourceDoc` + `containsString` (keep `LoadDocProducts` for the completion branch). - `internal/ingestion/knowledge_compile/dedup.go`: `NewLLMDeduper` takes `llmMaxTokens`; wires `SetMaxBatchTokens`/`SetSubmitter`. - `internal/ingestion/knowledge_compiler/{structure,merge}.go`, `mindmap/mindmap.go`, `pool_wiring.go`: token-budget split + submitter injection. - Tests: `structure_test.go` (token-budget split), `dedup_test.go`, `consumer_test.go` (tombstone + DocEngine deletion assertions) updated. ## Validation `bash build.sh --test -race ./internal/ingestion/knowledge_compile/... ./internal/ingestion/component/knowledge_compiler/...` passes (unit tier, no external services). 🤖 Generated with [CodeBuddy](https://www.codebuddy.ai) --------- Co-authored-by: yuzhichang <yuzhichang@infiniflow.ai>
2026-08-02 17:06:29 +08:00
func (KnowledgeCompileDataset) TableName() string { return "knowledge_compile_docs" }