Files
ragflow/internal/entity/knowledge_compile_doc.go
Zhichang Yu 01d667296d refactor(knowledge_compile): global compile pool, token-budget batching, and DocEngine-only deletion (#17679)
## Summary

This PR refactors the Go knowledge-compilation ingestion pipeline
(`internal/ingestion/knowledge_compile` +
`internal/ingestion/component/knowledge_compiler`) with three related
changes:

- **Token-budget batching for LLM merge decisions.**
`LLMMergeDecider.DecideBatch` previously stuffed every `(existing,
candidate)` pair into a single LLM call, risking `max_token` overflow.
It now splits pairs into token-bounded sub-batches (budget =
`llmMaxTokens * 0.85`) via `tokenizer.NumTokensFromString`, runs them
concurrently while preserving the global pair index, and never
reindexes.
- **Process-level global compile pool.** Introduces a single vCPU-sized
goroutine pool (`pool.go`, env `KC_COMPILE_CONCURRENCY`) dedicated to
*all* knowledge-compilation stages. KNN search loop, `DecideBatch`
sub-batches, `WriteMerged`/`DeleteMerged` internals, and the
component-level (structure/mindmap) per-call pools are all unified into
it via an injected submitter. No more per-job short-lived goroutines in
`runCompilerJobs` (futures are collected then awaited on the caller).
Fan-out stays bounded by the pool worker count; these stages are
docengine-bounded / LLM-bounded, not CPU-bounded.
- **DocEngine-only deletion.** `Consumer.processBatch` deletion no
longer loads the deleted docs' products into memory. Two sequential
DocEngine calls replace the old in-memory surgery:
- `DeleteDocLevelForDocs` — one `DeleteChunks` over `doc_id IN
deletedDocIDs` (merged rows carry `doc_id == kb`, so only per-doc
products match).
- `StripMergedSources` — one `Search` of `kc_merged=1` rows filtered by
`source_doc_ids IN deletedDocIDs` (intersection pushed down to the
engine), `UpdateChunks` the source array of survivors, and
`DeleteChunks` the rows whose array became empty.

## Changes

- `internal/ingestion/knowledge_compile/pool.go` (new): global
`compilerPool` +
`runCompilerJobs`/`SubmitCompilerJob`/`SubmitCompilerJobs`.
- `internal/ingestion/knowledge_compile/consumer.go`: deletion rewritten
to the two DocEngine calls;
`mergedBase`/`toDelete`/`stripDeletedSources` removed.
- `internal/ingestion/knowledge_compile/writer.go`:
`DeleteDocLevelForDocs` + `StripMergedSources` replace
`DeleteMergedForDoc`/`DeleteMerged`.
- `internal/ingestion/knowledge_compile/reader.go`: drop
`LoadMergedBySourceDoc` + `containsString` (keep `LoadDocProducts` for
the completion branch).
- `internal/ingestion/knowledge_compile/dedup.go`: `NewLLMDeduper` takes
`llmMaxTokens`; wires `SetMaxBatchTokens`/`SetSubmitter`.
- `internal/ingestion/knowledge_compiler/{structure,merge}.go`,
`mindmap/mindmap.go`, `pool_wiring.go`: token-budget split + submitter
injection.
- Tests: `structure_test.go` (token-budget split), `dedup_test.go`,
`consumer_test.go` (tombstone + DocEngine deletion assertions) updated.

## Validation

`bash build.sh --test -race ./internal/ingestion/knowledge_compile/...
./internal/ingestion/component/knowledge_compiler/...` passes (unit
tier, no external services).

🤖 Generated with [CodeBuddy](https://www.codebuddy.ai)

---------

Co-authored-by: yuzhichang <yuzhichang@infiniflow.ai>
2026-08-02 17:06:29 +08:00

51 lines
2.8 KiB
Go

//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package entity
import "time"
// KnowledgeCompileDataset is the MySQL scheduling row for the dataset-level
// post-processing consumer (knowledge_compile_design.md §11.4, Option E). It is
// the scheduling system of record: backlog_doc_ids holds the not-yet-processed
// doc entries for the KB, inflight_doc_ids the ones a worker has claimed (the
// closed batch), and the claim_* fields the owner/lease used for crash
// recovery. NATS notify is only a wake-up; same-KB serialization comes from
// these rows, not from the broker.
//
// The *_doc_ids columns store a JSON array of knowledge_compile.BacklogEntry
// (doc_id + event_type + seq) as TEXT so the consumer can re-apply the same
// out-of-order / tombstone guards as the broker-based design without re-reading
// the queue.
type KnowledgeCompileDataset struct {
DatasetID string `gorm:"primaryKey;column:dataset_id;size:64" json:"dataset_id"`
TenantID string `gorm:"column:tenant_id;size:64;not null;default:''" json:"tenant_id"`
// The *_doc_ids columns store a JSON array as TEXT. No DDL default is set:
// MySQL (8.0.13+) rejects a literal DEFAULT on TEXT/BLOB columns (Error
// 1101), and the application always writes "[]" explicitly on insert/update
// (scheduler.go FirstOrCreate / release paths), so the default is redundant.
BacklogDocIDs string `gorm:"column:backlog_doc_ids;type:text;not null" json:"backlog_doc_ids"`
InflightDocIDs string `gorm:"column:inflight_doc_ids;type:text;not null" json:"inflight_doc_ids"`
ClaimOwner string `gorm:"column:claim_owner;size:64;not null;default:''" json:"claim_owner"`
ClaimToken string `gorm:"column:claim_token;size:64;not null;default:''" json:"claim_token"`
ClaimExpiresAt *time.Time `gorm:"column:claim_expires_at;default:null" json:"claim_expires_at"`
Priority int `gorm:"column:priority;not null;default:0" json:"priority"`
CreatedAt time.Time `gorm:"column:created_at;autoCreateTime" json:"created_at"`
UpdatedAt time.Time `gorm:"column:updated_at;autoUpdateTime" json:"updated_at"`
}
// TableName pins the scheduling table name.
func (KnowledgeCompileDataset) TableName() string { return "knowledge_compile_docs" }