refactor(ingestion/task): emit kb_id as a single string in ProcessChunksForPipeline (#17802)

## Summary
- `ProcessChunksForPipeline` now sets `kb_id` to a plain string instead
of `[]string{kbID}`, removing an index-physical array shape from the
ingestion domain.
- Stored documents are byte-identical: Elasticsearch overrides `kb_id`
with `datasetID` on write, and Infinity's `transformChunkFields` already
accepts a plain string.
- Infinity is intentionally left unchanged — `service/chunk` paths still
feed `kb_id` as `[]string`, and Infinity handles both forms. The
`dataset` artifact merge (`dataset_artifact_service.go`) is out of scope
for this step.
- Unit assertion updated to expect a string.

## Scope / non-goals
This is the smallest first step (T1) of the index-schema leak cleanup
tracked in #17371. It does **not** move the other leaks (`docnm_kwd`,
`create_timestamp_flt`, position ints) to the engine boundary — those
are later steps behind a read-back golden test.

## Test plan
- `go test ./internal/ingestion/task/indexdoc/...` passes.
- The two `task` "Real" integration tests fail identically on a clean
tree (environment lacks real embedding/parsing); they are pre-existing,
unrelated to this change.

🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy)
This commit is contained in:
Jack
2026-08-04 19:50:53 +08:00
committed by GitHub
parent 744b3ea7c1
commit 5efdd2d795
7 changed files with 310 additions and 12 deletions

View File

@@ -53,12 +53,8 @@ func TestProcessChunksForPipeline_SetsDocIDAndKBID(t *testing.T) {
if chunks[0]["doc_id"] != "doc-1" {
t.Errorf("doc_id = %q, want \"doc-1\"", chunks[0]["doc_id"])
}
if kbIDs, ok := chunks[0]["kb_id"].([]string); ok {
if len(kbIDs) != 1 || kbIDs[0] != "kb-1" {
t.Errorf("kb_id = %v, want [\"kb-1\"]", chunks[0]["kb_id"])
}
} else {
t.Errorf("kb_id should be []string, got %T", chunks[0]["kb_id"])
if kbID, ok := chunks[0]["kb_id"].(string); !ok || kbID != "kb-1" {
t.Errorf("kb_id = %v, want \"kb-1\" (string)", chunks[0]["kb_id"])
}
}