Files
ragflow/internal/ingestion/task/debug.go
Jack 8546eb62bf Feat(ingestion): add canvas pipeline debug (dry-run) mode with View result log (#17538)
## Summary
Adds a side-effect-free DataFlow canvas **debug (dry-run) mode** plus a
**debug run log with a "View result" panel**, so a canvas can be
executed synchronously and inspected end-to-end (per-component progress
and parsed chunks) without persisting anything.

### Dry-run execution (inline parsed chunks)
- `task/debug.go`: `NewDebugTaskContext` builds an in-memory
`TaskContext` with **`KB.ID == ""` — the single debug signal used across
the ingestion pipeline**. A canvas debug run has no knowledgebase, and
production ingestion always supplies one, so `kb_id == ""` occurs ONLY
in debug mode. Components gate their own side effects on this signal
without any dedicated debug vocabulary (the former `CANVAS_DEBUG_DOC_ID`
marker constant is removed).
- `task/pipeline_executor.go`: `validateTaskContext` no longer requires
a KB when `KB.ID == ""` (debug); debug runs return `collectDebugOutput`
(chunks) instead of a no-op; uploaded bytes are delivered as
`inputs['binary']` for doc-less runs; `injectDebugPageCap` caps the
parser to the first pages for a fast preview via the production
`override_params` channel (Parser cpnID + family).
- `component/tokenizer.go`: `shouldHaveEmbedding` skips embedding when
`kb_id == ""` — the embedder is configured on the knowledgebase, so a
debug run has nothing to resolve against and stays side-effect free.
- `chunker/register.go`: chunk images are uploaded to MinIO only when a
KB is present (persist run). **In debug mode the raw image bytes are
intentionally dropped (`delete(ck, "image")`)** — the debug preview does
not render chunk images, and dropping the bytes keeps them out of memory
and out of the Redis-stored debug log. This is a deliberate trade-off,
not an oversight.
- `component/file.go`: pass through in-memory binary bytes, skipping
`doc_id` -> storage resolution.
- `handler/agent.go` + `agent_webhook.go`: detect `dataflow_canvas` and
run a sync debug returning chunks inline on the existing
chat/completions endpoint; reject DataFlow canvases from webhooks (fixes
the previously dead `== "DataFlow"` check; mirrors Python
`agent_api.py`).
- `parser_dispatch.go`: export `ParserFileFamily` for the executor's
page-cap injection.

### Debug run log + "View result"
Mirrors Python's debug-log contract so the front-end can replay each
component's progress and parsed output:
- `task/debug_log_sink.go`: a `DebugLogSink` records every component's
lifecycle into a `[{component_id, trace}]` array (each trace entry
carries `message`, `progress`, `timestamp`, `elapsed_time`). `Flush`
appends a terminal `END` marker whose first trace message is non-empty
so the front-end detects completion. On failure the END marker is
prefixed `[ERROR]` yet still carries the run, so the failure timeline
renders instead of being stuck empty. Timestamps and `elapsed_time` are
in seconds (matching the rest of the app).
- `task/debug_result_dsl.go`: `BuildDebugResultDSL` builds the `dsl` the
END marker carries — the Go analogue of Python's `Graph.__str__` +
END-marker `dsl` in `rag/flow/pipeline.py`. It combines the static DSL
structure (component_name / downstream / params / graph.nodes) with the
run output map (`output["state"][<id>]`) to emit, per component,
`obj.params.outputs[<format>].value` (chunks / text / json / html /
markdown) — the exact keys the front-end `dataflow-result` page reads to
render each step's parsed chunks. Raw embedding vectors (including the
dimension-scoped `q_<dim>_vec` keys) are stripped so the stored log
stays Python-scale.
- `task/pipeline_executor.go`: after the run, attach the built `dsl` to
the END marker via the `ResultSink` capability.
- `handler/agent.go`: `runCanvasPipelineDebug` generates a stable
`message_id` up-front and always flushes the log (success or failure);
`respondWithDebugResult` returns `message_id` in **both** the success
and the error envelope so the front-end can poll the log. The debug-log
endpoint `GET /agents/:id/logs/:message_id` serves the array.
- `web/src/pages/agent/hooks/use-run-dataflow.ts`: on a run failure,
also surface `message_id` via `setMessageId` so the log sheet renders
the failure timeline (the `[ERROR]` END marker is already written).
Guarded by `if (msgId)`, so it is a safe no-op when the back-end does
not return an id.

## Behavioral notes
- Debug parses only the first pages (`debugPageCapPages`) for a fast
preview; an explicit `pages` cap already present in the ParserConfig is
respected.
- Debug mode does not keep chunk images (see above) and does not compute
embeddings — it exercises parse + chunk only.

## Test plan
- Go: `debug_test.go`, `debug_log_sink_test.go` (trace pairing, END
marker, `[ERROR]` prefix, fractional-second timestamp/elapsed_time, size
caps, and a real-pipeline test asserting the END-marker `dsl` carries
non-empty per-component `params.outputs` with chunks),
`debug_result_dsl_test.go` (flat and real nested `output["state"]`
shapes, vector stripping, format priority),
`debug_pages_integration_test.go`, `pipeline_executor_persist_test.go`,
`handler/agent_pipeline_debug_test.go`, `handler/agent_logs_test.go`
(incl. `TestRunCanvasPipelineDebug_ErrorStillExposesMessageID` /
`TestRespondWithDebugResult_ErrorCarriesMessageID` locking `message_id`
on failure), plus updates to `agent_test.go` / `agent_webhook_test.go` /
`chunker/image_upload_test.go` / `tokenizer*.go`.
- `go build ./...` and `./build.sh --test` for affected packages.

🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy)

---------

Co-authored-by: CodeBuddy Code <noreply@tencent.com>
2026-07-31 13:10:08 +08:00

93 lines
3.7 KiB
Go

//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//
package task
import (
"strings"
"github.com/google/uuid"
"ragflow/internal/entity"
)
// NewDebugTaskContext builds an in-memory TaskContext for a canvas-debug
// (dataflow dry-run) request. The returned context is intentionally
// side-effect free:
//
// - It carries no knowledgebase: KB.ID and Doc.KbID are empty. A canvas
// debug run has no KB, so kb_id == "" — this is the single debug signal
// used throughout the ingestion pipeline (the tokenizer skips embedding,
// the chunker skips image upload, the executor skips the persist stage).
// kb_id == "" occurs ONLY in debug mode; production ingestion always
// supplies a KB, so this branch is never reached in normal operation.
// - Doc.ID is a fresh uuid (a throwaway marker, not a persisted row); it
// only needs to be non-empty for the executor's input validation.
// - The parser page cap (debug preview only parses the first few pages)
// is injected at run time by the executor via Run's override_params
// channel, keyed by the Parser component's cpnID and the file's family
// (see injectDebugPageCap). It deliberately does NOT live in
// ParserConfig as a flat key, because flat keys are dropped by
// override_params merging and never reach the parser.
//
// fileData is the raw bytes of the uploaded debug document (may be nil for a
// file-less dry-run). It is stored on TaskContext.File and threaded into the
// pipeline run as the "file" / "binary" input.
//
// The caller-supplied kb_id is deliberately ignored: debug never resolves an
// embedder or writes to a KB, so a debug context has none.
func NewDebugTaskContext(tenantID, canvasID, fileName string, fileData []byte) *TaskContext {
doc := entity.Document{
ID: uuid.New().String(),
KbID: "",
Name: &fileName,
}
if suffix, docType := deriveDocSuffixAndType(fileName); suffix != "" {
doc.Suffix = suffix
doc.Type = docType
}
return &TaskContext{
Doc: doc,
KB: entity.Knowledgebase{ID: ""},
Tenant: entity.Tenant{ID: tenantID},
PipelineID: canvasID,
File: fileData,
}
}
// IsDebug reports whether this context is a canvas-debug (dry-run) context.
// A debug context carries no knowledgebase (KB.ID == "", forced by
// NewDebugTaskContext) — the single debug signal used across the ingestion
// pipeline. Production ingestion always supplies a KB (LoadFromIngestionTask
// follows doc -> kb), so KB.ID == "" occurs ONLY in debug mode. A debug run
// must produce no persistent side effect: the executor skips the persist
// stage, the tokenizer skips embedding, and the chunker skips image upload.
func (t *TaskContext) IsDebug() bool {
return t.KB.ID == ""
}
// deriveDocSuffixAndType extracts the file extension (e.g. ".pdf") and a
// lower-cased document type from a file name. It is best-effort: when the
// name has no usable extension both return "".
func deriveDocSuffixAndType(fileName string) (suffix, docType string) {
dot := strings.LastIndex(fileName, ".")
if dot < 0 || dot == len(fileName)-1 {
return "", ""
}
ext := strings.ToLower(fileName[dot+1:])
return "." + ext, ext
}