mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-08-02 05:47:31 +08:00
## Summary
Adds a side-effect-free DataFlow canvas **debug (dry-run) mode** plus a
**debug run log with a "View result" panel**, so a canvas can be
executed synchronously and inspected end-to-end (per-component progress
and parsed chunks) without persisting anything.
### Dry-run execution (inline parsed chunks)
- `task/debug.go`: `NewDebugTaskContext` builds an in-memory
`TaskContext` with **`KB.ID == ""` — the single debug signal used across
the ingestion pipeline**. A canvas debug run has no knowledgebase, and
production ingestion always supplies one, so `kb_id == ""` occurs ONLY
in debug mode. Components gate their own side effects on this signal
without any dedicated debug vocabulary (the former `CANVAS_DEBUG_DOC_ID`
marker constant is removed).
- `task/pipeline_executor.go`: `validateTaskContext` no longer requires
a KB when `KB.ID == ""` (debug); debug runs return `collectDebugOutput`
(chunks) instead of a no-op; uploaded bytes are delivered as
`inputs['binary']` for doc-less runs; `injectDebugPageCap` caps the
parser to the first pages for a fast preview via the production
`override_params` channel (Parser cpnID + family).
- `component/tokenizer.go`: `shouldHaveEmbedding` skips embedding when
`kb_id == ""` — the embedder is configured on the knowledgebase, so a
debug run has nothing to resolve against and stays side-effect free.
- `chunker/register.go`: chunk images are uploaded to MinIO only when a
KB is present (persist run). **In debug mode the raw image bytes are
intentionally dropped (`delete(ck, "image")`)** — the debug preview does
not render chunk images, and dropping the bytes keeps them out of memory
and out of the Redis-stored debug log. This is a deliberate trade-off,
not an oversight.
- `component/file.go`: pass through in-memory binary bytes, skipping
`doc_id` -> storage resolution.
- `handler/agent.go` + `agent_webhook.go`: detect `dataflow_canvas` and
run a sync debug returning chunks inline on the existing
chat/completions endpoint; reject DataFlow canvases from webhooks (fixes
the previously dead `== "DataFlow"` check; mirrors Python
`agent_api.py`).
- `parser_dispatch.go`: export `ParserFileFamily` for the executor's
page-cap injection.
### Debug run log + "View result"
Mirrors Python's debug-log contract so the front-end can replay each
component's progress and parsed output:
- `task/debug_log_sink.go`: a `DebugLogSink` records every component's
lifecycle into a `[{component_id, trace}]` array (each trace entry
carries `message`, `progress`, `timestamp`, `elapsed_time`). `Flush`
appends a terminal `END` marker whose first trace message is non-empty
so the front-end detects completion. On failure the END marker is
prefixed `[ERROR]` yet still carries the run, so the failure timeline
renders instead of being stuck empty. Timestamps and `elapsed_time` are
in seconds (matching the rest of the app).
- `task/debug_result_dsl.go`: `BuildDebugResultDSL` builds the `dsl` the
END marker carries — the Go analogue of Python's `Graph.__str__` +
END-marker `dsl` in `rag/flow/pipeline.py`. It combines the static DSL
structure (component_name / downstream / params / graph.nodes) with the
run output map (`output["state"][<id>]`) to emit, per component,
`obj.params.outputs[<format>].value` (chunks / text / json / html /
markdown) — the exact keys the front-end `dataflow-result` page reads to
render each step's parsed chunks. Raw embedding vectors (including the
dimension-scoped `q_<dim>_vec` keys) are stripped so the stored log
stays Python-scale.
- `task/pipeline_executor.go`: after the run, attach the built `dsl` to
the END marker via the `ResultSink` capability.
- `handler/agent.go`: `runCanvasPipelineDebug` generates a stable
`message_id` up-front and always flushes the log (success or failure);
`respondWithDebugResult` returns `message_id` in **both** the success
and the error envelope so the front-end can poll the log. The debug-log
endpoint `GET /agents/:id/logs/:message_id` serves the array.
- `web/src/pages/agent/hooks/use-run-dataflow.ts`: on a run failure,
also surface `message_id` via `setMessageId` so the log sheet renders
the failure timeline (the `[ERROR]` END marker is already written).
Guarded by `if (msgId)`, so it is a safe no-op when the back-end does
not return an id.
## Behavioral notes
- Debug parses only the first pages (`debugPageCapPages`) for a fast
preview; an explicit `pages` cap already present in the ParserConfig is
respected.
- Debug mode does not keep chunk images (see above) and does not compute
embeddings — it exercises parse + chunk only.
## Test plan
- Go: `debug_test.go`, `debug_log_sink_test.go` (trace pairing, END
marker, `[ERROR]` prefix, fractional-second timestamp/elapsed_time, size
caps, and a real-pipeline test asserting the END-marker `dsl` carries
non-empty per-component `params.outputs` with chunks),
`debug_result_dsl_test.go` (flat and real nested `output["state"]`
shapes, vector stripping, format priority),
`debug_pages_integration_test.go`, `pipeline_executor_persist_test.go`,
`handler/agent_pipeline_debug_test.go`, `handler/agent_logs_test.go`
(incl. `TestRunCanvasPipelineDebug_ErrorStillExposesMessageID` /
`TestRespondWithDebugResult_ErrorCarriesMessageID` locking `message_id`
on failure), plus updates to `agent_test.go` / `agent_webhook_test.go` /
`chunker/image_upload_test.go` / `tokenizer*.go`.
- `go build ./...` and `./build.sh --test` for affected packages.
🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy)
---------
Co-authored-by: CodeBuddy Code <noreply@tencent.com>
214 lines
7.1 KiB
Go
214 lines
7.1 KiB
Go
//go:build integration
|
|
|
|
//
|
|
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
|
|
//
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
// you may not use this file except in compliance with the License.
|
|
// You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
package task
|
|
|
|
import (
|
|
"context"
|
|
"encoding/json"
|
|
"fmt"
|
|
"os"
|
|
"path/filepath"
|
|
"strings"
|
|
"testing"
|
|
"time"
|
|
|
|
"ragflow/internal/dao"
|
|
"ragflow/internal/entity"
|
|
"ragflow/internal/ingestion/component"
|
|
"ragflow/internal/ingestion/pipeline"
|
|
"ragflow/internal/storage"
|
|
)
|
|
|
|
// TestExecute_DebugViaEntry_HonorsPagesCap_Integration is the end-to-end
|
|
// proof that the canvas-debug page cap travels through Run's override_params
|
|
// channel (NOT pipeline inputs) and is honoured by the deepdoc/pdf parser.
|
|
//
|
|
// It parses a multi-page PDF twice through a real debug pipeline (real DSL,
|
|
// real pdfium parser, in-memory sqlite + storage):
|
|
//
|
|
// - Uncapped baseline: an explicit cpnID+family page cap of [[1, 1000000]]
|
|
// is supplied in ParserConfig. injectDebugPageCap must RESPECT it, so the
|
|
// parser reads every page.
|
|
// - Capped: no page cap is supplied, so injectDebugPageCap injects the
|
|
// debug default [[1, debugPageCapPages]], and the parser reads only the
|
|
// leading pages.
|
|
//
|
|
// The capped run must therefore yield strictly fewer chunk characters than
|
|
// the uncapped run. This exercises the full path the production code uses
|
|
// (Run's 3rd-arg override_params → applyOverrideParams → c.Setups →
|
|
// ConfigureFromSetup → deepdoc pdf Config.Pages).
|
|
//
|
|
// This is gated by //go:build integration because it depends on the native
|
|
// pdfium library, the 03_multipage.pdf fixture, and (optionally) the
|
|
// tokenizer resource dict.
|
|
func TestExecute_DebugViaEntry_HonorsPagesCap_Integration(t *testing.T) {
|
|
requireTokenizerPool(t)
|
|
mustLoadTaskTestConfig(t)
|
|
|
|
origDB := dao.DB
|
|
realDB := mustOpenTaskTestDB(t)
|
|
dao.DB = realDB
|
|
t.Cleanup(func() { dao.DB = origDB })
|
|
|
|
realStorage := storage.NewMemoryStorage()
|
|
origStorage := storage.GetStorageFactory().GetStorage()
|
|
storage.GetStorageFactory().SetStorage(realStorage)
|
|
t.Cleanup(func() { storage.GetStorageFactory().SetStorage(origStorage) })
|
|
|
|
templatePath := filepath.Join(taskRepoRoot(t), "internal", "ingestion", "pipeline", "template", "ingestion_pipeline_general.json")
|
|
templateBytes, err := os.ReadFile(templatePath)
|
|
if err != nil {
|
|
t.Fatalf("read template: %v", err)
|
|
}
|
|
templateBytes = disableTokenizerEmbeddingForTaskTemplate(t, templateBytes)
|
|
|
|
// envelopeDSL is stored verbatim on the canvas (this is what production
|
|
// persists), so loadDSLFromCanvas marshals the envelope and Run receives
|
|
// it. injectDebugPageCap must unwrap it to find the Parser cpnID.
|
|
var envelope struct {
|
|
DSL json.RawMessage `json:"dsl"`
|
|
}
|
|
if err := json.Unmarshal(templateBytes, &envelope); err != nil {
|
|
t.Fatalf("unmarshal template envelope: %v", err)
|
|
}
|
|
var canvasDSL entity.JSONMap
|
|
if err := json.Unmarshal(templateBytes, &canvasDSL); err != nil {
|
|
t.Fatalf("unmarshal template into canvas dsl: %v", err)
|
|
}
|
|
|
|
// Discover the Parser cpnID from the inner DSL.
|
|
schemas, err := pipeline.ExtractAllComponentParams(envelope.DSL)
|
|
if err != nil {
|
|
t.Fatalf("ExtractAllComponentParams: %v", err)
|
|
}
|
|
var parserCpnID string
|
|
for _, s := range schemas {
|
|
if s.ComponentName == component.ComponentNameParser {
|
|
parserCpnID = s.CpnID
|
|
}
|
|
}
|
|
if parserCpnID == "" {
|
|
t.Fatal("template has no Parser component")
|
|
}
|
|
|
|
pdfPath := filepath.Join(taskRepoRoot(t), "internal", "deepdoc", "parser", "pdf", "testdata", "pdfs", "03_multipage.pdf")
|
|
pdfBytes, err := os.ReadFile(pdfPath)
|
|
if err != nil {
|
|
t.Skipf("read pdf fixture %s: %v", pdfPath, err)
|
|
return
|
|
}
|
|
|
|
suffix := fmt.Sprintf("%d", time.Now().UnixNano())
|
|
tenantID := taskLimit32("it_tenant_" + suffix)
|
|
canvasID := taskLimit32("it_canvas_" + suffix)
|
|
|
|
if err := realDB.Create(&entity.UserCanvas{
|
|
ID: canvasID,
|
|
UserID: tenantID,
|
|
Permission: "me",
|
|
CanvasCategory: "agent_canvas",
|
|
DSL: canvasDSL,
|
|
}).Error; err != nil {
|
|
t.Fatalf("create user canvas: %v", err)
|
|
}
|
|
t.Cleanup(func() { _ = realDB.Where("id = ?", canvasID).Delete(&entity.UserCanvas{}).Error })
|
|
|
|
// Explicit "parse all pages" cap (JSON-decoded []any form, the shape the
|
|
// parser actually consumes). injectDebugPageCap must respect it and leave
|
|
// it untouched, so the parser reads every page of the PDF.
|
|
allPages := map[string]any{
|
|
parserCpnID: map[string]any{
|
|
"pdf": map[string]any{
|
|
"pages": []any{[]any{1, 1000000}},
|
|
},
|
|
},
|
|
}
|
|
|
|
newDebugCtx := func(parserConfig map[string]any) *TaskContext {
|
|
// A canvas-debug (dry-run) context carries no KB: KB.ID == "" is the
|
|
// single debug signal. The executor then skips the persist stage and
|
|
// injects the debug page cap (see injectDebugPageCap, gated on
|
|
// KB.ID == "").
|
|
return &TaskContext{
|
|
Doc: entity.Document{
|
|
ID: "debug-run-1",
|
|
KbID: "",
|
|
Name: taskStrPtr("03_multipage.pdf"),
|
|
Type: "pdf",
|
|
ParserConfig: parserConfig,
|
|
},
|
|
KB: entity.Knowledgebase{
|
|
ID: "",
|
|
TenantID: tenantID,
|
|
EmbdID: "embd-1",
|
|
},
|
|
Tenant: entity.Tenant{ID: tenantID},
|
|
File: pdfBytes,
|
|
PipelineID: canvasID,
|
|
}
|
|
}
|
|
|
|
// Uncapped baseline: explicit pages=all → executor respects override.
|
|
uncappedExec, err := NewPipelineExecutor(newDebugCtx(allPages), canvasID, 0)
|
|
if err != nil {
|
|
t.Fatalf("NewPipelineExecutor (uncapped): %v", err)
|
|
}
|
|
uncapped, err := uncappedExec.Execute(context.Background())
|
|
if err != nil {
|
|
t.Fatalf("Execute (uncapped): %v", err)
|
|
}
|
|
|
|
// Capped: no pages override → executor injects [[1, debugPageCapPages]].
|
|
cappedExec, err := NewPipelineExecutor(newDebugCtx(nil), canvasID, 0)
|
|
if err != nil {
|
|
t.Fatalf("NewPipelineExecutor (capped): %v", err)
|
|
}
|
|
capped, err := cappedExec.Execute(context.Background())
|
|
if err != nil {
|
|
t.Fatalf("Execute (capped): %v", err)
|
|
}
|
|
|
|
uncappedLen := len(joinedChunks(uncapped.Chunks))
|
|
cappedLen := len(joinedChunks(capped.Chunks))
|
|
if uncappedLen == 0 {
|
|
t.Fatal("uncapped baseline produced no chunks; fixture/pipeline misconfigured")
|
|
}
|
|
if cappedLen == 0 {
|
|
t.Fatal("pages cap produced no chunks; pages wiring broken")
|
|
}
|
|
if cappedLen >= uncappedLen {
|
|
t.Errorf("pages cap had no effect: capped text len %d >= uncapped %d (pages not honoured via override_params)", cappedLen, uncappedLen)
|
|
}
|
|
}
|
|
|
|
// joinedChunks concatenates the text of every chunk into a single string so
|
|
// the page-cap assertion does not depend on how the chunker splits pages.
|
|
func joinedChunks(chunks []map[string]any) string {
|
|
var b strings.Builder
|
|
for _, ck := range chunks {
|
|
for _, key := range []string{"content_with_weight", "content", "text"} {
|
|
if v, ok := ck[key].(string); ok && v != "" {
|
|
b.WriteString(v)
|
|
b.WriteString("\n")
|
|
break
|
|
}
|
|
}
|
|
}
|
|
return b.String()
|
|
}
|