Files
ragflow/internal/service/parser_mode.go
Jack d12fd3b79d feat: parser pages range and parse type validation for dataset/document (#17293)
## Summary

Adds page-range parsing support to the Go-native pipeline path and
introduces strict `parse_type` validation for both dataset and document
update endpoints.

## What changed

### Pages range parsing
- **`internal/utility/pdf_pages.go`** — `NormalizePDFPages`: normalizes
raw page ranges (list of `[from,to]` 1-indexed inclusive ranges) into
sorted, merged, deduplicated `[][]int`. Invalid ranges are dropped.
- **`internal/ingestion/pipeline/pdf_pages.go`** —
`NormalizeParserConfigPages`: walks any parser_config map and normalizes
`"pages"` values under every component → filetype setup, so the
persisted config always carries clean, merged ranges.
- **`internal/deepdoc/parser/pdf/parser.go`** — integrates
`resolvePagesToProcess` to filter parsed PDF pages by the configured
ranges.
- Pipeline integration (parser pages):
`internal/parser/parser/pdf_parser_common.go`, `chunk_process.go`, plus
associated e2e and unit tests.

### Parse type validation (shared logic)
- **`internal/service/parser_mode.go`** (new) — `ValidateParseTypeMode`:
shared function that validates `parse_type` (1=BuiltIn/parser_id,
2=Pipeline/pipeline_id) and ensures the corresponding field is present.
Used by both dataset and document update endpoints.
- **`internal/service/dataset/crud.go`** / `update.go` — replaces inline
`isPipelineMode`/`isBuiltinMode` computation with the shared
`service.ValidateParseTypeMode`.
- **`internal/service/document/document_dataset_update.go`** — adds
strict `parse_type` validation in `validateDatasetDocumentUpdate`,
simplifies the reparse logic to a two-way switch (isBuiltin/isPipeline)
now that parse_type is always valid.
- **`internal/service/document/document.go`** — adds `ParseType` field
to `UpdateDatasetDocumentRequest`.
- **`internal/service/document/document_dataset_update.go`** —
`updateDocumentParserConfig` fallback path when DSL loading fails.
- **`internal/service/parser_mode_test.go`** (new) — test coverage for
nil, invalid, and missing-field scenarios.

### Frontend
- **`web/src/interfaces/request/document.ts`** — adds `parseType` to
`IChangeParserRequestBody`.
- **`web/src/hooks/use-document-request.ts`** —
`useSetDocumentPipelineParser` sends `parse_type` in the PATCH payload.
- **`web/src/pages/dataset/dataset/use-change-document-parser.ts`** —
Go/Python branching for the document parser config dialog.
-
**`web/src/components/document-pipeline-dialog/use-document-pipeline-form.ts`**
— `buildSubmitData` returns `parseType` (bugfix: was dropped from the
return value).

### Test changes
- **Removed**: 2 tests that verified the old "mutually exclusive" error
(replaced by `ValidateParseTypeMode` coverage).
- **Modified**: 6 tests across document and dataset packages to include
`ParseType` in request structs.
- **Added**: new e2e tests for pages parsing (`pages_e2e_test.go`,
`pdf_parser_pages_e2e_test.go`) and unit tests for `NormalizePDFPages`,
`NormalizeParserConfigPages`, `resolvePagesToProcess`.

## Backward compatibility
- The `parse_type` field is **required** when `parser_id` or
`pipeline_id` is sent. This changes the contract for both dataset and
document PATCH endpoints, but aligns the Go backend with the existing
frontend behavior (the frontend already sends `parse_type`). Callers
that omit `parse_type` when updating parser/pipeline selections will
receive a clear error message.
- Existing callers that only update fields like `name`, `enabled`, or
`meta_fields` are unaffected.
- Test updates ensure all known call sites are compliant.
2026-07-23 19:57:27 +08:00

90 lines
3.2 KiB
Go

package service
import (
"errors"
"fmt"
"strings"
)
// ValidateParseTypeMode validates parse_type and ensures the corresponding
// identifier field is present. Returns (isBuiltin, isPipeline, error).
// - parse_type must be 1 (BuiltIn) or 2 (Pipeline); nil is an error.
// - BuiltIn mode requires parserID to be present and non-empty.
// - Pipeline mode requires pipelineID to be present and non-empty.
func ValidateParseTypeMode(parseType *int, parserID, pipelineID *string) (isBuiltin bool, isPipeline bool, err error) {
if parseType == nil {
return false, false, errors.New("parse_type is required")
}
switch *parseType {
case 1:
if parserID == nil || strings.TrimSpace(*parserID) == "" {
return false, false, errors.New("parser_id is required when parse_type is BuiltIn")
}
return true, false, nil
case 2:
if pipelineID == nil || strings.TrimSpace(*pipelineID) == "" {
return false, false, errors.New("pipeline_id is required when parse_type is Pipeline")
}
return false, true, nil
default:
return false, false, fmt.Errorf("invalid parse_type: %d (must be 1 or 2)", *parseType)
}
}
// ParseModeState carries the persisted IDs that ResolveParseMode falls back to
// when the request does not switch modes.
type ParseModeState struct {
ParserID string
PipelineID *string
}
// ResolveParseMode returns the effective (isPipeline, parserID, pipelineID)
// after applying an update. It is the single source of truth for "which DSL
// should parser_config be cleaned against" and "which mode should reparse
// target", shared by the dataset and document update paths.
//
// parseType is authoritative when present:
// - 1 (BuiltIn): effParserID = reqParserID (or current); effPipelineID = nil
// so any prior canvas is cleared. reqPipelineID is ignored.
// - 2 (Pipeline): effPipelineID = reqPipelineID (or current); effParserID =
// current.ParserID (parser_id is not applicable in pipeline mode).
// reqParserID is ignored.
//
// When parseType is nil (no mode switch, e.g. only parser_config changed),
// inherit current mode and apply any incremental ID updates from the request;
// isPipeline is derived from whether effPipelineID is non-empty.
func ResolveParseMode(parseType *int, reqParserID, reqPipelineID *string,
current ParseModeState) (isPipeline bool, effParserID string, effPipelineID *string) {
if parseType != nil {
switch *parseType {
case 1: // BuiltIn
effParserID = current.ParserID
if reqParserID != nil {
if p := strings.TrimSpace(*reqParserID); p != "" {
effParserID = p
}
}
return false, effParserID, nil
case 2: // Pipeline
effPipelineID = current.PipelineID
if reqPipelineID != nil {
effPipelineID = reqPipelineID
}
return true, current.ParserID, effPipelineID
}
}
// No mode switch — inherit current state, apply incremental ID updates.
effParserID = current.ParserID
if reqParserID != nil {
if p := strings.TrimSpace(*reqParserID); p != "" {
effParserID = p
}
}
effPipelineID = current.PipelineID
if reqPipelineID != nil {
effPipelineID = reqPipelineID
}
isPipeline = effPipelineID != nil && strings.TrimSpace(*effPipelineID) != ""
return isPipeline, effParserID, effPipelineID
}