Commit Graph

815 Commits

Author SHA1 Message Date
euvre
91d9bf7fc4 fix: strip YAML frontmatter and defer API reference markdown rendering (#17272) 2026-07-24 11:37:10 +08:00
chanx
3dc298dbf5 fix(web): prevent multi-select popover from closing on first selection (#17310) 2026-07-24 09:29:40 +08:00
Jack
d12fd3b79d feat: parser pages range and parse type validation for dataset/document (#17293)
## Summary

Adds page-range parsing support to the Go-native pipeline path and
introduces strict `parse_type` validation for both dataset and document
update endpoints.

## What changed

### Pages range parsing
- **`internal/utility/pdf_pages.go`** — `NormalizePDFPages`: normalizes
raw page ranges (list of `[from,to]` 1-indexed inclusive ranges) into
sorted, merged, deduplicated `[][]int`. Invalid ranges are dropped.
- **`internal/ingestion/pipeline/pdf_pages.go`** —
`NormalizeParserConfigPages`: walks any parser_config map and normalizes
`"pages"` values under every component → filetype setup, so the
persisted config always carries clean, merged ranges.
- **`internal/deepdoc/parser/pdf/parser.go`** — integrates
`resolvePagesToProcess` to filter parsed PDF pages by the configured
ranges.
- Pipeline integration (parser pages):
`internal/parser/parser/pdf_parser_common.go`, `chunk_process.go`, plus
associated e2e and unit tests.

### Parse type validation (shared logic)
- **`internal/service/parser_mode.go`** (new) — `ValidateParseTypeMode`:
shared function that validates `parse_type` (1=BuiltIn/parser_id,
2=Pipeline/pipeline_id) and ensures the corresponding field is present.
Used by both dataset and document update endpoints.
- **`internal/service/dataset/crud.go`** / `update.go` — replaces inline
`isPipelineMode`/`isBuiltinMode` computation with the shared
`service.ValidateParseTypeMode`.
- **`internal/service/document/document_dataset_update.go`** — adds
strict `parse_type` validation in `validateDatasetDocumentUpdate`,
simplifies the reparse logic to a two-way switch (isBuiltin/isPipeline)
now that parse_type is always valid.
- **`internal/service/document/document.go`** — adds `ParseType` field
to `UpdateDatasetDocumentRequest`.
- **`internal/service/document/document_dataset_update.go`** —
`updateDocumentParserConfig` fallback path when DSL loading fails.
- **`internal/service/parser_mode_test.go`** (new) — test coverage for
nil, invalid, and missing-field scenarios.

### Frontend
- **`web/src/interfaces/request/document.ts`** — adds `parseType` to
`IChangeParserRequestBody`.
- **`web/src/hooks/use-document-request.ts`** —
`useSetDocumentPipelineParser` sends `parse_type` in the PATCH payload.
- **`web/src/pages/dataset/dataset/use-change-document-parser.ts`** —
Go/Python branching for the document parser config dialog.
-
**`web/src/components/document-pipeline-dialog/use-document-pipeline-form.ts`**
— `buildSubmitData` returns `parseType` (bugfix: was dropped from the
return value).

### Test changes
- **Removed**: 2 tests that verified the old "mutually exclusive" error
(replaced by `ValidateParseTypeMode` coverage).
- **Modified**: 6 tests across document and dataset packages to include
`ParseType` in request structs.
- **Added**: new e2e tests for pages parsing (`pages_e2e_test.go`,
`pdf_parser_pages_e2e_test.go`) and unit tests for `NormalizePDFPages`,
`NormalizeParserConfigPages`, `resolvePagesToProcess`.

## Backward compatibility
- The `parse_type` field is **required** when `parser_id` or
`pipeline_id` is sent. This changes the contract for both dataset and
document PATCH endpoints, but aligns the Go backend with the existing
frontend behavior (the frontend already sends `parse_type`). Callers
that omit `parse_type` when updating parser/pipeline selections will
receive a clear error message.
- Existing callers that only update fields like `name`, `enabled`, or
`meta_fields` are unaffected.
- Test updates ensure all known call sites are compliant.
2026-07-23 19:57:27 +08:00
buua436
d4a8c91f3c refa: improve tree clustering (#17285) 2026-07-23 17:49:13 +08:00
euvre
418d3c8cef fix: fetch all knowledge bases via pagination in link-to-dataset dialog (#17170) 2026-07-23 16:48:05 +08:00
balibabu
c48a59db96 Feat: Add page rank to ParserForm (#17267) 2026-07-23 14:23:07 +08:00
balibabu
b3d394954d Feat: Support pipeline-related configuration at the document level. (#17212)
### Summary

Feat: Support pipeline-related configuration at the document level.
2026-07-22 22:24:50 +08:00
euvre
d757303fa4 fix(web): show completed label for think block after streaming ends (#17149) 2026-07-22 19:30:09 +08:00
euvre
5bc28d4d15 fix: exclude TTS models from agent component model selector (#17231) 2026-07-22 19:27:12 +08:00
euvre
cfaef879fe fix: MinerU.Net and PaddleOCR.Net missing icons (#17229) 2026-07-22 19:24:23 +08:00
euvre
2f7dc60337 fix: restore download button when file type is selected in agent message (#17184) 2026-07-22 19:21:17 +08:00
euvre
dabe426c36 fix: render documentation URLs as hyperlinks in model tooltips (#17224) 2026-07-22 16:03:03 +08:00
euvre
a81b5057bb fix(web): show builtin chunk methods in file pipeline dialog under Go backend (#17169) 2026-07-22 15:14:02 +08:00
euvre
ade2f3a5ab fix(web): use outline-none instead of outline-0 on focusable controls (#17165) 2026-07-22 14:57:23 +08:00
balibabu
20b760f266 Feat: Highlight wiki force graph. (#17160) 2026-07-21 16:48:52 +08:00
balibabu
b288888050 Feat: Add a "None" option for reasoning intensity in the chat message box. (#17006)
### Summary

Feat: Add a "None" option for reasoning intensity in the chat message
box.
2026-07-21 11:50:34 +08:00
euvre
ee8aa03fa0 fix(web): regenerate in multi-model chat uses the card's own state and model (#17126) 2026-07-21 10:40:59 +08:00
euvre
33765c1a64 Fix: add missing model provider icons (#17122) 2026-07-20 19:30:26 +08:00
chanx
b3f731d4de Fix: model selection format compat and provider api_key wrapping (#17112) 2026-07-20 19:12:59 +08:00
balibabu
392b249404 Feat: Render the skills list using a tree view. (#17115)
### Summary
Feat: Render the skills list using a tree view.
2026-07-20 19:00:08 +08:00
euvre
3fa890294a fix: selected destination appears white in file move dialog in dark mode (#17085) 2026-07-20 16:47:48 +08:00
balibabu
9d850e782b Feat: Configure the relevant pipeline parameters on the dataset configuration page. (#17100) 2026-07-20 16:05:03 +08:00
euvre
67a2310ea4 fix: prevent share pages from overriding the user's theme preference (#17098) 2026-07-20 15:46:47 +08:00
balibabu
677960716e Feat: Enable the wiki's Markdown editor to navigate to a new Markdown file when a link is clicked. (#17019) 2026-07-17 11:22:18 +08:00
Jack
3d65e8d124 Revert frontend changes from PR #16873 (#16948) 2026-07-15 17:49:37 +08:00
Jack
e543ff02c4 Feat: nats message processing refactor and support built in DSL (#16873) 2026-07-15 14:53:16 +08:00
chanx
e5bc644c18 fix(tree-select): avoid incorrect node selection when value is empty (#16898)
### Summary

fix(tree-select): avoid incorrect node selection when value is empty

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-14 15:53:54 +08:00
balibabu
654fc8851a Feat: Add a three-level dropdown menu to the talk page message box. (#16886)
### Summary

Feat: Add a three-level dropdown menu to the talk page message box.
2026-07-14 13:45:31 +08:00
balibabu
999eb533a9 Feat: Adjust the color and size of the graph based on the data. (#16868) 2026-07-14 10:20:31 +08:00
balibabu
45862bf95d Feat: If the interval between two outputs exceeds 600ms, a loading state is displayed at the end. (#16861)
### Summary

Feat: If the interval between two outputs exceeds 600ms, a loading state
is displayed at the end.
2026-07-13 18:02:29 +08:00
euvre
3bfad1f00e fix: correct model type mappings and improve system setting persistence (#16501) 2026-07-13 16:42:55 +08:00
chanx
547bc86141 Fixed an issue where cited webpages could not be opened during online searches. (#16840) 2026-07-13 15:25:57 +08:00
balibabu
2b9569ff51 Feat: Added support for session graph and session essence templates. (#16851)
### Summary

Feat: Added support for session graph and session essence templates.
2026-07-13 14:48:02 +08:00
balibabu
d317742975 Feat: Rewrite wiki template with reui (#16797) 2026-07-10 15:44:04 +08:00
chanx
868e524f29 fix: pass ownerTenantId to LLMLabel and related components for improved model fetching (#16800) 2026-07-10 13:27:14 +08:00
chanx
8236f2cabb Fix: update LayoutRecognizeFormField to accept ownerTenantId and refactor model handling in LLM requests (#16781) 2026-07-10 09:36:25 +08:00
chanx
0095fa048f fix: Show full text on hover when text overflows the cards on the list page. (#16787) 2026-07-10 09:25:05 +08:00
balibabu
0083ad0deb Feat: Add a data compilation layer. (#16777)
### Summary

Feat: Add a data compilation layer.
2026-07-09 17:49:16 +08:00
chanx
a9420a7832 Fix: resolve shared embedding/LLM model selection errors (#16773) 2026-07-09 15:17:49 +08:00
Lynn
1430d0e431 Fix: provider name (#16733) 2026-07-09 10:19:10 +08:00
balibabu
575984877f Fix: Rapid clicking results in multiple message requests being sent. (#16739) 2026-07-09 09:57:54 +08:00
euvre
3ec9187cd2 fix(web): prevent 'last saved at' label from vertical stacking in agent home card (#16756) 2026-07-09 09:46:03 +08:00
chanx
080dd84fed Feat: apply prose typography styling to markdown preview (#16752) 2026-07-09 09:45:52 +08:00
euvre
a41fef49d0 fix(web): hide folder tab in agent JSON import uploader (#16754) 2026-07-08 20:11:55 +08:00
chanx
dd2f27d6a3 fix: Restrict the agent to using memory compatible with the embedding model. (#16699) 2026-07-07 16:28:58 +08:00
chanx
f082675e6f Fix: Prevent text overflow in confirm delete dialog (#16689) 2026-07-07 14:50:25 +08:00
Rodger Blom
d8cefcf052 feat: add native Dutch language support for BM25 tokenization (#14140)
## Summary
- Add language-aware Snowball stemmer to `RagTokenizer` supporting 16
languages (Dutch, German, French, Spanish, etc.)
- Thread the KB `language` parameter through the full tokenization
pipeline (14 parser modules + task executor)
- Add Dutch to the frontend language lists and cross-language form

## Problem
RAGFlow uses the English Porter stemmer + WordNet lemmatizer for **all**
BM25 tokenization, regardless of the knowledge base language setting.
This produces incorrect stems for non-English text. For example:

| Dutch word | Dutch stemmer | English Porter |
|---|---|---|
| documenten | document | documenten (unchanged!) |
| gebruikers | gebruiker | gebruik (over-stemmed) |
| instellingen | instell | instellingen (unchanged!) |

This degrades BM25 recall for any non-English knowledge base.

## Solution
NLTK already ships Snowball stemmers for 16 languages. This PR:

1. **`rag/nlp/rag_tokenizer.py`**: Overrides `tokenize()` with
`set_language()` and `_normalize_token()` that selects the correct NLTK
Snowball stemmer. Falls back to Porter for unmapped languages (Chinese,
Japanese, Korean, etc. — these use character-based tokenization anyway).
2. **`rag/nlp/__init__.py`** + **14 `rag/app/*.py` parsers** +
**`rag/svr/task_executor.py`**: Threads the `language` parameter through
`tokenize()`, `tokenize_chunks()`, `tokenize_table()`, and all callers.
3. **Frontend**: Adds Dutch (`Nederlands`) to `LanguageList`,
`LanguageMap`, `LanguageAbbreviationMap`, `LanguageTranslationMap`,
cross-language form field, and `en.ts` locale.

## Backward Compatibility
- Default language is `"English"`, preserving existing behavior for all
current users
- Languages without a Snowball stemmer mapping fall back to Porter (no
change)
- No new dependencies — NLTK Snowball is already bundled
2026-07-06 23:39:56 +08:00
euvre
81cfcdf2d3 feat(frontend): add AuthenticatedImg component for authorized image requests (#16525) 2026-07-01 17:02:44 +08:00
Lynn
400476f0b3 Feat: SoMark (#16482)
Follow #15486
Co-authored-by: limuting <limuting233@gmail.com>
Co-authored-by: lutianyi <lutianyi233@163.com>
Co-authored-by: justinychuang <huangyicheng@soulcode.cn>
Co-authored-by: maybehokori <138367708+maybehokori@users.noreply.github.com>
2026-07-01 13:29:28 +08:00
Lynn
b53b693f22 Fix: CI (#16504)
### Summary

Fix race condition in parallel lefthook hooks causing ETXTBSY error
2026-06-30 22:14:11 +08:00