mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-07-21 07:01:04 +08:00
ddd7fc02eec8a5812e889715f10214439d2c813f
13 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8f87e19120 |
feat(llm): add AWS Bedrock reranker connector (#16960)
## What Adds a reranker connector for the **Bedrock** factory, which previously offered chat/embedding/CV models but no reranker — selecting a Bedrock rerank model raised `Factory not in rerank model`. ## How `BedrockRerank` calls the `bedrock-agent-runtime` Rerank API. It reuses the same JSON key protocol as `BedrockEmbed` (`auth_mode` / `bedrock_region` / `bedrock_ak` / `bedrock_sk`, with `access_key_secret` / `iam_role` / `assume_role` modes). Documents are truncated to the model window (Cohere Rerank v3.5 ~2k of its shared 4k window, Amazon Rerank v1 8k) on top of Bedrock's own internal truncation. Scores are returned in `[0, 1]`, so the shared `Base.similarity` normalization applies unchanged. Verified against `amazon.rerank-v1:0` and `cohere.rerank-v3-5:0` in `eu-central-1`. > Note: this PR adds the connector only. Bedrock rerank models can be selected by > adding the relevant entries to `conf/llm_factories.json` under the Bedrock > provider; that catalog change is intentionally left out of this PR. ## Tests `test/unit_test/rag/llm/test_bedrock_rerank.py` — boto3 is mocked (no AWS call): score-by-index mapping, per-model document truncation, model ARN construction, auth-mode validation and the empty-input short-circuit. `pytest` green alongside the existing reranker normalization suite. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
eeb59ec4f2 |
feat(stt): support Fun-ASR-Flash in Tongyi-Qianwen provider (#16844)
## What this PR does Adds support for Alibaba Cloud's hosted Fun-ASR-Flash snapshots to the existing Tongyi-Qianwen speech-to-text provider. - registers `fun-asr-flash-2026-06-15` as a speech-to-text model; - routes only `fun-asr-flash*` models to the documented workspace-native multimodal-generation endpoint; - supports local audio through size-checked data URIs as well as URL/data-URI inputs; - uses the documented SSE response mode for incremental streaming transcription; - closes the streamed HTTP response on completion, failure, or early consumer cancellation; - preserves the existing `dashscope.MultiModalConversation` path for all other Qwen audio models; - keeps RAGFlow's existing synchronous and streaming adapter interfaces. ## Why Fun-ASR-Flash does not use the legacy Qwen audio request shape currently used by `QWenSeq2txt`. Its synchronous API expects `input_audio` at: `/api/v1/services/aigc/multimodal-generation/generation` Without a narrowly scoped adapter path, the hosted model cannot be selected successfully through RAGFlow's Tongyi-Qianwen speech-to-text provider. Closes #16843. ## Compatibility The new behavior is gated by the `fun-asr-flash` model-name prefix. Existing Qwen audio models continue through the original code path unchanged. ## Validation - `pytest test/unit_test/rag/llm/test_sequence2txt_model.py`: 10 passed - Ruff check: passed - Ruff format check: passed - `llm_factories.json` validation: passed - Real hosted-API validation with WAV audio - Real RAGFlow upload/indexing validation with MP3 audio The unit tests cover the native Fun-ASR-Flash request, regression behavior for the legacy Qwen path, SSE streaming, and early response cleanup. ## Documentation - https://help.aliyun.com/document_detail/2979031.html - https://help.aliyun.com/document_detail/2869541.html ### Why a dedicated adapter path is necessary (official evidence) Alibaba Cloud's [Fun-ASR RESTful API reference](https://help.aliyun.com/en/model-studio/fun-asr-recorded-speech-recognition-http-api) makes the incompatibilities with RAGFlow's existing Qwen audio path explicit: | Adapter change | Official API requirement | Why the existing path is insufficient | | --- | --- | --- | | Call the workspace-native HTTP endpoint | The Fun-ASR-Flash synchronous section states that SDK calls are not supported and specifies `POST /api/v1/services/aigc/multimodal-generation/generation`. | The existing adapter calls `dashscope.MultiModalConversation`, so a direct HTTP path is required. | | Use the `input_audio` message shape | `input.messages`, `content`, `type: input_audio`, `input_audio`, and `input_audio.data` are documented as required for an audio request. | The existing Qwen path sends the legacy `audio` content shape, which does not match this API contract. | | Send `parameters.format` | The request schema marks `parameters` and `format` as **Required**, and says the value must match the actual audio format. | The legacy request has no Fun-ASR-Flash `parameters.format` field, so the adapter must derive and send it. | | Encode local files as Data URIs | `input_audio.data` accepts either a public URL or a Base64 Data URI; the reference gives the exact `data:{MIME_TYPE};base64,...` form. | RAGFlow supplies local file paths, which the remote API cannot read directly. | | Parse `output.text` | The documented non-streaming response returns the accumulated transcription in `output.text`. | The legacy Qwen response parser reads `output.choices[].message.content`, so a separate response parser is required. | | Enforce the Base64 input limit | The reference requires the Base64-encoded audio to remain within the 10 MB input limit. | The adapter checks encoded size before reading/sending local audio and directs oversized inputs to the existing public-URL path. | | Use SSE for streaming | The reference specifies `X-DashScope-SSE: enable` and documents intermediate and final SSE events. | The adapter parses those events instead of wrapping one blocking response as a synthetic stream. | | Release streamed responses | Streaming responses must be closed when iteration completes or stops early. | A `finally` cleanup releases the HTTP response on completion, errors, and consumer cancellation. | `sample_rate` is documented as **Optional**. The implementation omits it instead of declaring a fixed value that may not match remote or compressed audio. The [official speech-to-text model list](https://help.aliyun.com/en/model-studio/asr-model/) separately confirms that `fun-asr-flash-2026-06-15` is an offline HTTP model with a five-minute audio limit. --------- Signed-off-by: LauraGPT <LauraGPT@users.noreply.github.com> Co-authored-by: openhands <openhands@all-hands.dev> Co-authored-by: LauraGPT <LauraGPT@users.noreply.github.com> |
||
|
|
06e36d24f4 |
feat(stt): add FunASR / SenseVoice provider (#16473)
### Summary Adds FunASR as a self-hosted speech-to-text provider through its OpenAI-compatible `/v1/audio/transcriptions` endpoint. This is a focused replacement for #15526 by @Rene0422 and relates to #15448. The unrelated Markdown parser changes from the previous branch are intentionally removed so this PR contains only the FunASR provider integration. - register FunASR as a `SPEECH2TEXT` factory; - add `FunASRSeq2txt` with `sensevoice` and `http://localhost:8000/v1` defaults, an optional API key, URL normalization, and inherited transcription handling; - wire FunASR into the current local-provider schema with a prefilled local URL and official documentation link; - discover the server's `/v1/models` dynamically and expose every returned model as speech-to-text in the model picker; - use RAGFlow's existing default provider icon fallback instead of referencing a missing `funasr` asset; - list FunASR in the supported-provider documentation; - add focused backend and frontend regression tests. ### Validation - focused backend pytest suite -> `7 passed` - real CPU `funasr-server` + RAGFlow provider smoke test -> discovered `fun-asr-nano`, `sensevoice`, and `paraformer`; transcribed a real WAV as `我现在在录一段测试音频` (`10` tokens, `0.504s`) - `ruff check` and `ruff format --check` on the changed Python files - `python3 -m py_compile` on the provider and its test - JSON parse and a semantic assertion for exactly one enabled FunASR `SPEECH2TEXT` factory - focused frontend Jest test -> `2 passed` - ESLint and Prettier on all changed TypeScript files - `npm run build` -> production build succeeded (`14,181` modules transformed) - `git diff --check` ### Deployment Run FunASR separately and point the RAGFlow provider at it: ```bash pip install funasr funasr-server --device cuda --model sensevoice ``` The API key remains optional because the stock local server does not require authentication. A key can still be supplied when the endpoint is protected by a gateway. --------- Signed-off-by: LauraGPT <LauraGPT@users.noreply.github.com> Co-authored-by: LauraGPT <LauraGPT@users.noreply.github.com> |
||
|
|
6a4b9be426 | Refactor: reformat all code for lefthook using ruff and gofmt (#16585) | ||
|
|
0d7ad0ed0c |
Feat/agent thinking switch (#15446)
### What problem does this PR solve? This PR adds an Agent LLM setting to control thinking mode for official providers that expose a thinking switch. Related to #12842. Closes #15445. Some providers expose thinking controls through provider-specific request fields, but Agent LLM settings did not have a unified option for users to enable or disable thinking mode. This PR adds a `Thinking` selector with: - System default - Enabled - Disabled <img width="452" height="278" alt="8566b0b4-0546-4c8a-913d-f9bbd38319f6" src="https://github.com/user-attachments/assets/25b497f7-1ba0-4bfe-940d-6fe79287d6ab" /> <img width="471" height="971" alt="8a0a6bee-f45f-48d5-bd83-17af260de3db" src="https://github.com/user-attachments/assets/41ad43c1-5087-48f1-bf37-f2ca14c2be2f" /> Initial support is limited to the verified official providers: - Qwen / DashScope: `enable_thinking` - Kimi / Moonshot: `thinking.type` - GLM / ZHIPU-AI: `thinking.type` For LiteLLM-based providers, provider-specific fields are forwarded through `extra_body` before `drop_params` filtering so the request parameters are preserved. ### Type of change - [x] New Feature (non-breaking change which adds functionality) --------- Co-authored-by: jiashi <jiashi19@outlook.com> Co-authored-by: Zhichang Yu <yuzhichang@gmail.com> |
||
|
|
bde2b1fc6d |
fix(llm): correct error handling, token accounting, and truncation in embedding providers (#15424)
### Summary Closes #15423 `rag/llm/embedding_model.py` hosts about 40 embedding providers that shared several defects affecting indexing reliability, cost accounting, and error visibility. This PR fixes four concrete bugs. **Masked, inconsistent errors (27 sites).** Nearly every provider ran `log_exception(_e, res)` followed by `raise Exception(f"Error: {res}")`. Because `log_exception` always raises, the second line was dead code, and the surfaced exception varied with whether the SDK response exposed a `.text` attribute. Every failure path now raises a single `EmbeddingError` that includes the underlying response detail, so the cause of a failed embedding is consistent and visible. **Fabricated token counts.** `LocalAIEmbed` returned a hardcoded `1024` and `OllamaEmbed` added `128` per text. These values feed `used_tokens` and therefore billing and usage tracking. Both now report the real count from the API (Ollama `prompt_eval_count`, LocalAI `usage`) and fall back to a local token count only when the server omits it. **Truncation overshoot.** The `8196` limit used by Mistral and Bedrock exceeded the standard `8192` ceiling and could push boundary sized inputs past the model limit. Limits are corrected to `8192` and made intentional per provider, and providers that rely on server side truncation now request it explicitly (Ollama `truncate=True`, Cohere `truncate="END"`). **Missing batching on Zhipu and Ollama.** Both issued one request per text. They now batch like the other OpenAI compatible providers, turning N round trips into `ceil(N / batch_size)`. Batched results are realigned by response `index` so a chunk always keeps its own vector. A shared `Base._batched_encode` helper owns the batch loop, optional truncation, result accumulation, and the single error path. It is the mechanism that lets these fixes live in one place instead of across 27 duplicated sites. The public `encode()` and `encode_queries()` contract stays the same, so existing callers are unaffected. Tests covering all four fixes are added under `test/unit_test/rag/llm/test_embedding_model.py`. ### Type of change - [x] Bug Fix (non-breaking change which fixes an issue) |
||
|
|
0d3e410826 |
fix: strip Ollama-style tag suffix from LocalAI model names (#15908)
## Summary
LocalAI exposes two API surfaces with conflicting naming conventions:
- `GET /api/tags` returns model names with `:latest` suffix (Ollama
format)
- `POST /v1/chat/completions` expects names without `:latest` (OpenAI
format)
RAGFlow discovered models via `/api/tags` and stored the tagged name,
then used it with `/v1/chat/completions`, causing a 404 error because
LocalAI didn't recognize `model:latest`.
## Fix
In `LocalAI.get_model_list()`, strip the tag suffix from model names
using `model["name"].rsplit(":", 1)[0]`, so stored names match what the
OpenAI-compatible endpoints expect.
|
||
|
|
38f9ea5fec |
fix(rerank): normalize reranker scores onto a single scale before hybrid blend (#15429)
### What problem does this PR solve? Closes #15428 The hybrid score in `rag/nlp/search.py` (`rerank_by_model`) blends reranker similarity with token similarity on a fixed `[0, 1]` scale: ```python return tkweight * np.array(tksim) + vtweight * vtsim + rank_fea # tkweight=0.3, vtweight=0.7 ``` The reranker implementations did not agree on that scale. Only three of roughly 17 providers normalized their output, and `NvidiaRerank` returned raw, unbounded logits. Weighted at `0.7`, a negative logit could push a genuinely relevant chunk below pure keyword matches, and its magnitude swamped `tksim`, which lives in `[0, 1]`. The practical effect was that the same query produced differently scaled scores depending on the configured reranker, and logit based providers degraded retrieval quality instead of improving it. This PR enforces a single scoring contract in one place: - `Base.similarity` is now the only public entry point. It short-circuits empty input and guarantees a normalized result. Each provider implements its raw scoring in `_compute_rank`, which removes sixteen duplicated empty input guards and the three scattered normalization calls. - Normalization is range aware. Providers that already return calibrated `[0, 1]` relevance scores (Cohere, Jina, Voyage, and others) keep their absolute magnitudes, so `similarity_threshold` filtering and the reported `vector_similarity` stay meaningful. Only out-of-range output such as NVIDIA logits is min-max rescaled into `[0, 1]`. - The twelve leftover `[DEBUG ...]` prints in `rerank_by_model`, introduced in #14231, are removed. They ran on every retrieval, added per chunk overhead, and leaked queries, keywords, and document content to stdout and logs. A new regression suite in `test/unit_test/rag/llm/test_rerank_normalization.py` covers logit rescaling (positive, negative, and flat batches), preservation of already calibrated scores, ordering, empty input handling, and the per provider HTTP path. It also asserts that no provider overrides `similarity()`, so the contract cannot silently drift. ### Type of change - [x] Bug Fix (non-breaking change which fixes an issue) |
||
|
|
bebf6ed244 |
fix(llm): strip non-generation keys from gen_conf for LiteLLM providers (#15427) (#15432)
### What problem does this PR solve? Fixes #15427. All LiteLLM-routed chats fail with: - Anthropic: `litellm.BadRequestError: AnthropicException - {"type":"invalid_request_error","message":"model_type: Extra inputs are not permitted"}` - OpenAI: `litellm.BadRequestError: OpenAIException - Unknown parameter: 'model_type'` This is a regression from v0.25.4. #### Root cause A chat assistant's `llm_setting` is forwarded to the model as `gen_conf`. `llm_setting` can legitimately carry RAGFlow-internal metadata such as `model_type` (the chat REST APIs in `api/apps/restful_apis/` read it back out of `llm_setting`), so that key ends up inside `gen_conf`. `Base._clean_conf` (OpenAI-compatible providers) already **whitelists** the keys it forwards, so direct-OpenAI providers were unaffected. `LiteLLMBase._clean_conf` only dropped `max_tokens` and passed everything else straight through to `litellm.acompletion`, which forwarded `model_type` to the upstream provider — and Anthropic / OpenAI reject it. Because both Claude and GPT route through LiteLLM, every chat broke. #### Fix - Extract the allowed-key set into a shared `ALLOWED_GEN_CONF_KEYS` constant and reuse it in `Base._clean_conf`. - Apply the same whitelist in `LiteLLMBase._clean_conf`, plus the LiteLLM-specific reasoning params (`thinking`, `reasoning_effort`, `extra_body`) that the model-family policies inject for reasoning models. This covers all four LiteLLM completion paths (`async_chat`, `async_chat_streamly`, `async_chat_with_tools`, `async_chat_streamly_with_tools`), since they all route through `_clean_conf`. #### Tests Adds `test/unit_test/rag/llm/test_clean_conf_whitelist.py` covering both backends: `model_type` (and other stray keys) are dropped, genuine generation params and `thinking` survive, `max_tokens` is removed, and the whitelist invariants hold. ### Type of change - [x] Bug Fix (non-breaking change which fixes an issue) - [x] Added test cases |
||
|
|
1a6df01b53 |
Bug fix: Enhance embeding model to give better error message (#15346)
To resolve https://github.com/infiniflow/ragflow/issues/15343 enhance the model embedding message to give extact failure message to customer. # QWen ## Retrieval <img width="3321" height="1033" alt="image" src="https://github.com/user-attachments/assets/6b82921a-a3a7-4a33-a383-1cf316398ee2" /> ## Chat <img width="2241" height="311" alt="image" src="https://github.com/user-attachments/assets/ec311365-62d5-407a-8915-5c8d72be9716" /> # SiliconFlow ## Retrieval <img width="3321" height="1033" alt="image" src="https://github.com/user-attachments/assets/ee2cd191-a27d-4729-b53d-2fbdb4e352cd" /> ## Chat <img width="1562" height="210" alt="image" src="https://github.com/user-attachments/assets/10376a8e-a3f4-422f-bc2e-96f2a8a96448" /> # Baichuan ## Retrieval <img width="3321" height="1107" alt="image" src="https://github.com/user-attachments/assets/dcb5409d-f7fc-4804-b186-5e1ee11e09c4" /> ## Chat <img width="2241" height="311" alt="image" src="https://github.com/user-attachments/assets/ec311365-62d5-407a-8915-5c8d72be9716" /> # Zhipu zhipu is good. |
||
|
|
e7544562cc |
Feat: @tool decorator for chat-model tool registration (#15047)
## Summary
- Adds a lightweight `@tool` decorator and `FunctionToolSession` adapter
in `rag/llm/tool_decorator.py` that let callers register plain Python
functions as LLM tools without hand-writing OpenAI function schemas or
building an MCP-style session.
- Refactors `Base.bind_tools` and `LiteLLMBase.bind_tools` in
`rag/llm/chat_model.py` to accept either the new decorator form
`bind_tools(tools=[fn1, fn2])` or the existing `(toolcall_session,
tools_schemas)` form, so existing agent/dialog call-sites in
`agent/component/agent_with_tools.py`, `api/db/services/llm_service.py`,
and `api/db/services/dialog_service.py` are unaffected.
- Adds 8 unit tests in `test/unit_test/rag/llm/test_tool_decorator.py`
covering schema shape, required/optional inference, sync + async
dispatch, and bad-input rejection.
## Usage
```python
from rag.llm.tool_decorator import tool
@tool
def get_weather(city: str) -> str:
"""Get current weather for a city.
:param city: City name to look up.
"""
return f"{city}: 21 C, partly cloudy"
chat_mdl.bind_tools(tools=[get_weather])
ans, tk = await chat_mdl.async_chat_with_tools(system, history)
```
The decorator introspects `inspect.signature` + type hints + the
docstring (`:param name:` style) and attaches an OpenAI-format
`openai_schema` to the callable. `FunctionToolSession` duck-types the
existing `ToolCallSession` protocol, dispatching async callables
directly and sync ones through `thread_pool_exec` so the event loop is
never blocked.
## Design notes
- `tool_decorator.py` deliberately does **not** live inside
`rag/llm/__init__.py` to avoid forcing every consumer through the heavy
provider auto-discovery loop and to sidestep a circular import
(`__init__.py` imports `chat_model`, which would otherwise need symbols
from `__init__.py`).
- `FunctionToolSession` is duck-typed against
`common.mcp_tool_call_conn.ToolCallSession` rather than explicitly
inheriting from it, so importing the decorator doesn't pull the MCP
client SDK into the import graph.
- Docstring parsing is intentionally minimal (`:param name:` only) to
keep this dependency-free; Google/NumPy styles can be added later via
`docstring_parser` if needed.
## Test plan
- [x] `python -m pytest test/unit_test/rag/llm/test_tool_decorator.py
-v` — 8 passed
- [x] `python -m pytest test/unit_test/rag/llm/
--ignore=test/unit_test/rag/llm/test_perplexity_embed.py` — 11 passed
(the ignored test has a pre-existing `numpy` import that's unrelated)
- [ ] Reviewer: smoke-test the new path end-to-end with a live model via
`chat_mdl.bind_tools(tools=[my_fn])` to confirm the OpenAI-format
schemas pass through unchanged
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1046042e01 |
fix(llm): replace mutable default gen_conf={} with None + defensive copy (#14566)
### What
19 methods across `rag/llm/chat_model.py` and `rag/llm/cv_model.py`
declare `gen_conf={}` (or `gen_conf: dict = {}`) as a parameter default
and then mutate `gen_conf` in place — typically `del
gen_conf["max_tokens"]`, `gen_conf["penalty_score"] = ...`, or
`gen_conf.pop(...)` as part of provider-specific normalization.
### The two bugs in this pattern
**1. Mutable default argument (Python footgun).** Python evaluates
default values **once** at function-definition time, so the single `{}`
dict is *shared* across every caller that doesn't pass `gen_conf`. The
first such call's mutations leak into the default seen by every
subsequent call.
```python
# Before
def chat_streamly(self, system, history, gen_conf={}, **kwargs):
if "max_tokens" in gen_conf:
del gen_conf["max_tokens"] # mutates the SHARED default dict
...
```
After call N with `max_tokens` set, call N+1 that omits `gen_conf` no
longer sees `max_tokens` — even though the caller never touched it.
**2. Caller-dict pollution.** When the caller *does* pass a `gen_conf`
dict, the same in-place mutations modify the caller's dict. A reused
`gen_conf` (very common for chat-loop callers that build the config once
and pass it on every turn) silently loses `max_tokens`,
`presence_penalty`, etc. after the first round.
### The fix
In every affected method:
- Change `gen_conf={}` (or `gen_conf: dict = {}`) → `gen_conf=None`.
- Add `gen_conf = dict(gen_conf or {})` as the first statement of the
body so all subsequent mutations operate on a fresh local copy.
```python
# After
def chat_streamly(self, system, history, gen_conf=None, **kwargs):
gen_conf = dict(gen_conf or {})
if "max_tokens" in gen_conf:
del gen_conf["max_tokens"] # local copy — safe
...
```
This is byte-for-byte identical provider-side behavior for callers that
already pass a fresh `gen_conf` per call. The new `dict(...)` copy is
O(small constant) per call.
### Files changed
- `rag/llm/chat_model.py` — 17 methods
- `rag/llm/cv_model.py` — 2 methods
### Tests
Adds `test/unit_test/rag/llm/test_gen_conf_no_mutable_default.py` — an
`ast`-based regression guard that walks both modules and asserts no
parameter named `gen_conf` ever has a mutable literal (`{}` or `[]`) as
its default. The test caught **five additional `gen_conf: dict = {}`
sites** that an initial `gen_conf={}` text grep had missed (annotated
parameters with whitespace), and would fail again if the pattern is ever
reintroduced.
```
$ pytest test/unit_test/rag/llm/test_gen_conf_no_mutable_default.py -v
============================== 3 passed in 0.04s ===============================
```
`ruff check` passes on all touched files.
### Notes
- This PR is intentionally focused on **just** the `gen_conf` default +
copy fix. There's a related (but separate) `history.insert(0, ...)`
pattern in the same files that mutates the caller's history list in 12
places — left for a follow-up so this PR stays mechanical and easy to
review.
### Latest revision (`700bb54a7`) — addresses CodeRabbit review
- Type annotation: `gen_conf: dict = None` → `gen_conf: dict | None =
None` (5 occurrences in `chat_model.py`). The old annotation was a
static-checker mismatch since `None` isn't a `dict`.
- Regression test: the AST check accessed `default.keys` directly.
`ast.List` has no `.keys` attribute — a future `gen_conf=[]` would crash
with `AttributeError` instead of being caught. Use `getattr` for both
`.keys` (Dict) and `.elts` (List). Manually verified the updated check
correctly catches both `gen_conf={}` and `gen_conf=[]` while ignoring
`gen_conf=None` and non-empty literals.
---------
Co-authored-by: Ricardo <ricardo@example.com>
|
||
|
|
13d0df1562 |
feat: add Perplexity contextualized embeddings API as a new model provider (#13709)
### What problem does this PR solve? Adds Perplexity contextualized embeddings API as a new model provider, as requested in #13610. - `PerplexityEmbed` provider in `rag/llm/embedding_model.py` supporting both standard (`/v1/embeddings`) and contextualized (`/v1/contextualizedembeddings`) endpoints - All 4 Perplexity embedding models registered in `conf/llm_factories.json`: `pplx-embed-v1-0.6b`, `pplx-embed-v1-4b`, `pplx-embed-context-v1-0.6b`, `pplx-embed-context-v1-4b` - Frontend entries (enum, icon mapping, API key URL) in `web/src/constants/llm.ts` - Updated `docs/guides/models/supported_models.mdx` - 22 unit tests in `test/unit_test/rag/llm/test_perplexity_embed.py` Perplexity's API returns `base64_int8` encoded embeddings (not OpenAI-compatible), so this uses a custom `requests`-based implementation. Contextualized vs standard model is auto-detected from the model name. Closes #13610 ### Type of change - [x] New Feature (non-breaking change which adds functionality) - [x] Documentation Update |