Files
ragflow/conf/models/bedrock.json

172 lines
3.4 KiB
JSON
Raw Permalink Normal View History

{
"name": "Bedrock",
"url_suffix": {
"chat": "converse",
Go: implement Bedrock embeddings (#15543) ### What problem does this PR solve? Fixes #15542. AWS Bedrock support for the Go model provider layer was added in #15166, but embedding support was intentionally left out of scope and `BedrockModel.Embed(...)` still returned the `no such method` sentinel. This PR implements Bedrock text embeddings under the umbrella provider tracker #14736. ### What this PR includes - `internal/entity/models/bedrock.go`: implement `BedrockModel.Embed(...)` through Bedrock Runtime `InvokeModel` with existing SigV4 auth, region resolution, and runtime URL helpers. - Titan embeddings: supports `amazon.titan-embed-text-v1` and `amazon.titan-embed-text-v2:0`; v2 forwards `EmbeddingConfig.Dimension` as `dimensions` when provided, while v1 keeps the payload minimal. - Cohere embeddings: supports `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, and `cohere.embed-v4:0`; batches input texts and maps returned vectors to RAGFlow `EmbeddingData` in input order. - `conf/models/bedrock.json`: adds the `embedding` URL suffix (`invoke`) and Bedrock embedding model entries. - `internal/entity/models/bedrock_test.go`: adds unit tests for Titan, Cohere, typed Cohere responses, validation, empty input, unsupported models, and HTTP error propagation. Reference docs: - Bedrock InvokeModel API: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModel.html - Titan Text Embeddings: https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html - Cohere Embed models on Bedrock: https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed.html ### Type of change - [x] New Feature (non-breaking change which adds functionality) ### How was this tested? - [x] `jq empty conf/models/bedrock.json` - [x] `git diff --check` - [x] `go test ./internal/entity/models/... -run Bedrock -count=1` - [x] `go test ./internal/entity/models/... -run '^$' -count=1` - [x] `go test ./internal/entity/models/... -run Bedrock -race -count=1` Note: `go test ./internal/entity/models/... -count=1` currently fails in unrelated existing Astraflow coverage (`TestAstraflowEmbedReturnsNoSuchMethod` panics in `internal/entity/models/astraflow.go`). The Bedrock-specific tests and compile-only package check pass.
2026-06-04 19:26:32 -10:00
"models": "foundation-models",
"embedding": "invoke"
},
"class": "bedrock",
"models": [
{
"name": "anthropic.claude-3-5-sonnet-20241022-v2:0",
"content_length": 200000,
"max_output": 8192,
"model_types": [
"chat"
]
},
{
"name": "anthropic.claude-3-5-haiku-20241022-v1:0",
"content_length": 200000,
"max_output": 8192,
"model_types": [
"chat"
]
},
{
"name": "anthropic.claude-3-opus-20240229-v1:0",
"content_length": 200000,
"max_output": 4096,
"model_types": [
"chat"
]
},
{
"name": "anthropic.claude-3-sonnet-20240229-v1:0",
"content_length": 200000,
"max_output": 4096,
"model_types": [
"chat"
]
},
{
"name": "anthropic.claude-3-haiku-20240307-v1:0",
"content_length": 200000,
"max_output": 4096,
"model_types": [
"chat"
]
},
{
"name": "meta.llama3-1-405b-instruct-v1:0",
"content_length": 131072,
"max_output": 8192,
"model_types": [
"chat"
]
},
{
"name": "meta.llama3-1-70b-instruct-v1:0",
"content_length": 131072,
"max_output": 8192,
"model_types": [
"chat"
]
},
{
"name": "meta.llama3-1-8b-instruct-v1:0",
"content_length": 131072,
"max_output": 8192,
"model_types": [
"chat"
]
},
{
"name": "mistral.mistral-large-2407-v1:0",
"content_length": 128000,
"max_output": 4096,
"model_types": [
"chat"
]
},
{
"name": "mistral.mixtral-8x7b-instruct-v0:1",
"content_length": 32000,
"max_output": 4096,
"model_types": [
"chat"
]
},
{
"name": "amazon.nova-pro-v1:0",
"content_length": 300000,
"max_output": 5000,
"model_types": [
"chat"
]
},
{
"name": "amazon.nova-lite-v1:0",
"content_length": 300000,
"max_output": 5000,
"model_types": [
"chat"
]
},
{
"name": "amazon.nova-micro-v1:0",
"content_length": 128000,
"max_output": 5000,
"model_types": [
"chat"
]
},
{
"name": "cohere.command-r-plus-v1:0",
"content_length": 131072,
"max_output": 4096,
"model_types": [
"chat"
]
},
{
"name": "cohere.command-r-v1:0",
"content_length": 131072,
"max_output": 4096,
"model_types": [
"chat"
]
Go: implement Bedrock embeddings (#15543) ### What problem does this PR solve? Fixes #15542. AWS Bedrock support for the Go model provider layer was added in #15166, but embedding support was intentionally left out of scope and `BedrockModel.Embed(...)` still returned the `no such method` sentinel. This PR implements Bedrock text embeddings under the umbrella provider tracker #14736. ### What this PR includes - `internal/entity/models/bedrock.go`: implement `BedrockModel.Embed(...)` through Bedrock Runtime `InvokeModel` with existing SigV4 auth, region resolution, and runtime URL helpers. - Titan embeddings: supports `amazon.titan-embed-text-v1` and `amazon.titan-embed-text-v2:0`; v2 forwards `EmbeddingConfig.Dimension` as `dimensions` when provided, while v1 keeps the payload minimal. - Cohere embeddings: supports `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, and `cohere.embed-v4:0`; batches input texts and maps returned vectors to RAGFlow `EmbeddingData` in input order. - `conf/models/bedrock.json`: adds the `embedding` URL suffix (`invoke`) and Bedrock embedding model entries. - `internal/entity/models/bedrock_test.go`: adds unit tests for Titan, Cohere, typed Cohere responses, validation, empty input, unsupported models, and HTTP error propagation. Reference docs: - Bedrock InvokeModel API: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModel.html - Titan Text Embeddings: https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html - Cohere Embed models on Bedrock: https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed.html ### Type of change - [x] New Feature (non-breaking change which adds functionality) ### How was this tested? - [x] `jq empty conf/models/bedrock.json` - [x] `git diff --check` - [x] `go test ./internal/entity/models/... -run Bedrock -count=1` - [x] `go test ./internal/entity/models/... -run '^$' -count=1` - [x] `go test ./internal/entity/models/... -run Bedrock -race -count=1` Note: `go test ./internal/entity/models/... -count=1` currently fails in unrelated existing Astraflow coverage (`TestAstraflowEmbedReturnsNoSuchMethod` panics in `internal/entity/models/astraflow.go`). The Bedrock-specific tests and compile-only package check pass.
2026-06-04 19:26:32 -10:00
},
{
"name": "amazon.titan-embed-text-v2:0",
"max_tokens": 8192,
"model_types": [
"embedding"
feat: add batch_size to all embedding model configs (#17877) ## Summary Add a `batch_size` field to every embedding model entry in `conf/models/*.json`. The field represents the maximum number of text inputs that can be submitted to the embedding API in a single request. **75 embedding models across 30 config files** now carry a `batch_size`. Values were verified against each provider's official documentation (see the verification table at `Desktop/embedding_models_verified.md`). ## Distribution | batch_size | # models | Provider / Model | |---|---|---| | 1 | 2 | AWS Bedrock `amazon.titan-embed-text-v1/v2:0` — Bedrock `invoke` accepts a single input per call | | 10 | 3 | Aliyun `text-embedding-v3/v4`, Volcengine `doubao-embedding-vision-251215` | | 16 | 5 | BaiChuan `Baichuan-Text-Embedding`, Baidu Qianfan `embedding-v1`, Mistral `mistral-embed`, Replicate (x2) | | 32 | 10 | NVIDIA NIM (x3), SILICONFLOW (x2), PPIO (x3), GiteeAI `bge-m3`, HuaweiCloud `bge-m3` | | 50 | 4 | Tencent Hunyuan `kinfra` embeddings (x4) — `InputList.N` max 50 | | 96 | 8 | Cohere embed-v3/v4 (x5), Bedrock Cohere (x3) | | 100 | 3 | Google Gemini `text-embedding-004`, Upstage (x2) | | 512 | 5 | Zhipu GLM `embedding-2/3` (x2), Perplexity `pplx-embed` (x2), Astraflow `text-embedding-3-large` | | 1000 | 9 | Voyage AI (x9) — API reference max | | 1024 | 1 | DeepInfra `Qwen/Qwen3-Embedding-4B` | | 2048 | 15 | OpenAI (x3) + OpenAI-API-compatible proxies (CometAPI, n1n, Jiekou.AI, GreenPT, TogetherAI, NovitaAI) — OpenAI contract limit | | 16384 | 10 | Jina (x8), 302.AI, GiteeAI `jina-clip-v2` — no documented Jina batch limit, safe high cap | --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-06 10:47:38 +08:00
],
"max_batch_size": 1
Go: implement Bedrock embeddings (#15543) ### What problem does this PR solve? Fixes #15542. AWS Bedrock support for the Go model provider layer was added in #15166, but embedding support was intentionally left out of scope and `BedrockModel.Embed(...)` still returned the `no such method` sentinel. This PR implements Bedrock text embeddings under the umbrella provider tracker #14736. ### What this PR includes - `internal/entity/models/bedrock.go`: implement `BedrockModel.Embed(...)` through Bedrock Runtime `InvokeModel` with existing SigV4 auth, region resolution, and runtime URL helpers. - Titan embeddings: supports `amazon.titan-embed-text-v1` and `amazon.titan-embed-text-v2:0`; v2 forwards `EmbeddingConfig.Dimension` as `dimensions` when provided, while v1 keeps the payload minimal. - Cohere embeddings: supports `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, and `cohere.embed-v4:0`; batches input texts and maps returned vectors to RAGFlow `EmbeddingData` in input order. - `conf/models/bedrock.json`: adds the `embedding` URL suffix (`invoke`) and Bedrock embedding model entries. - `internal/entity/models/bedrock_test.go`: adds unit tests for Titan, Cohere, typed Cohere responses, validation, empty input, unsupported models, and HTTP error propagation. Reference docs: - Bedrock InvokeModel API: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModel.html - Titan Text Embeddings: https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html - Cohere Embed models on Bedrock: https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed.html ### Type of change - [x] New Feature (non-breaking change which adds functionality) ### How was this tested? - [x] `jq empty conf/models/bedrock.json` - [x] `git diff --check` - [x] `go test ./internal/entity/models/... -run Bedrock -count=1` - [x] `go test ./internal/entity/models/... -run '^$' -count=1` - [x] `go test ./internal/entity/models/... -run Bedrock -race -count=1` Note: `go test ./internal/entity/models/... -count=1` currently fails in unrelated existing Astraflow coverage (`TestAstraflowEmbedReturnsNoSuchMethod` panics in `internal/entity/models/astraflow.go`). The Bedrock-specific tests and compile-only package check pass.
2026-06-04 19:26:32 -10:00
},
{
"name": "amazon.titan-embed-text-v1",
"max_tokens": 8192,
"model_types": [
"embedding"
feat: add batch_size to all embedding model configs (#17877) ## Summary Add a `batch_size` field to every embedding model entry in `conf/models/*.json`. The field represents the maximum number of text inputs that can be submitted to the embedding API in a single request. **75 embedding models across 30 config files** now carry a `batch_size`. Values were verified against each provider's official documentation (see the verification table at `Desktop/embedding_models_verified.md`). ## Distribution | batch_size | # models | Provider / Model | |---|---|---| | 1 | 2 | AWS Bedrock `amazon.titan-embed-text-v1/v2:0` — Bedrock `invoke` accepts a single input per call | | 10 | 3 | Aliyun `text-embedding-v3/v4`, Volcengine `doubao-embedding-vision-251215` | | 16 | 5 | BaiChuan `Baichuan-Text-Embedding`, Baidu Qianfan `embedding-v1`, Mistral `mistral-embed`, Replicate (x2) | | 32 | 10 | NVIDIA NIM (x3), SILICONFLOW (x2), PPIO (x3), GiteeAI `bge-m3`, HuaweiCloud `bge-m3` | | 50 | 4 | Tencent Hunyuan `kinfra` embeddings (x4) — `InputList.N` max 50 | | 96 | 8 | Cohere embed-v3/v4 (x5), Bedrock Cohere (x3) | | 100 | 3 | Google Gemini `text-embedding-004`, Upstage (x2) | | 512 | 5 | Zhipu GLM `embedding-2/3` (x2), Perplexity `pplx-embed` (x2), Astraflow `text-embedding-3-large` | | 1000 | 9 | Voyage AI (x9) — API reference max | | 1024 | 1 | DeepInfra `Qwen/Qwen3-Embedding-4B` | | 2048 | 15 | OpenAI (x3) + OpenAI-API-compatible proxies (CometAPI, n1n, Jiekou.AI, GreenPT, TogetherAI, NovitaAI) — OpenAI contract limit | | 16384 | 10 | Jina (x8), 302.AI, GiteeAI `jina-clip-v2` — no documented Jina batch limit, safe high cap | --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-06 10:47:38 +08:00
],
"max_batch_size": 1
Go: implement Bedrock embeddings (#15543) ### What problem does this PR solve? Fixes #15542. AWS Bedrock support for the Go model provider layer was added in #15166, but embedding support was intentionally left out of scope and `BedrockModel.Embed(...)` still returned the `no such method` sentinel. This PR implements Bedrock text embeddings under the umbrella provider tracker #14736. ### What this PR includes - `internal/entity/models/bedrock.go`: implement `BedrockModel.Embed(...)` through Bedrock Runtime `InvokeModel` with existing SigV4 auth, region resolution, and runtime URL helpers. - Titan embeddings: supports `amazon.titan-embed-text-v1` and `amazon.titan-embed-text-v2:0`; v2 forwards `EmbeddingConfig.Dimension` as `dimensions` when provided, while v1 keeps the payload minimal. - Cohere embeddings: supports `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, and `cohere.embed-v4:0`; batches input texts and maps returned vectors to RAGFlow `EmbeddingData` in input order. - `conf/models/bedrock.json`: adds the `embedding` URL suffix (`invoke`) and Bedrock embedding model entries. - `internal/entity/models/bedrock_test.go`: adds unit tests for Titan, Cohere, typed Cohere responses, validation, empty input, unsupported models, and HTTP error propagation. Reference docs: - Bedrock InvokeModel API: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModel.html - Titan Text Embeddings: https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html - Cohere Embed models on Bedrock: https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed.html ### Type of change - [x] New Feature (non-breaking change which adds functionality) ### How was this tested? - [x] `jq empty conf/models/bedrock.json` - [x] `git diff --check` - [x] `go test ./internal/entity/models/... -run Bedrock -count=1` - [x] `go test ./internal/entity/models/... -run '^$' -count=1` - [x] `go test ./internal/entity/models/... -run Bedrock -race -count=1` Note: `go test ./internal/entity/models/... -count=1` currently fails in unrelated existing Astraflow coverage (`TestAstraflowEmbedReturnsNoSuchMethod` panics in `internal/entity/models/astraflow.go`). The Bedrock-specific tests and compile-only package check pass.
2026-06-04 19:26:32 -10:00
},
{
"name": "cohere.embed-english-v3",
"max_tokens": 512,
"model_types": [
"embedding"
feat: add batch_size to all embedding model configs (#17877) ## Summary Add a `batch_size` field to every embedding model entry in `conf/models/*.json`. The field represents the maximum number of text inputs that can be submitted to the embedding API in a single request. **75 embedding models across 30 config files** now carry a `batch_size`. Values were verified against each provider's official documentation (see the verification table at `Desktop/embedding_models_verified.md`). ## Distribution | batch_size | # models | Provider / Model | |---|---|---| | 1 | 2 | AWS Bedrock `amazon.titan-embed-text-v1/v2:0` — Bedrock `invoke` accepts a single input per call | | 10 | 3 | Aliyun `text-embedding-v3/v4`, Volcengine `doubao-embedding-vision-251215` | | 16 | 5 | BaiChuan `Baichuan-Text-Embedding`, Baidu Qianfan `embedding-v1`, Mistral `mistral-embed`, Replicate (x2) | | 32 | 10 | NVIDIA NIM (x3), SILICONFLOW (x2), PPIO (x3), GiteeAI `bge-m3`, HuaweiCloud `bge-m3` | | 50 | 4 | Tencent Hunyuan `kinfra` embeddings (x4) — `InputList.N` max 50 | | 96 | 8 | Cohere embed-v3/v4 (x5), Bedrock Cohere (x3) | | 100 | 3 | Google Gemini `text-embedding-004`, Upstage (x2) | | 512 | 5 | Zhipu GLM `embedding-2/3` (x2), Perplexity `pplx-embed` (x2), Astraflow `text-embedding-3-large` | | 1000 | 9 | Voyage AI (x9) — API reference max | | 1024 | 1 | DeepInfra `Qwen/Qwen3-Embedding-4B` | | 2048 | 15 | OpenAI (x3) + OpenAI-API-compatible proxies (CometAPI, n1n, Jiekou.AI, GreenPT, TogetherAI, NovitaAI) — OpenAI contract limit | | 16384 | 10 | Jina (x8), 302.AI, GiteeAI `jina-clip-v2` — no documented Jina batch limit, safe high cap | --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-06 10:47:38 +08:00
],
"max_batch_size": 96
Go: implement Bedrock embeddings (#15543) ### What problem does this PR solve? Fixes #15542. AWS Bedrock support for the Go model provider layer was added in #15166, but embedding support was intentionally left out of scope and `BedrockModel.Embed(...)` still returned the `no such method` sentinel. This PR implements Bedrock text embeddings under the umbrella provider tracker #14736. ### What this PR includes - `internal/entity/models/bedrock.go`: implement `BedrockModel.Embed(...)` through Bedrock Runtime `InvokeModel` with existing SigV4 auth, region resolution, and runtime URL helpers. - Titan embeddings: supports `amazon.titan-embed-text-v1` and `amazon.titan-embed-text-v2:0`; v2 forwards `EmbeddingConfig.Dimension` as `dimensions` when provided, while v1 keeps the payload minimal. - Cohere embeddings: supports `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, and `cohere.embed-v4:0`; batches input texts and maps returned vectors to RAGFlow `EmbeddingData` in input order. - `conf/models/bedrock.json`: adds the `embedding` URL suffix (`invoke`) and Bedrock embedding model entries. - `internal/entity/models/bedrock_test.go`: adds unit tests for Titan, Cohere, typed Cohere responses, validation, empty input, unsupported models, and HTTP error propagation. Reference docs: - Bedrock InvokeModel API: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModel.html - Titan Text Embeddings: https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html - Cohere Embed models on Bedrock: https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed.html ### Type of change - [x] New Feature (non-breaking change which adds functionality) ### How was this tested? - [x] `jq empty conf/models/bedrock.json` - [x] `git diff --check` - [x] `go test ./internal/entity/models/... -run Bedrock -count=1` - [x] `go test ./internal/entity/models/... -run '^$' -count=1` - [x] `go test ./internal/entity/models/... -run Bedrock -race -count=1` Note: `go test ./internal/entity/models/... -count=1` currently fails in unrelated existing Astraflow coverage (`TestAstraflowEmbedReturnsNoSuchMethod` panics in `internal/entity/models/astraflow.go`). The Bedrock-specific tests and compile-only package check pass.
2026-06-04 19:26:32 -10:00
},
{
"name": "cohere.embed-multilingual-v3",
"max_tokens": 512,
"model_types": [
"embedding"
feat: add batch_size to all embedding model configs (#17877) ## Summary Add a `batch_size` field to every embedding model entry in `conf/models/*.json`. The field represents the maximum number of text inputs that can be submitted to the embedding API in a single request. **75 embedding models across 30 config files** now carry a `batch_size`. Values were verified against each provider's official documentation (see the verification table at `Desktop/embedding_models_verified.md`). ## Distribution | batch_size | # models | Provider / Model | |---|---|---| | 1 | 2 | AWS Bedrock `amazon.titan-embed-text-v1/v2:0` — Bedrock `invoke` accepts a single input per call | | 10 | 3 | Aliyun `text-embedding-v3/v4`, Volcengine `doubao-embedding-vision-251215` | | 16 | 5 | BaiChuan `Baichuan-Text-Embedding`, Baidu Qianfan `embedding-v1`, Mistral `mistral-embed`, Replicate (x2) | | 32 | 10 | NVIDIA NIM (x3), SILICONFLOW (x2), PPIO (x3), GiteeAI `bge-m3`, HuaweiCloud `bge-m3` | | 50 | 4 | Tencent Hunyuan `kinfra` embeddings (x4) — `InputList.N` max 50 | | 96 | 8 | Cohere embed-v3/v4 (x5), Bedrock Cohere (x3) | | 100 | 3 | Google Gemini `text-embedding-004`, Upstage (x2) | | 512 | 5 | Zhipu GLM `embedding-2/3` (x2), Perplexity `pplx-embed` (x2), Astraflow `text-embedding-3-large` | | 1000 | 9 | Voyage AI (x9) — API reference max | | 1024 | 1 | DeepInfra `Qwen/Qwen3-Embedding-4B` | | 2048 | 15 | OpenAI (x3) + OpenAI-API-compatible proxies (CometAPI, n1n, Jiekou.AI, GreenPT, TogetherAI, NovitaAI) — OpenAI contract limit | | 16384 | 10 | Jina (x8), 302.AI, GiteeAI `jina-clip-v2` — no documented Jina batch limit, safe high cap | --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-06 10:47:38 +08:00
],
"max_batch_size": 96
Go: implement Bedrock embeddings (#15543) ### What problem does this PR solve? Fixes #15542. AWS Bedrock support for the Go model provider layer was added in #15166, but embedding support was intentionally left out of scope and `BedrockModel.Embed(...)` still returned the `no such method` sentinel. This PR implements Bedrock text embeddings under the umbrella provider tracker #14736. ### What this PR includes - `internal/entity/models/bedrock.go`: implement `BedrockModel.Embed(...)` through Bedrock Runtime `InvokeModel` with existing SigV4 auth, region resolution, and runtime URL helpers. - Titan embeddings: supports `amazon.titan-embed-text-v1` and `amazon.titan-embed-text-v2:0`; v2 forwards `EmbeddingConfig.Dimension` as `dimensions` when provided, while v1 keeps the payload minimal. - Cohere embeddings: supports `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, and `cohere.embed-v4:0`; batches input texts and maps returned vectors to RAGFlow `EmbeddingData` in input order. - `conf/models/bedrock.json`: adds the `embedding` URL suffix (`invoke`) and Bedrock embedding model entries. - `internal/entity/models/bedrock_test.go`: adds unit tests for Titan, Cohere, typed Cohere responses, validation, empty input, unsupported models, and HTTP error propagation. Reference docs: - Bedrock InvokeModel API: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModel.html - Titan Text Embeddings: https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html - Cohere Embed models on Bedrock: https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed.html ### Type of change - [x] New Feature (non-breaking change which adds functionality) ### How was this tested? - [x] `jq empty conf/models/bedrock.json` - [x] `git diff --check` - [x] `go test ./internal/entity/models/... -run Bedrock -count=1` - [x] `go test ./internal/entity/models/... -run '^$' -count=1` - [x] `go test ./internal/entity/models/... -run Bedrock -race -count=1` Note: `go test ./internal/entity/models/... -count=1` currently fails in unrelated existing Astraflow coverage (`TestAstraflowEmbedReturnsNoSuchMethod` panics in `internal/entity/models/astraflow.go`). The Bedrock-specific tests and compile-only package check pass.
2026-06-04 19:26:32 -10:00
},
{
"name": "cohere.embed-v4:0",
"max_tokens": 128000,
"model_types": [
"embedding"
feat: add batch_size to all embedding model configs (#17877) ## Summary Add a `batch_size` field to every embedding model entry in `conf/models/*.json`. The field represents the maximum number of text inputs that can be submitted to the embedding API in a single request. **75 embedding models across 30 config files** now carry a `batch_size`. Values were verified against each provider's official documentation (see the verification table at `Desktop/embedding_models_verified.md`). ## Distribution | batch_size | # models | Provider / Model | |---|---|---| | 1 | 2 | AWS Bedrock `amazon.titan-embed-text-v1/v2:0` — Bedrock `invoke` accepts a single input per call | | 10 | 3 | Aliyun `text-embedding-v3/v4`, Volcengine `doubao-embedding-vision-251215` | | 16 | 5 | BaiChuan `Baichuan-Text-Embedding`, Baidu Qianfan `embedding-v1`, Mistral `mistral-embed`, Replicate (x2) | | 32 | 10 | NVIDIA NIM (x3), SILICONFLOW (x2), PPIO (x3), GiteeAI `bge-m3`, HuaweiCloud `bge-m3` | | 50 | 4 | Tencent Hunyuan `kinfra` embeddings (x4) — `InputList.N` max 50 | | 96 | 8 | Cohere embed-v3/v4 (x5), Bedrock Cohere (x3) | | 100 | 3 | Google Gemini `text-embedding-004`, Upstage (x2) | | 512 | 5 | Zhipu GLM `embedding-2/3` (x2), Perplexity `pplx-embed` (x2), Astraflow `text-embedding-3-large` | | 1000 | 9 | Voyage AI (x9) — API reference max | | 1024 | 1 | DeepInfra `Qwen/Qwen3-Embedding-4B` | | 2048 | 15 | OpenAI (x3) + OpenAI-API-compatible proxies (CometAPI, n1n, Jiekou.AI, GreenPT, TogetherAI, NovitaAI) — OpenAI contract limit | | 16384 | 10 | Jina (x8), 302.AI, GiteeAI `jina-clip-v2` — no documented Jina batch limit, safe high cap | --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-06 10:47:38 +08:00
],
"max_batch_size": 96
}
]
}