2026-06-09 22:48:50 +08:00
|
|
|
//
|
|
|
|
|
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
|
|
|
|
|
//
|
|
|
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
|
|
|
// you may not use this file except in compliance with the License.
|
|
|
|
|
// You may obtain a copy of the License at
|
|
|
|
|
//
|
|
|
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
|
|
|
//
|
|
|
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
|
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
|
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
|
|
|
// See the License for the specific language governing permissions and
|
|
|
|
|
// limitations under the License.
|
|
|
|
|
//
|
|
|
|
|
|
|
|
|
|
package common
|
|
|
|
|
|
2026-06-15 10:10:14 +08:00
|
|
|
import (
|
2026-06-20 02:31:07 +08:00
|
|
|
"encoding/base64"
|
2026-06-15 10:10:14 +08:00
|
|
|
"fmt"
|
|
|
|
|
"regexp"
|
2026-07-07 11:12:38 +08:00
|
|
|
"slices"
|
2026-07-03 19:37:53 +08:00
|
|
|
"strconv"
|
2026-06-15 10:10:14 +08:00
|
|
|
"strings"
|
2026-07-03 19:37:53 +08:00
|
|
|
"time"
|
Feat(ingestion): align image to MinIO upload, unify chunk-id computation and add PPT parsing support (#17111)
## Summary
Align the Go ingestion pipeline with Python's `image` → `img_id`
persistence semantics, and unify the chunk-id computation across all
paths.
### Changes
**1. Image upload at chunker stage **
- Add `ImageUploader` type and `DefaultImageUploader` in
`internal/ingestion/component/image_uploader.go` — the write-side
counterpart to `FetchBinary`, storing raw image bytes at `(bucket=kbID,
key=chunkID)`, no re-encoding.
- Add `uploadOneImage` — pure upload primitive (bytes in, `img_id` out),
does not touch chunk maps.
- Add `uploadChunkImages` / `uploadChunkImage` — caller-side helper:
decodes `image` from a chunk, uploads bytes, writes `ck["img_id"]`,
`delete(ck,"image")` , bounded by a process-wide semaphore (default 10,
env `MAX_CONCURRENT_MINIO`).
- Wire via `imageUploadDecorator` in `register.go`: every chunker runs
the upload pass at invocation time, writing `ck["id"]` before upload and
dropping image bytes right after — peak memory = single chunk image
lifetime.
**2. Unify chunk-id computation**
- Consolidate three separate id-computation paths (`component.ChunkID`,
`task.ChunkID`, inline `FormatUint` in API) into one:
`common.ChunkID(docID, text string)`, using `%016x` +
`xxhash.Sum64String(text+docID)` (matching Python `hexdigest()`).
- The chunker decorator writes `ck["id"]` via `common.ChunkID`; the
persist stage (`ProcessChunksForPipeline`) falls back to the same
function (`if !exists id`).
- The API AddChunk path now also calls `common.ChunkID` instead of the
divergent `FormatUint(xxhash.Sum64(...))` — fixing a pre-existing
inconsistency.
- Delete `internal/ingestion/component/chunk_id.go` and
`internal/ingestion/task/chunk_builder.go` (both were pure forwarding
shells).
**3. Preserve `img_id` (never deleted)**
- `img_id` is a persistent index field (Infinity, OB) and the only
consumer-side reference for image retrieval; it is NEVER removed from
the chunk map. Only `image` (raw data URL) is dropped after upload.
**4. PPT parser support**
Previously PPT parsing failed. Add support to parse.
### Key design decisions
| Decision | Choice |
|----------|--------|
| Upload timing | Chunker stage (not persist), so image bytes are
dropped immediately — bounds peak memory to one chunk image |
| Upload concurrency | Process-wide semaphore, default 10 (matches
Python `minio_limiter`), env `MAX_CONCURRENT_MINIO` |
| Image encoding | Store as-is, no JPEG re-encoding (unlike Python) |
| `img_id` format | `"<kb_id>-<chunk_id>"` — matches Python
task_executor path |
| id function | Single `common.ChunkID(docID, text)`, concatenation
`text+docID` inside hash (matching Python) |
| `removeInternalChunkFields` | Retains `delete(ck,"image")` as
defensive fallback for non-chunker paths |
### Files touched
| File | Change |
|------|--------|
| `internal/common/format.go` | Add `ChunkID(docID, text)` |
| `internal/common/format_test.go` | Add ChunkID golden-value test |
| `internal/ingestion/component/image_uploader.go` | Add `ImageUploader`
type + `DefaultImageUploader` |
| `internal/ingestion/component/chunker/image_upload.go` | Add
`uploadOneImage`, `uploadChunkImages`, `uploadChunkImage`,
`decodeChunkImage`, semaphore |
| `internal/ingestion/component/chunker/image_upload_test.go` | Tests:
upload/drop, skip, no-image, concurrency, missing-id error |
| `internal/ingestion/component/chunker/register.go` | Add
`imageUploadDecorator` (writes `ck["id"]`, runs upload) |
| `internal/ingestion/task/chunk_process.go` | Use `common.ChunkID` for
persist fallback |
| `internal/service/chunk/chunk.go` | Use `common.ChunkID` instead of
`FormatUint` |
| `internal/ingestion/component/chunk_id.go` | **Deleted** (moved to
`common`) |
| `internal/ingestion/task/chunk_builder.go` | **Deleted** (shell, no
callers left) |
| `internal/ingestion/task/chunk_builder_test.go` | **Deleted** (test
migrated to `common/format_test.go`) |
### Verification
```
bash build.sh --test ./internal/service/chunk/... ./internal/common/... ./internal/ingestion/component/... ./internal/ingestion/task/...
→ ok service/chunk / common / component / chunker / schema / task
```
2026-07-20 19:33:51 +08:00
|
|
|
|
|
|
|
|
"github.com/cespare/xxhash/v2"
|
2026-06-15 10:10:14 +08:00
|
|
|
)
|
2026-06-09 22:48:50 +08:00
|
|
|
|
|
|
|
|
// PtrString formats a pointer value as a string for debug/log output.
|
|
|
|
|
// Returns "<nil>" for nil pointers.
|
|
|
|
|
func PtrString[T any](p *T) string {
|
|
|
|
|
if p == nil {
|
|
|
|
|
return "<nil>"
|
|
|
|
|
}
|
|
|
|
|
return fmt.Sprintf("%v", *p)
|
|
|
|
|
}
|
2026-06-15 10:10:14 +08:00
|
|
|
|
|
|
|
|
// composite model name format: model_name@instance_name@provider_name
|
|
|
|
|
func IsCompositeModelName(modelName string) bool {
|
|
|
|
|
parts := strings.Split(modelName, "@")
|
|
|
|
|
if len(parts) != 3 {
|
|
|
|
|
return false
|
|
|
|
|
}
|
2026-07-07 11:12:38 +08:00
|
|
|
return !slices.Contains(parts, "")
|
2026-06-15 10:10:14 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
func IsUUID(uuid string) bool {
|
|
|
|
|
// only lower case letters and numbers, length is 32
|
|
|
|
|
if len(uuid) != 32 {
|
|
|
|
|
return false
|
|
|
|
|
}
|
|
|
|
|
uuidRegex := regexp.MustCompile(`^[a-z0-9]+$`)
|
|
|
|
|
if uuidRegex.MatchString(uuid) {
|
|
|
|
|
return true
|
|
|
|
|
}
|
|
|
|
|
return false
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// ExtractCompositeName splits a composite model name into three parts.
|
|
|
|
|
// Returns (modelName, instanceName, providerName, true) on success,
|
|
|
|
|
// or ("", "", "", false) if the name is not a valid composite name.
|
|
|
|
|
func ExtractCompositeName(modelName string) (string, string, string, error) {
|
|
|
|
|
parts := strings.Split(modelName, "@")
|
|
|
|
|
if len(parts) != 3 {
|
|
|
|
|
return "", "", "", fmt.Errorf("invalid model name format")
|
|
|
|
|
}
|
2026-07-07 11:12:38 +08:00
|
|
|
if slices.Contains(parts, "") {
|
|
|
|
|
return "", "", "", fmt.Errorf("invalid model name format")
|
2026-06-15 10:10:14 +08:00
|
|
|
}
|
|
|
|
|
return parts[0], parts[1], parts[2], nil
|
|
|
|
|
}
|
2026-06-20 02:31:07 +08:00
|
|
|
|
2026-06-23 19:29:06 +08:00
|
|
|
func EncodeToBase64(email string) string {
|
2026-06-20 02:31:07 +08:00
|
|
|
return base64.StdEncoding.EncodeToString([]byte(email))
|
|
|
|
|
}
|
|
|
|
|
|
2026-06-23 19:29:06 +08:00
|
|
|
func DecodeFromBase64(encoded string) (string, error) {
|
2026-06-20 02:31:07 +08:00
|
|
|
decoded, err := base64.StdEncoding.DecodeString(encoded)
|
|
|
|
|
if err != nil {
|
|
|
|
|
return "", err
|
|
|
|
|
}
|
|
|
|
|
return string(decoded), nil
|
|
|
|
|
}
|
2026-07-03 17:00:43 +08:00
|
|
|
|
2026-07-03 19:37:53 +08:00
|
|
|
func FormatBytes(bytes int64) string {
|
|
|
|
|
const (
|
|
|
|
|
KB = 1024
|
|
|
|
|
MB = KB * 1024
|
|
|
|
|
GB = MB * 1024
|
|
|
|
|
TB = GB * 1024
|
|
|
|
|
)
|
|
|
|
|
switch {
|
|
|
|
|
case bytes >= TB:
|
|
|
|
|
return fmt.Sprintf("%.1f TB", float64(bytes)/float64(TB))
|
|
|
|
|
case bytes >= GB:
|
|
|
|
|
return fmt.Sprintf("%.1f GB", float64(bytes)/float64(GB))
|
|
|
|
|
case bytes >= MB:
|
|
|
|
|
return fmt.Sprintf("%.1f MB", float64(bytes)/float64(MB))
|
|
|
|
|
case bytes >= KB:
|
|
|
|
|
return fmt.Sprintf("%.1f KB", float64(bytes)/float64(KB))
|
|
|
|
|
default:
|
|
|
|
|
return fmt.Sprintf("%d B", bytes)
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
func FormatNumber(n int64) string {
|
|
|
|
|
s := fmt.Sprintf("%d", n)
|
|
|
|
|
parts := []string{}
|
|
|
|
|
for i := len(s); i > 0; i -= 3 {
|
|
|
|
|
start := i - 3
|
|
|
|
|
if start < 0 {
|
|
|
|
|
start = 0
|
|
|
|
|
}
|
|
|
|
|
parts = append([]string{s[start:i]}, parts...)
|
|
|
|
|
}
|
|
|
|
|
return strings.Join(parts, ",")
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
func ParseBytesString(s string) int64 {
|
|
|
|
|
s = strings.TrimSpace(strings.ToLower(s))
|
|
|
|
|
if s == "" || s == "-" || s == "0" {
|
|
|
|
|
return 0
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
var multiplier int64 = 1
|
|
|
|
|
switch {
|
|
|
|
|
case strings.HasSuffix(s, "tb"):
|
|
|
|
|
multiplier = 1024 * 1024 * 1024 * 1024
|
|
|
|
|
s = strings.TrimSuffix(s, "tb")
|
|
|
|
|
case strings.HasSuffix(s, "gb"):
|
|
|
|
|
multiplier = 1024 * 1024 * 1024
|
|
|
|
|
s = strings.TrimSuffix(s, "gb")
|
|
|
|
|
case strings.HasSuffix(s, "mb"):
|
|
|
|
|
multiplier = 1024 * 1024
|
|
|
|
|
s = strings.TrimSuffix(s, "mb")
|
|
|
|
|
case strings.HasSuffix(s, "kb"):
|
|
|
|
|
multiplier = 1024
|
|
|
|
|
s = strings.TrimSuffix(s, "kb")
|
|
|
|
|
case strings.HasSuffix(s, "b"):
|
|
|
|
|
s = strings.TrimSuffix(s, "b")
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
s = strings.TrimSpace(s)
|
|
|
|
|
val, err := strconv.ParseFloat(s, 64)
|
|
|
|
|
if err != nil {
|
|
|
|
|
return 0
|
|
|
|
|
}
|
|
|
|
|
return int64(val * float64(multiplier))
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
func FormatTime(t *int64) string {
|
|
|
|
|
if t == nil {
|
|
|
|
|
return "N/A"
|
|
|
|
|
}
|
|
|
|
|
return time.UnixMilli(*t).Format("2006-01-02 15:04:05")
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-03 17:00:43 +08:00
|
|
|
func IsValidString(v interface{}) bool {
|
|
|
|
|
str, ok := v.(string)
|
|
|
|
|
return ok && str != ""
|
|
|
|
|
}
|
Feat(ingestion): align image to MinIO upload, unify chunk-id computation and add PPT parsing support (#17111)
## Summary
Align the Go ingestion pipeline with Python's `image` → `img_id`
persistence semantics, and unify the chunk-id computation across all
paths.
### Changes
**1. Image upload at chunker stage **
- Add `ImageUploader` type and `DefaultImageUploader` in
`internal/ingestion/component/image_uploader.go` — the write-side
counterpart to `FetchBinary`, storing raw image bytes at `(bucket=kbID,
key=chunkID)`, no re-encoding.
- Add `uploadOneImage` — pure upload primitive (bytes in, `img_id` out),
does not touch chunk maps.
- Add `uploadChunkImages` / `uploadChunkImage` — caller-side helper:
decodes `image` from a chunk, uploads bytes, writes `ck["img_id"]`,
`delete(ck,"image")` , bounded by a process-wide semaphore (default 10,
env `MAX_CONCURRENT_MINIO`).
- Wire via `imageUploadDecorator` in `register.go`: every chunker runs
the upload pass at invocation time, writing `ck["id"]` before upload and
dropping image bytes right after — peak memory = single chunk image
lifetime.
**2. Unify chunk-id computation**
- Consolidate three separate id-computation paths (`component.ChunkID`,
`task.ChunkID`, inline `FormatUint` in API) into one:
`common.ChunkID(docID, text string)`, using `%016x` +
`xxhash.Sum64String(text+docID)` (matching Python `hexdigest()`).
- The chunker decorator writes `ck["id"]` via `common.ChunkID`; the
persist stage (`ProcessChunksForPipeline`) falls back to the same
function (`if !exists id`).
- The API AddChunk path now also calls `common.ChunkID` instead of the
divergent `FormatUint(xxhash.Sum64(...))` — fixing a pre-existing
inconsistency.
- Delete `internal/ingestion/component/chunk_id.go` and
`internal/ingestion/task/chunk_builder.go` (both were pure forwarding
shells).
**3. Preserve `img_id` (never deleted)**
- `img_id` is a persistent index field (Infinity, OB) and the only
consumer-side reference for image retrieval; it is NEVER removed from
the chunk map. Only `image` (raw data URL) is dropped after upload.
**4. PPT parser support**
Previously PPT parsing failed. Add support to parse.
### Key design decisions
| Decision | Choice |
|----------|--------|
| Upload timing | Chunker stage (not persist), so image bytes are
dropped immediately — bounds peak memory to one chunk image |
| Upload concurrency | Process-wide semaphore, default 10 (matches
Python `minio_limiter`), env `MAX_CONCURRENT_MINIO` |
| Image encoding | Store as-is, no JPEG re-encoding (unlike Python) |
| `img_id` format | `"<kb_id>-<chunk_id>"` — matches Python
task_executor path |
| id function | Single `common.ChunkID(docID, text)`, concatenation
`text+docID` inside hash (matching Python) |
| `removeInternalChunkFields` | Retains `delete(ck,"image")` as
defensive fallback for non-chunker paths |
### Files touched
| File | Change |
|------|--------|
| `internal/common/format.go` | Add `ChunkID(docID, text)` |
| `internal/common/format_test.go` | Add ChunkID golden-value test |
| `internal/ingestion/component/image_uploader.go` | Add `ImageUploader`
type + `DefaultImageUploader` |
| `internal/ingestion/component/chunker/image_upload.go` | Add
`uploadOneImage`, `uploadChunkImages`, `uploadChunkImage`,
`decodeChunkImage`, semaphore |
| `internal/ingestion/component/chunker/image_upload_test.go` | Tests:
upload/drop, skip, no-image, concurrency, missing-id error |
| `internal/ingestion/component/chunker/register.go` | Add
`imageUploadDecorator` (writes `ck["id"]`, runs upload) |
| `internal/ingestion/task/chunk_process.go` | Use `common.ChunkID` for
persist fallback |
| `internal/service/chunk/chunk.go` | Use `common.ChunkID` instead of
`FormatUint` |
| `internal/ingestion/component/chunk_id.go` | **Deleted** (moved to
`common`) |
| `internal/ingestion/task/chunk_builder.go` | **Deleted** (shell, no
callers left) |
| `internal/ingestion/task/chunk_builder_test.go` | **Deleted** (test
migrated to `common/format_test.go`) |
### Verification
```
bash build.sh --test ./internal/service/chunk/... ./internal/common/... ./internal/ingestion/component/... ./internal/ingestion/task/...
→ ok service/chunk / common / component / chunker / schema / task
```
2026-07-20 19:33:51 +08:00
|
|
|
|
|
|
|
|
// ChunkID generates a deterministic chunk identifier matching Python's:
|
|
|
|
|
//
|
|
|
|
|
// xxhash.xxh64((content_with_weight + str(doc_id)).encode("utf-8", "surrogatepass")).hexdigest()
|
|
|
|
|
//
|
|
|
|
|
// The concatenation inside the hash is text+docID (matching Python); the
|
|
|
|
|
// parameter order is (docID, text) so the scope/identity argument comes first.
|
|
|
|
|
// This is the single shared implementation used by the ingestion pipeline (for
|
|
|
|
|
// image object keys and index document ids) and by the API (for directly
|
|
|
|
|
// created chunks), so the two paths always produce the same id from the same
|
|
|
|
|
// (docID, text) pair.
|
|
|
|
|
func ChunkID(docID, text string) string {
|
|
|
|
|
return fmt.Sprintf("%016x", xxhash.Sum64String(text+docID))
|
|
|
|
|
}
|