mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-08-04 23:00:30 +08:00
## Summary RAGFlow's Go tokenizer silently returned **0 tokens for every string** whenever the `cl100k_base` BPE table could not be loaded — which is the normal case for an offline/air-gapped Go server. This PR makes the loader resolve the table from disk (where RAGFlow actually ships it) and fail loudly when it is genuinely missing. ## Root cause `tiktoken-go`'s stock loader downloads the encoding table over HTTP and caches it under `TIKTOKEN_CACHE_DIR`. That does not work for RAGFlow: - `TIKTOKEN_CACHE_DIR` is exported **only inside the Python process** (`common/token_utils.py`). `docker/entrypoint.sh` launches the Go binary (`bin/ragflow_server`) from a shell, so the Go process never inherits the variable. - The Dockerfile *does* ship the table (under its sha1 name in the working directory), but nothing told the Go side to look there. - Reaching `openaipublic.blob.core.windows.net` at runtime is not an option for air-gapped installs, and is unreliable where that host is blocked. The failure was **silent**: `NumTokensFromString` returns `0` when the encoder fails to build, and a `sync.Once` memoizes that error for the process lifetime. Every token count became `0`, so chunk merging never crossed its token budget and an entire document collapsed into a single chunk. Python has no such failure mode because its encoder is built at import time (a missing table aborts startup instead of degrading). ## Fix Register a local-only `BpeLoader` via `tiktoken.SetBpeLoader` (`internal/tokenizer/bpe_loader.go`) that resolves the table from disk **only**, in priority order: 1. `TIKTOKEN_CACHE_DIR` / `DATA_GYM_CACHE_DIR` (honored so operators who already configured one keep working). 2. The working directory, the executable's directory, and all of their ancestors — matching the Dockerfile layout (table under its sha1 name in the install root). 3. A `ragflow_deps/<basename>` checkout produced by `ragflow_deps/download_deps.py`. It **never performs network I/O**. When nothing is found it returns an error listing every path it tried (pointing at `download_deps.py` or `TIKTOKEN_CACHE_DIR`), so a genuinely missing table fails loudly instead of degrading to zero. ## Test plan - `internal/tokenizer/bpe_loader_test.go` (unit tier, runs under `bash build.sh --test ./internal/tokenizer/...`): - Loader reads from `TIKTOKEN_CACHE_DIR`, `DATA_GYM_CACHE_DIR`, the sha1-named file in the working dir, and the bundled `ragflow_deps/` name. - Explicit cache dir wins over the bundled vocab. - A malformed table is reported as an error rather than skipped. - A genuinely missing table reports the candidates it tried (no network attempt). - `NumTokensFromString` matches Python-derived anchors (`""`→0, `"hello"`→1, `"hello world"`→2, `"hello, world!"`→4, `"世界"`→3, `"Hello 世界 🌍"`→8, `"RAGFlow"`→3). ## Notes - `.github/workflows/tests.yml` currently excludes `internal/tokenizer` from `go test`, so these tests do not run in CI. The tokenizer fix is exercised in CI indirectly via the chunker package once a token-count-sensitive parity case lands (tracked separately). Consider including `internal/tokenizer` in CI as a follow-up. - Supported deployments already ship the table (`download_deps.py` → `ragflow_deps/cl100k_base.tiktoken`; Dockerfile → `<sha1>` in cwd), so no `ENV` change is required for the fix to take effect. Setting `ENV TIKTOKEN_CACHE_DIR` in the Dockerfile remains a cheap belt-and-suspenders hardening that can be done separately. 🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy) --------- Co-authored-by: CodeBuddy <noreply@codebuddy.ai> Co-authored-by: CodeBuddy Code <noreply@cnb.cool> Co-authored-by: CodeBuddy <noreply@tencent.com>
78 lines
2.8 KiB
Go
78 lines
2.8 KiB
Go
//
|
|
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
|
|
//
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
// you may not use this file except in compliance with the License.
|
|
// You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
//
|
|
|
|
//go:build manual
|
|
|
|
package tokenizer
|
|
|
|
import "testing"
|
|
|
|
// TestNumTokensFromString_MatchesPythonAnchors pins exact counts taken from the
|
|
// Python reference suite (test/unit_test/common/test_token_utils.py:28-49) and
|
|
// from common.token_utils.num_tokens_from_string for the CJK cases. The corpus
|
|
// was later expanded to ~24 entries spanning ASCII, punctuation, digits,
|
|
// whitespace, newlines, CJK, emoji, mixed-language, and code-like strings, all
|
|
// recomputed against tiktoken's cl100k encoder.
|
|
//
|
|
// Exact values matter more than they look. NumTokensFromString swallows loader
|
|
// errors and returns 0, so an assertion of the form "> 0" passes for an empty
|
|
// string and fails to notice a dead encoder — which is precisely how the
|
|
// offline breakage stayed invisible. Pinning the numbers also catches loading a
|
|
// structurally valid but wrong table.
|
|
//
|
|
// This test needs the real cl100k_base table on disk (TIKTOKEN_CACHE_DIR,
|
|
// the Dockerfile's /ragflow/<sha1> file, or ragflow_deps/cl100k_base.tiktoken),
|
|
// so it is tagged `manual` and runs only under `build.sh --test-manual`,
|
|
// which the docker builder provisions with /usr/share/infinity/resource.
|
|
func TestNumTokensFromString_MatchesPythonAnchors(t *testing.T) {
|
|
anchors := []struct {
|
|
in string
|
|
want int
|
|
}{
|
|
{"", 0},
|
|
{"hello", 1},
|
|
{"hello world", 2},
|
|
{"hello, world!", 4},
|
|
{"世界", 3},
|
|
{"Hello 世界 🌍", 8},
|
|
{"RAGFlow", 3},
|
|
{"1234567890", 4},
|
|
{"a b", 3},
|
|
{"hello\nworld", 3},
|
|
{"user@example.com", 3},
|
|
{"https://example.com/path?x=1", 9},
|
|
{"func main() {}", 4},
|
|
{"aaaaaaaaaa", 2},
|
|
{"中文字符测试", 4},
|
|
{"🚀🔥", 6},
|
|
{"state-of-the-art", 4},
|
|
{`"quoted"`, 3},
|
|
{"The quick brown fox jumps over the lazy dog.", 10},
|
|
{"Café naïve résumé", 8},
|
|
{"x² + y² = z²", 8},
|
|
{"混合 English 和 中文 的 sentence。", 10},
|
|
{"tokenization is the process of splitting text into tokens", 10},
|
|
{"人工智能正在改变世界,这是毫无疑问的事实。", 25},
|
|
{"SELECT * FROM users WHERE id = 1;", 10},
|
|
{"こんにちは世界", 4},
|
|
}
|
|
for _, tc := range anchors {
|
|
if got := NumTokensFromString(tc.in); got != tc.want {
|
|
t.Errorf("NumTokensFromString(%q) = %d, want %d", tc.in, got, tc.want)
|
|
}
|
|
}
|
|
}
|