Files
ragflow/internal/tokenizer/bpe_loader_anchors_test.go

78 lines
2.8 KiB
Go
Raw Normal View History

fix(tokenizer): load cl100k BPE table from disk instead of failing silently offline (#17712) ## Summary RAGFlow's Go tokenizer silently returned **0 tokens for every string** whenever the `cl100k_base` BPE table could not be loaded — which is the normal case for an offline/air-gapped Go server. This PR makes the loader resolve the table from disk (where RAGFlow actually ships it) and fail loudly when it is genuinely missing. ## Root cause `tiktoken-go`'s stock loader downloads the encoding table over HTTP and caches it under `TIKTOKEN_CACHE_DIR`. That does not work for RAGFlow: - `TIKTOKEN_CACHE_DIR` is exported **only inside the Python process** (`common/token_utils.py`). `docker/entrypoint.sh` launches the Go binary (`bin/ragflow_server`) from a shell, so the Go process never inherits the variable. - The Dockerfile *does* ship the table (under its sha1 name in the working directory), but nothing told the Go side to look there. - Reaching `openaipublic.blob.core.windows.net` at runtime is not an option for air-gapped installs, and is unreliable where that host is blocked. The failure was **silent**: `NumTokensFromString` returns `0` when the encoder fails to build, and a `sync.Once` memoizes that error for the process lifetime. Every token count became `0`, so chunk merging never crossed its token budget and an entire document collapsed into a single chunk. Python has no such failure mode because its encoder is built at import time (a missing table aborts startup instead of degrading). ## Fix Register a local-only `BpeLoader` via `tiktoken.SetBpeLoader` (`internal/tokenizer/bpe_loader.go`) that resolves the table from disk **only**, in priority order: 1. `TIKTOKEN_CACHE_DIR` / `DATA_GYM_CACHE_DIR` (honored so operators who already configured one keep working). 2. The working directory, the executable's directory, and all of their ancestors — matching the Dockerfile layout (table under its sha1 name in the install root). 3. A `ragflow_deps/<basename>` checkout produced by `ragflow_deps/download_deps.py`. It **never performs network I/O**. When nothing is found it returns an error listing every path it tried (pointing at `download_deps.py` or `TIKTOKEN_CACHE_DIR`), so a genuinely missing table fails loudly instead of degrading to zero. ## Test plan - `internal/tokenizer/bpe_loader_test.go` (unit tier, runs under `bash build.sh --test ./internal/tokenizer/...`): - Loader reads from `TIKTOKEN_CACHE_DIR`, `DATA_GYM_CACHE_DIR`, the sha1-named file in the working dir, and the bundled `ragflow_deps/` name. - Explicit cache dir wins over the bundled vocab. - A malformed table is reported as an error rather than skipped. - A genuinely missing table reports the candidates it tried (no network attempt). - `NumTokensFromString` matches Python-derived anchors (`""`→0, `"hello"`→1, `"hello world"`→2, `"hello, world!"`→4, `"世界"`→3, `"Hello 世界 🌍"`→8, `"RAGFlow"`→3). ## Notes - `.github/workflows/tests.yml` currently excludes `internal/tokenizer` from `go test`, so these tests do not run in CI. The tokenizer fix is exercised in CI indirectly via the chunker package once a token-count-sensitive parity case lands (tracked separately). Consider including `internal/tokenizer` in CI as a follow-up. - Supported deployments already ship the table (`download_deps.py` → `ragflow_deps/cl100k_base.tiktoken`; Dockerfile → `<sha1>` in cwd), so no `ENV` change is required for the fix to take effect. Setting `ENV TIKTOKEN_CACHE_DIR` in the Dockerfile remains a cheap belt-and-suspenders hardening that can be done separately. 🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy) --------- Co-authored-by: CodeBuddy <noreply@codebuddy.ai> Co-authored-by: CodeBuddy Code <noreply@cnb.cool> Co-authored-by: CodeBuddy <noreply@tencent.com>
2026-08-03 19:03:08 +08:00
//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//
//go:build manual
package tokenizer
import "testing"
// TestNumTokensFromString_MatchesPythonAnchors pins exact counts taken from the
// Python reference suite (test/unit_test/common/test_token_utils.py:28-49) and
// from common.token_utils.num_tokens_from_string for the CJK cases. The corpus
// was later expanded to ~24 entries spanning ASCII, punctuation, digits,
// whitespace, newlines, CJK, emoji, mixed-language, and code-like strings, all
// recomputed against tiktoken's cl100k encoder.
//
// Exact values matter more than they look. NumTokensFromString swallows loader
// errors and returns 0, so an assertion of the form "> 0" passes for an empty
// string and fails to notice a dead encoder — which is precisely how the
// offline breakage stayed invisible. Pinning the numbers also catches loading a
// structurally valid but wrong table.
//
// This test needs the real cl100k_base table on disk (TIKTOKEN_CACHE_DIR,
// the Dockerfile's /ragflow/<sha1> file, or ragflow_deps/cl100k_base.tiktoken),
// so it is tagged `manual` and runs only under `build.sh --test-manual`,
// which the docker builder provisions with /usr/share/infinity/resource.
func TestNumTokensFromString_MatchesPythonAnchors(t *testing.T) {
anchors := []struct {
in string
want int
}{
{"", 0},
{"hello", 1},
{"hello world", 2},
{"hello, world!", 4},
{"世界", 3},
{"Hello 世界 🌍", 8},
{"RAGFlow", 3},
{"1234567890", 4},
{"a b", 3},
{"hello\nworld", 3},
{"user@example.com", 3},
{"https://example.com/path?x=1", 9},
{"func main() {}", 4},
{"aaaaaaaaaa", 2},
{"中文字符测试", 4},
{"🚀🔥", 6},
{"state-of-the-art", 4},
{`"quoted"`, 3},
{"The quick brown fox jumps over the lazy dog.", 10},
{"Café naïve résumé", 8},
{"x² + y² = z²", 8},
{"混合 English 和 中文 的 sentence。", 10},
{"tokenization is the process of splitting text into tokens", 10},
{"人工智能正在改变世界,这是毫无疑问的事实。", 25},
{"SELECT * FROM users WHERE id = 1;", 10},
{"こんにちは世界", 4},
}
for _, tc := range anchors {
if got := NumTokensFromString(tc.in); got != tc.want {
t.Errorf("NumTokensFromString(%q) = %d, want %d", tc.in, got, tc.want)
}
}
}