Files
ragflow/internal/tokenizer/bpe_loader_anchors_test.go
Jack e997fd655a fix(tokenizer): load cl100k BPE table from disk instead of failing silently offline (#17712)
## Summary

RAGFlow's Go tokenizer silently returned **0 tokens for every string**
whenever the `cl100k_base` BPE table could not be loaded — which is the
normal case for an offline/air-gapped Go server. This PR makes the
loader resolve the table from disk (where RAGFlow actually ships it) and
fail loudly when it is genuinely missing.

## Root cause

`tiktoken-go`'s stock loader downloads the encoding table over HTTP and
caches it under `TIKTOKEN_CACHE_DIR`. That does not work for RAGFlow:

- `TIKTOKEN_CACHE_DIR` is exported **only inside the Python process**
(`common/token_utils.py`). `docker/entrypoint.sh` launches the Go binary
(`bin/ragflow_server`) from a shell, so the Go process never inherits
the variable.
- The Dockerfile *does* ship the table (under its sha1 name in the
working directory), but nothing told the Go side to look there.
- Reaching `openaipublic.blob.core.windows.net` at runtime is not an
option for air-gapped installs, and is unreliable where that host is
blocked.

The failure was **silent**: `NumTokensFromString` returns `0` when the
encoder fails to build, and a `sync.Once` memoizes that error for the
process lifetime. Every token count became `0`, so chunk merging never
crossed its token budget and an entire document collapsed into a single
chunk. Python has no such failure mode because its encoder is built at
import time (a missing table aborts startup instead of degrading).

## Fix

Register a local-only `BpeLoader` via `tiktoken.SetBpeLoader`
(`internal/tokenizer/bpe_loader.go`) that resolves the table from disk
**only**, in priority order:

1. `TIKTOKEN_CACHE_DIR` / `DATA_GYM_CACHE_DIR` (honored so operators who
already configured one keep working).
2. The working directory, the executable's directory, and all of their
ancestors — matching the Dockerfile layout (table under its sha1 name in
the install root).
3. A `ragflow_deps/<basename>` checkout produced by
`ragflow_deps/download_deps.py`.

It **never performs network I/O**. When nothing is found it returns an
error listing every path it tried (pointing at `download_deps.py` or
`TIKTOKEN_CACHE_DIR`), so a genuinely missing table fails loudly instead
of degrading to zero.

## Test plan

- `internal/tokenizer/bpe_loader_test.go` (unit tier, runs under `bash
build.sh --test ./internal/tokenizer/...`):
- Loader reads from `TIKTOKEN_CACHE_DIR`, `DATA_GYM_CACHE_DIR`, the
sha1-named file in the working dir, and the bundled `ragflow_deps/`
name.
  - Explicit cache dir wins over the bundled vocab.
  - A malformed table is reported as an error rather than skipped.
- A genuinely missing table reports the candidates it tried (no network
attempt).
- `NumTokensFromString` matches Python-derived anchors (`""`→0,
`"hello"`→1, `"hello world"`→2, `"hello, world!"`→4, `"世界"`→3, `"Hello
世界 🌍"`→8, `"RAGFlow"`→3).

## Notes

- `.github/workflows/tests.yml` currently excludes `internal/tokenizer`
from `go test`, so these tests do not run in CI. The tokenizer fix is
exercised in CI indirectly via the chunker package once a
token-count-sensitive parity case lands (tracked separately). Consider
including `internal/tokenizer` in CI as a follow-up.
- Supported deployments already ship the table (`download_deps.py` →
`ragflow_deps/cl100k_base.tiktoken`; Dockerfile → `<sha1>` in cwd), so
no `ENV` change is required for the fix to take effect. Setting `ENV
TIKTOKEN_CACHE_DIR` in the Dockerfile remains a cheap
belt-and-suspenders hardening that can be done separately.

🤖 Generated with [CodeBuddy Code](https://cnb.cool/codebuddy)

---------

Co-authored-by: CodeBuddy <noreply@codebuddy.ai>
Co-authored-by: CodeBuddy Code <noreply@cnb.cool>
Co-authored-by: CodeBuddy <noreply@tencent.com>
2026-08-03 19:03:08 +08:00

78 lines
2.8 KiB
Go

//
// Copyright 2026 The InfiniFlow Authors. All Rights Reserved.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//
//go:build manual
package tokenizer
import "testing"
// TestNumTokensFromString_MatchesPythonAnchors pins exact counts taken from the
// Python reference suite (test/unit_test/common/test_token_utils.py:28-49) and
// from common.token_utils.num_tokens_from_string for the CJK cases. The corpus
// was later expanded to ~24 entries spanning ASCII, punctuation, digits,
// whitespace, newlines, CJK, emoji, mixed-language, and code-like strings, all
// recomputed against tiktoken's cl100k encoder.
//
// Exact values matter more than they look. NumTokensFromString swallows loader
// errors and returns 0, so an assertion of the form "> 0" passes for an empty
// string and fails to notice a dead encoder — which is precisely how the
// offline breakage stayed invisible. Pinning the numbers also catches loading a
// structurally valid but wrong table.
//
// This test needs the real cl100k_base table on disk (TIKTOKEN_CACHE_DIR,
// the Dockerfile's /ragflow/<sha1> file, or ragflow_deps/cl100k_base.tiktoken),
// so it is tagged `manual` and runs only under `build.sh --test-manual`,
// which the docker builder provisions with /usr/share/infinity/resource.
func TestNumTokensFromString_MatchesPythonAnchors(t *testing.T) {
anchors := []struct {
in string
want int
}{
{"", 0},
{"hello", 1},
{"hello world", 2},
{"hello, world!", 4},
{"世界", 3},
{"Hello 世界 🌍", 8},
{"RAGFlow", 3},
{"1234567890", 4},
{"a b", 3},
{"hello\nworld", 3},
{"user@example.com", 3},
{"https://example.com/path?x=1", 9},
{"func main() {}", 4},
{"aaaaaaaaaa", 2},
{"中文字符测试", 4},
{"🚀🔥", 6},
{"state-of-the-art", 4},
{`"quoted"`, 3},
{"The quick brown fox jumps over the lazy dog.", 10},
{"Café naïve résumé", 8},
{"x² + y² = z²", 8},
{"混合 English 和 中文 的 sentence。", 10},
{"tokenization is the process of splitting text into tokens", 10},
{"人工智能正在改变世界,这是毫无疑问的事实。", 25},
{"SELECT * FROM users WHERE id = 1;", 10},
{"こんにちは世界", 4},
}
for _, tc := range anchors {
if got := NumTokensFromString(tc.in); got != tc.want {
t.Errorf("NumTokensFromString(%q) = %d, want %d", tc.in, got, tc.want)
}
}
}