mirror of
https://github.com/infiniflow/ragflow.git
synced 2026-08-22 08:13:21 +08:00
### Summary Closes #18414. `rag/res/term.freq` is not shipped, and both term-weight implementations therefore assigned the same `300` fallback frequency to every lowercase Latin token. With no tokenizer frequency, NER, or POS signal, function words and content words received identical lexical boosts. This PR adds the same bounded out-of-vocabulary prior to Python and Go: - Use it only when the explicit DF dictionary or tokenizer has no frequency. - Count Latin, Greek, and Cyrillic letters, including uppercase and accented forms. - Keep the existing frequency of `300` for words up to three letters, halve it every two additional letters, and clamp it at `10`. - Reject digits, underscores, and logographic terms so Chinese and other existing fine-grained-tokenizer paths are unchanged. - Treat an absent optional `term.freq` as the supported fallback path without a startup warning, while still logging inaccessible or malformed dictionaries. A corpus-derived table was intentionally not added: that would require provenance/licensing decisions, language detection, and handling cross-language homographs. The bounded prior is deterministic, dependency-free, and fixes the equal-weight degradation for whitespace-delimited alphabetic languages without claiming corpus-specific precision. Python and Go consume one shared fixture covering ASCII, uppercase, accented Latin, Greek, Cyrillic, separators, invalid mixed tokens, and a CJK non-match. Both sides also verify the issue's ordering (`was < largest < supplier < equipment`) and that an explicit dictionary entry still takes precedence. Co-authored-by: Loong <184861530+yzl0ng@users.noreply.github.com>
(1). Deploy RAGFlow services and images
https://ragflow.io/docs/build_docker_image
(2). Configure the required environment for testing
Install Python dependencies (including test dependencies):
uv sync --python 3.13 --only-group test --no-default-groups --frozen
Activate the environment:
source .venv/bin/activate
Install SDK:
uv pip install sdk/python
Modify the .env file: Add the following code:
COMPOSE_PROFILES=${COMPOSE_PROFILES},tei-cpu
TEI_MODEL=BAAI/bge-small-en-v1.5
RAGFLOW_IMAGE=infiniflow/ragflow:v0.27.0 #Replace with the image you are using
Start the container(wait two minutes):
docker compose -f docker/docker-compose.yml up -d
(3). Test Elasticsearch
a) Run sdk tests against Elasticsearch:
export HTTP_API_TEST_LEVEL=p2
export HOST_ADDRESS=http://127.0.0.1:9380 # Ensure that this port is the API port mapped to your localhost
pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_sdk_api
b) Run http api tests against Elasticsearch:
pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_http_api
(4). Test Infinity
Modify the .env file:
DOC_ENGINE=${DOC_ENGINE:-infinity}
Start the container:
docker compose -f docker/docker-compose.yml down -v
docker compose -f docker/docker-compose.yml up -d
a) Run sdk tests against Infinity:
DOC_ENGINE=infinity pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_sdk_api
b) Run http api tests against Infinity:
DOC_ENGINE=infinity pytest -s --tb=short --level=${HTTP_API_TEST_LEVEL} test/testcases/test_http_api