mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
468cc4b7d7
* fix(embeddings): honor query and document prompts locally * fix(embeddings): require sentence-transformers >=5.0 for local asymmetric encoding encode_query()/encode_document() only exist from sentence-transformers 5.0 onwards. The local-ml extra pinned >=3.3.0, so on 4.x the new code path was an AttributeError at the first encode (recall/retain), not at startup. The extra was only accidentally safe because it also pins transformers>=5.5.0, which ST <5 caps out; docker/docker-compose/custom-models/Dockerfile mirrors the pins with transformers>=4.53.0 and could genuinely resolve to ST 4.x. Also: - assert the real SentenceTransformer class exposes both entry points; the existing test drives a MagicMock, so it passes on any version - explain why the model's own entry points are used instead of prefixing here, and note that prompt-less models are unaffected - document the one case that needs a re-index: a local model that instructs the stored side as well as the search side --------- Co-authored-by: jpmf33 <265638852+jpmf33@users.noreply.github.com> Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
35 lines
1.6 KiB
Docker
35 lines
1.6 KiB
Docker
# Example: custom Hindsight image with non-default local models baked in.
|
|
#
|
|
# Use this pattern in production when you run a non-default embedder or
|
|
# reranker. Baking models into the image removes the runtime dependency on
|
|
# HuggingFace and lets the container registry handle caching per node, so
|
|
# you don't need a model-cache PVC.
|
|
#
|
|
# Built on top of the slim image so only the deps and models you actually
|
|
# use end up in the final image.
|
|
FROM ghcr.io/vectorize-io/hindsight:latest-slim
|
|
|
|
# Install the local-ml deps required to load sentence-transformers /
|
|
# cross-encoder models at runtime. Pinned ranges mirror hindsight-api-slim's
|
|
# `local-ml` extra in hindsight-api-slim/pyproject.toml. Use `uv pip
|
|
# install` against the image's venv explicitly: the slim image's venv was
|
|
# created by `uv sync` and does not ship its own `pip`, so a bare
|
|
# `pip install` would fall back to user site-packages and not be visible
|
|
# to the runtime python.
|
|
RUN uv pip install --python /app/api/.venv/bin/python --no-cache \
|
|
'sentence-transformers>=5.0.0' \
|
|
'transformers>=4.53.0' \
|
|
'torch>=2.6.0'
|
|
|
|
# Pre-download the models you want to use. Replace these with your own.
|
|
# The defaults bundled in the full image are BAAI/bge-small-en-v1.5 and
|
|
# cross-encoder/ms-marco-MiniLM-L-6-v2; here we pick multilingual variants
|
|
# as a concrete non-default example.
|
|
ARG EMBEDDER=sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
|
|
ARG RERANKER=cross-encoder/mmarco-mMiniLMv2-L12-H384-v1
|
|
ENV HF_HUB_DOWNLOAD_TIMEOUT=600
|
|
RUN python -c "\
|
|
from sentence_transformers import SentenceTransformer, CrossEncoder; \
|
|
SentenceTransformer('${EMBEDDER}'); \
|
|
CrossEncoder('${RERANKER}')"
|