* docs(docker): add guide and recipe for building CUDA standalone image
Provide a Docker Compose example and standalone Dockerfile recipe for
building a CUDA-enabled PyTorch image for NVIDIA GPU-accelerated local
embedding and reranker models. Document prerequisites, build steps, and
NVIDIA Container Toolkit configuration.
* docs(docker): correct CUDA recipe verification, drop x86-only pin, state size
Review follow-ups on the CUDA recipe:
- The README told users to verify GPU placement by grepping the logs for both
the embedding and cross-encoder device, but only the embedder logged one.
LocalSTCrossEncoder resolved _device_type and never reported it, so the
documented check showed a single line and looked like the reranker was still
on CPU. Log the device on both reranker init paths and show the real log
output in the README.
- Dropped the hardcoded `--platform=linux/amd64`. PyTorch ships cu126 wheels for
aarch64 too, and on an arm64 host the pin silently produced an emulated amd64
image that cannot reach the GPU at all. Documented instead that the build must
be native.
- Dropped `--index-strategy unsafe-best-match`. It relaxed the index isolation
that hindsight-api-slim/pyproject.toml deliberately sets up, and it was not
needed: resolving without it succeeds and yields the same package set.
- Documented the image size (~11 GB vs ~9 GB for the base) in installation.md
and the recipe README, since that cost is the reason no CUDA image is
published.
- Added a HINDSIGHT_VERSION build arg so the base tag can be pinned, fixed the
manual `docker build` context, and enabled RERANKER_LOCAL_FP16 in the compose
file (faster on GPU, quality-identical).
Verified: `uv pip install` inside the base image replaces only torch
(2.10.0+cpu -> 2.10.0+cu126) and adds the nvidia/cuda runtime wheels; compose
config validates; docs skill regenerates clean; test_local_cross_encoder.py
passes (21).
Claude-Session: https://claude.ai/code/session_017ufCz6qrNxn36Stug7ek8A
---------
Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
Managed PostgreSQL distributions install the pg_search functions (score, boolean,
match) under a schema of their choosing — pgsearch rather than paradedb on the
distribution in #3757 — so the hardcoded paradedb. qualification made the
pg_search backend unusable there without a custom image or wrapper functions.
Adds HINDSIGHT_API_TEXT_SEARCH_EXTENSION_PG_SEARCH_FUNCTION_SCHEMA (default
paradedb, validated as a plain PG identifier since it is interpolated into raw
SQL) and threads it through both builders that qualify those calls: the
memory-recall BM25 arm and the knowledge-pages arm. Index creation needs no
change — USING bm25(...) resolves the access method globally. The pdb tokenizer
pseudo-types stay fixed; they come from ParadeDB's own install DDL and are
independent of the extension's target schema.
Fixes#3757.
* docs(examples): add TEI embeddings + reranker docker-compose example
Adds docker/docker-compose/tei/ — a runnable Compose stack that serves
embeddings and reranking from two HuggingFace Text Embeddings Inference
(TEI) sidecars, with the slim Hindsight image talking to them via
HINDSIGHT_API_{EMBEDDINGS,RERANKER}_PROVIDER=tei. Serves Hindsight's
default models so it's a drop-in 'move embeddings/reranking onto TEI'
demo. Links it from the TEI section of the models docs.
Verified end-to-end: retain + recall return the expected memory with
both TEI semantic and reranker scores populated.
* docs(examples): prod-like TEI tuning — bge-reranker-base + throughput flags
Swap the reranker to BAAI/bge-reranker-base (the cross-encoder commonly
paired with bge-small embeddings on dedicated inference servers) and carry
prod-like TEI throughput flags on both services (--max-concurrent-requests,
--max-batch-tokens, --max-client-batch-size) instead of TEI's bare defaults,
so the example doubles as a starting point for real deployments.
* fix(compose): use HINDSIGHT_API_LLM_API_KEY env var
* docs(compose): point run examples at HINDSIGHT_API_LLM_API_KEY
The compose files no longer read OPENAI_API_KEY, so every doc telling
users to export it was left describing a variable nothing reads.
- custom-models/README.md + compose header: export the correct var
- timescale/README.md quick start, prereq and env-var table
- timescale/.env.example: compose's project directory is the compose
file's own directory, so this file IS auto-loaded - naming the wrong
var here silently dropped the key on the documented happy path
---------
Co-authored-by: Ish Fuseini <me@ishfuesini.com>
Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
The PGroonga example referenced groonga/pgroonga:latest-debian-pg17, which has no registry manifest, so the documented Compose stack could not build.\n\nPin the current PGroonga 4.0.8 PostgreSQL 17 image and install pgvector 0.8.6 from the PGDG repository already configured by that base image. This removes the source clone and build toolchain while keeping the example on PostgreSQL 17 for parity with neighboring recipes.\n\nContext:\n- Fixes #3311.\n- Verified on arm64 with a disposable PostgreSQL instance.\n- CREATE EXTENSION vector and pgroonga both succeed.\n- Mixed Korean/English PGroonga search and pgvector distance queries pass.\n- PostgreSQL 18 deployment work remains a separate operational concern.
* fix(embeddings): honor query and document prompts locally
* fix(embeddings): require sentence-transformers >=5.0 for local asymmetric encoding
encode_query()/encode_document() only exist from sentence-transformers 5.0
onwards. The local-ml extra pinned >=3.3.0, so on 4.x the new code path was an
AttributeError at the first encode (recall/retain), not at startup. The extra
was only accidentally safe because it also pins transformers>=5.5.0, which ST
<5 caps out; docker/docker-compose/custom-models/Dockerfile mirrors the pins
with transformers>=4.53.0 and could genuinely resolve to ST 4.x.
Also:
- assert the real SentenceTransformer class exposes both entry points; the
existing test drives a MagicMock, so it passes on any version
- explain why the model's own entry points are used instead of prefixing here,
and note that prompt-less models are unaffected
- document the one case that needs a re-index: a local model that instructs the
stored side as well as the search side
---------
Co-authored-by: jpmf33 <265638852+jpmf33@users.noreply.github.com>
Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
Remove unreferenced backend helpers, stale UI/docs components, and
unused imports across the API, control plane, clients, and integrations.
Drop obsolete consolidated-observation helpers and unused scoring code,
clean orphaned React/docs components, and remove stale Radix dependencies.
Align release scripts, Helm docs, lockfiles, generated clients, and
current API examples with the package and endpoint surface still in use.
Hindsight's published image deliberately omits llama-cpp-python to keep
the image small, so setting HINDSIGHT_API_LLM_PROVIDER=llamacpp directly
against ghcr.io/vectorize-io/hindsight fails with ModuleNotFoundError.
Adds a docker-compose recipe that runs the official llama.cpp server
container as a sidecar and points Hindsight's openai provider at it via
HINDSIGHT_API_LLM_BASE_URL. Verified end-to-end against
ghcr.io/ggml-org/llama.cpp:server pulling Gemma 4 E2B from HuggingFace.
The named volume is mounted at /root/.cache/huggingface (where
llama-server actually caches downloads) so the GGUF survives stack
recreation. README documents the CPU perf reality and how to flip the
relevant blocks for NVIDIA GPU acceleration.
Also links the recipe from the "Built-in llama.cpp" tip in the models
docs so users following the docs find the Docker setup.
* docs(models): add claude-code Docker recipe with host Max Plan auth
Adds a 'Running with host Max Plan auth in Docker (Linux)' subsection
under the existing Claude Code Setup docs. Documents the bind-mount
surface required to run HINDSIGHT_API_LLM_PROVIDER=claude-code inside
the standalone image: host claude CLI, single-file credential mounts,
the v2.1.128+ binary override for the bundled-binary protocol issue,
and the post-run chown/symlink steps.
Restates the personal-use-only constraint inline so the Docker recipe
isn't read as a production pattern. Verified on linux/amd64 per the
contributor's report; macOS and Windows paths are noted as not yet
covered.
Closes#1480
* refactor: move claude-code Docker recipe from docs to docker/docker-compose/
Instead of documenting the Docker recipe inline in models.mdx, create a
dedicated docker/docker-compose/claude-code/ setup following the existing
pattern (custom-models, external-pg, etc.).
- docker-compose.yaml: converts the docker run command into a Compose service
with all bind mounts, env vars, and ports
- README.md: full documentation including prerequisites, quick start,
post-setup steps, and detailed notes on every bind mount
- Reverts the models.mdx addition per review feedback
Allow ParadeDB pg_search BM25 indexes to be created with a configured
tokenizer via HINDSIGHT_API_TEXT_SEARCH_EXTENSION_PG_SEARCH_TOKENIZER.
Validate supported tokenizer values and thread the setting through
startup reconciliation, Alembic index creation paths, Docker examples,
docs, generated docs, and tests.
The default remains unset so existing pg_search deployments continue to
use ParadeDB's default tokenizer unless explicitly configured. Changing
the value for an existing database still requires rebuilding the
pg_search indexes or recreating the database.
* feat(api): add ParadeDB pg_search as Citus-compatible BM25 backend
Adds a fourth value (`pg_search`) for `HINDSIGHT_API_TEXT_SEARCH_EXTENSION`
alongside the existing `native`, `vchord`, and `pg_textsearch`. ParadeDB
pg_search is the only true-BM25 backend that works on a Citus distributed
Postgres cluster, so this unblocks horizontally scaled deployments.
The retrieval arm builds the @@@ predicate via paradedb.boolean(should =>
ARRAY[paradedb.match('text', $4), ...]) since @@@ on the key_field requires
field-qualified terms; this preserves multi-field coverage (text + context
+ text_signals) without needing query string interpolation.
Includes a docker-compose example under docker/docker-compose/pg_search/
based on the official paradedb/paradedb:latest-pg17 image.
Closes#1754
* fix: accept pgroonga in n9i0 migration; clarify consolidator search_vector comment
- n9i0 (learnings + pinned_reflections) validation now permits 'pgroonga',
treating it as native at this migration stage. ensure_text_search_extension()
at startup converts the reflections table (renamed from pinned_reflections in
p1k2l3m4n5o6) to pgroonga structures; the learnings table is dropped in the
same later migration so its transient native column never reaches steady state.
Without this, pgroonga users hit ValueError on a fresh install.
- consolidator.py single-observation INSERT: the previous comment claimed
search_vector was GENERATED ALWAYS, but migration p4q5r6s7t8u9 dropped that
expression. Updated to reflect current behavior and flag the resulting gap
for native (observations land with NULL search_vector and are not BM25-
searchable until reflected/re-ingested) so a follow-up can address it.
* chore: regenerate hindsight-docs skill after rebase
Rebasing onto main pulled in hindsight-docs/ changes from #1704
(Codex OAuth embeddings) and #1538 (pgroonga). Re-run the
generate-docs-skill.sh generator so the cached
skills/hindsight-docs/references/developer/configuration.md mirror
matches the current developer docs and verify-generated-files passes.
* feat(bm25): make native language configurable + opt-in pgroonga backend
Adds two new env-level config knobs and a new opt-in BM25 backend so users
can serve non-English banks (especially CJK) out of the box.
- HINDSIGHT_API_BM25_LANGUAGE drives the PostgreSQL text search dictionary
used by the native tsvector backend (default: english). Validated as a
PG identifier so it can be safely embedded in to_tsvector('<lang>', ...).
- HINDSIGHT_API_RETAIN_OUTPUT_LANGUAGE forces the fact extractor to emit
facts in the specified language regardless of source content's language.
Independent from bm25_language so users can mix indexing/extraction
languages deliberately.
- New 'pgroonga' option for HINDSIGHT_API_TEXT_SEARCH_EXTENSION. Uses
TokenBigram + NormalizerNFKC150 — single polyglot index handles English,
CJK, etc. simultaneously. Ships with a docker-compose recipe.
To support a per-deployment language, the GENERATED ALWAYS expression on
memory_units.search_vector (and reflections.search_vector) is dropped via
new alembic migration p4q5r6s7t8u9. The application now populates these
columns at INSERT time using the configured bm25_language.
* docs(bm25): rename env var to scope it to native; move multilingual content to dedicated page
- Rename HINDSIGHT_API_BM25_LANGUAGE → HINDSIGHT_API_TEXT_SEARCH_EXTENSION_NATIVE_LANGUAGE.
The setting only applies to the "native" backend (vchord/pg_textsearch/pgroonga
use their own tokenizers), so the env var name now reflects that scope. Field
renamed to text_search_extension_native_language.
- Trim configuration.md back to a brief env-var table + link. The expanded
multilingual / CJK / pgroonga content moves to the dedicated multilingual.md
page, alongside the existing LLM / embedding / reranker multilingual guidance.
* feat(llm-output-language): rename and broaden to cover retain + consolidation + reflect
Renames HINDSIGHT_API_RETAIN_OUTPUT_LANGUAGE → HINDSIGHT_API_LLM_OUTPUT_LANGUAGE
(field llm_output_language) and applies the same "respond exclusively in {lang}"
directive across every LLM-generated artifact:
- retain (fact extraction) — already wired, just renamed.
- consolidation (observations / mental models) — appended to the batch
consolidation prompt via a new llm_output_language parameter.
- reflect (response synthesis) — appended to the final-system prompt via a
new parameter threaded through run_reflect_agent and memory_engine.
The shared directive lives in engine/prompt_utils.output_language_directive
so all three pipelines build the same instruction from a single source.
* docs(multilingual): drop the backfill-after-language-change section
* docs(installation): bake custom models into image instead of PVC
Add a runnable example under `docker/docker-compose/custom-models/` that
extends the slim image and pre-downloads non-default embedder/reranker
models at build time. Document this as the recommended pattern for
production over enabling the Helm `modelCache` PVC: image layers cache
per node for free, while a PVC adds storage cost, pins pods to a node,
and needs lifecycle management on uninstall/upgrade. Add pointers from
the api/worker `modelCache` values in the chart to the new section.
Refs vectorize-io/hindsight#1383
* fix(docker/custom-models): install local-ml deps via uv into the venv
The slim image's venv at /app/api/.venv was created by uv sync and does
not ship its own pip, so a bare `pip install` falls through to the
system pip and lands the packages in /home/hindsight/.local — invisible
to the venv python that runs hindsight-api at runtime. Use
`uv pip install --python /app/api/.venv/bin/python` to install into the
venv directly. Verified the resulting image loads both baked-in models
with HF_HUB_OFFLINE=1.
* docs(installation): trim custom-models section to a tip and pointer
The Dockerfile/compose example in docker/docker-compose/custom-models/
already has its own README explaining when to use it and why it beats
the modelCache PVC. The installation page only needs to point readers
there.
* feat: add AlloyDB ScaNN vector index support
* fix(hindsight_api): resolved SCANN index mismatch by deferring creation
- Added SCANN-aware vector index helpers with a 10k minimum-row threshold.
- Updated bank index generation to skip per-bank clauses and index creation when unsupported.
- Updated vector migrations to validate extension names and skip SCANN-specific index creation or drops.
- Updated migration reconciliation to use row counts and defer SCANN index recreation instead of mismatch errors.
- Added tests for SCANN deferral, per-bank index ineligibility, and migration SQL freeze behavior.
* docs: add AlloyDB Omni compose example
* feat: accept pdf, images and office files
* refactor: rename FileConverter to FileParser, simplify file retain API
- Rename engine/converters/ → engine/parsers/, FileConverter → FileParser,
ConverterRegistry → FileParserRegistry, MarkitdownConverter → MarkitdownParser
- Rename env var HINDSIGHT_API_FILE_CONVERTER → HINDSIGHT_API_FILE_PARSER
- Remove async/document_tags params from FileRetainRequest (always async now)
- Add retain_files() to Python Hindsight client and retainFiles() to TypeScript client
- Add sample.pdf to doc examples for working file upload demonstrations
- Update test_file_retain.py to use new parser names and always-async behavior
- Fix Go client missing os import in api_files.go
- Simplify postgresql.py storage to minimal schema
* fix: update rust CLI tests to use is_supported_file instead of is_text_file
* fix: patch Go api_files.go to add missing 'os' import after generation
* fix: insert 'os' import after 'net/url' in api_files.go patch for correct position
* chore: regenerate OpenAPI spec and clients (converter→parser description update)
* feat: support timescale pg_textsearch as text search extension
* refactor: deduplicate text search query in retrieve_semantic_bm25_combined
Instead of maintaining 3 complete query copies (native, vchord, pg_textsearch),
now we:
- Build backend-specific parts (score_expr, order_by, where_filter)
- Use a single query template with injected backend-specific parts
This makes maintenance easier - changes to the semantic CTE or overall structure
only need to be made once.
* feat: support for other text and vector search pg extensions
* test: increase timeout for test_batch_chunking_behavior to account for VectorChord BM25 tokenization overhead
* feat: support for other text and vector search pg extensions
* feat: add reverse proxy support
* improve
* improve
* improve
* improve
* improve
* fix: update integration test to use modern 'docker compose' command
- Replace 'docker-compose' with 'docker compose' (Docker Compose v2+)
- Add fallback to legacy docker-compose command for compatibility
- Fixes test failures on systems using Docker Compose plugin
* ci: trigger test rerun
* fix: make docker-compose detection more robust for CI
- Add get_docker_compose_command() to detect available command
- Use shutil.which() to check command availability
- Dynamically use correct command (docker compose vs docker-compose)
- Should work in both modern and legacy Docker environments
* fix: docker-compose networking in base path integration test
Fix connection refused error in test_reverse_proxy_simple_config by
handling host vs bridge networking modes correctly:
- Linux (host mode): nginx listens on 18080 directly, no port mapping
- Mac/Windows (bridge mode): nginx listens on 80, mapped to 18080
With host networking, port mappings in docker-compose don't work since
the container binds directly to the host's network namespace.