* test(api): guard against silently losing free-threading
A free-threaded CPython re-enables the GIL the moment it imports a C extension
that has not declared `Py_MOD_GIL_NOT_USED`, and says so only with a
RuntimeWarning. One new module-scope import therefore reverts the whole server
to single-threaded execution while every existing test still passes and the
process still serves traffic -- which is exactly how the three imports fixed in
the previous commit went unnoticed.
Adds a subprocess probe that imports the full API surface and asserts the GIL is
still disabled, plus a second case pinning the diagnostic form:
PYTHONWARNINGS="error:The global interpreter lock:RuntimeWarning"
which turns the warning fatal at the offending import (naming the module and the
import chain) while leaving unrelated RuntimeWarnings alone.
The probe runs in a subprocess because the GIL can only be re-enabled once per
interpreter, so an import already done in the pytest parent would mask a
regression.
Skipped unless `Py_GIL_DISABLED`, so it is inert on the 3.11 matrix and only
bites on a python3.14t job.
Verified both ways on python3.14t: passes on the fixed tree, and a negative
control that adds `import psycopg2` to the probe fails with the module named in
the assertion message.
* feat(api): fail loudly when a free-threaded build loses the GIL
A free-threaded CPython re-enables the GIL for the whole process the moment it
imports a C extension that has not declared `Py_MOD_GIL_NOT_USED`, and says so only
with a RuntimeWarning. Nothing crashes and nothing degrades visibly: the server
starts, serves traffic, and passes its tests, having quietly reverted to
single-threaded execution. One new module-scope import is enough.
Adds `hindsight_api/_free_threading.py` and calls it eagerly from the package
`__init__`, next to `apply_default_thread_limits()` and for the same class of reason:
both configure how the process executes, and both are worthless once the libraries
they govern have loaded.
`HINDSIGHT_API_FREE_THREADING` selects the mode:
strict (default on a free-threaded build) — the GIL re-enable warning becomes an
exception, so the offending import raises with the module and full import
chain in the traceback. A GIL already on at startup raises.
warn — log and continue, for bringing up a deployment whose dependencies are not
all ready.
off — no guard; used by the migration subprocess, which imports psycopg2 on
purpose.
Every mode is a no-op on a normal build, so this is inert on 3.11.
* ci(api): run the test suite on free-threaded CPython 3.14
Adds a `test-api (free-threaded 3.14)` job plus the two pieces of packaging it
needs, so a regression that silently re-enables the GIL fails CI instead of
quietly costing the deployment its parallelism.
The job asserts free-threading before running anything -- a bad build or an
already-taken GIL shows up as its own red step rather than as a mass of confusing
downstream failures -- and runs pytest under
PYTHONWARNINGS="error:The global interpreter lock:RuntimeWarning"
so a regression fails at the offending import with the module named.
It is deliberately NOT gated on `has_secrets`. It uses the mock LLM, so it needs
no provider credentials and therefore also runs on fork PRs, which skip every
secret-gated test-api job today.
Packaging:
* `overrides-freethreaded.txt` drops orjson. PEP 508 has no marker for
"free-threaded build", so a never-true marker is the mechanism; the alternative
is forking pyproject.toml for one interpreter.
* `scripts/ci/build-freethreaded-quicktok.sh` builds quicktok with
pybind11>=2.13 and `py::mod_gil_not_used()`. Building the published sdist
unmodified is not enough: the extension would not declare free-threading
support and would re-enable the GIL on import. Delete this step once the
change is upstream and released.
local-ml is excluded from the job. Importing `sentence_transformers` re-enables
the GIL, so a process that loads the local models cannot stay free-threaded
(torch, tokenizers, safetensors and transformers are each fine on their own --
measured, not assumed). `tests/conftest.py` now skips the `embeddings` and
`cross_encoder` fixtures when that stack is absent, instead of collapsing every
DB-backed test into a misleading "sentence-transformers is required for
LocalSTEmbeddings" ImportError. On 3.11, where local-ml is installed, nothing
changes.
Measured on python3.14t (Linux/aarch64, Postgres 18 + pgvector, mock LLM): the
non-LLM suite is 5954 passed / 1719 skipped. Of the residue, the migration-test
errors are this harness missing the `embedded-db` extra (pg0), which the CI job
installs; the remaining ~28 failures are not yet attributed to the interpreter
and need a same-container 3.11 control run before any are called real.
NOT yet validated: the job's `uv pip install --group dev` and quicktok script
invocation exactly as written. The quicktok patch-and-build was verified by hand
(producing a cp314t wheel that imports GIL-free and tokenizes correctly), but the
container host died mid-run before the scripted forms were exercised end to end.
Refs: initiative kp-ac23cdb434a144bf947371b6e6e4f5e8
* ci(api): drop the quicktok wheel build from the free-threaded job
quicktok had no free-threaded wheel, so the free-threaded CI job patched and built
one (pybind11>=2.13 plus py::mod_gil_not_used()) before it could run anything.
#4022 moved token counting to toktok-rs, which publishes cp3XXt wheels and declares
gil_used = false, so none of that is needed: a plain `uv pip install` from PyPI now
yields a working free-threaded install. Removes the build step and
scripts/ci/build-freethreaded-quicktok.sh.
overrides-freethreaded.txt keeps only orjson, which still publishes no cp3XXt wheel
and whose build script refuses free-threading outright.
* fix(api): make the shared caches and budget manager loop-agnostic
Two process-wide singletons held an `asyncio.Lock`. An asyncio.Lock binds to the
loop that first *waits* on it, so with several event loops in one process
(free-threaded uvicorn) the first contended acquire claims it and every other loop
then fails with
RuntimeError: <asyncio.locks.Lock object ...> is bound to a different event loop
It passes a single request and collapses under load, which is the worst possible
shape: recall returned HTTP 500 from every loop but one.
* engine/bank_stats_cache.py — the TTL cache behind bank config resolution.
* engine/db_budget.py — ConnectionBudgetManager, a `_default_manager` singleton.
Both now use a threading.Lock. That is not a workaround: every critical section
either guards is await-free dict work, so the lock is never held across a
suspension point and cannot block a loop, and unlike asyncio.Lock it is
loop-agnostic.
The cache needed one thing more. Its in-flight map coalesces concurrent loads
behind an asyncio.Future, and a Future belongs to the loop that created it, so a
caller on another loop must never await it. In-flight slots are now keyed by
(running loop, cache key): coalescing happens within a loop, while the cached DATA
stays shared across all of them, which is the part worth having. `invalidate()`
detaches the key on every loop, since the slots are per-loop.
Measured on python3.14t, real recall against Postgres 18 + pgvector, 8 event loops
in one process, 64 concurrent:
before 0 successful requests, 57 cross-loop errors
after 124-165 rps at 5.6-5.9 cores, p50 354-438ms, 0 cross-loop errors
For reference the same workload on 3.11 (one loop, as uvicorn runs today) does
~40 rps at 0.96 cores with p50 ~1390ms.
Verified on 3.11 and python3.14t: the bank stats/info cache suites (12 passed,
13 skipped) pass identically on both.
Refs: initiative kp-ac23cdb434a144bf947371b6e6e4f5e8
* fix(api): replace the remaining process-wide asyncio primitives
Four module- or class-level `asyncio` primitives were left after the cache and
budget-manager fixes. Each binds to the loop that first waits on it, so with
several event loops in one process the first contended acquire claims it and every
other loop fails with "is bound to a different event loop" — under load only, which
is why none of them showed up in tests:
* llm_wrapper `_global_llm_semaphore` and `_per_op_llm_semaphores` (module scope,
built at import before any loop exists)
* cross_encoder `RemoteTEICrossEncoder._global_semaphore` (class attribute)
* llamacpp `_shared_server_lock` (module scope)
Adds `_cross_loop.py` with `CrossLoopSemaphore` / `CrossLoopLock`. The counter lives
in a threading primitive, which is loop-agnostic, and waiting is a short async
backoff so a loop is never blocked while it queues — unlike a bare threading.Lock,
these are safe to hold across `await`, which llamacpp's server start/stop needs.
The caps stay PROCESS-wide rather than becoming per-loop. That preserves the
existing contract: `--workers N` has always meant N independent caps, one per
process, so `HINDSIGHT_API_LLM_MAX_CONCURRENT=8` keeps meaning 8 in flight per
process instead of silently becoming 8 x loops against the provider.
Polling rather than a cross-loop future handoff is deliberate and documented: it
only runs while a cap is saturated, costs at most 20ms of extra latency acquiring a
slot, and carries none of the per-loop waiter-registry state that
`call_soon_threadsafe` would need. These gate LLM calls and subprocess startup, so
that is not measurable. The uncontended path does not yield at all — a test pins
that, since it is on every LLM call.
Also documents the whole class of bug in the code-review skill: a "Concurrency"
standards section (which lock, and why the choice is ownership rather than style)
and review step 11d with the greps to catch it.
tests/test_cross_loop_primitives.py covers cross-loop use, that the cap really is
process-wide, exclusivity held across an await, and the uncontended fast path. It
also pins the failure mode of a plain asyncio.Semaphore, so the reason this module
exists cannot quietly stop applying. These tests need no free-threaded build — two
event loops in one process reproduce it on 3.11.
* fix(api): expose CrossLoopSemaphore's cap instead of a private counter
test_llm_per_op_concurrency asserted the configured cap by reading
asyncio.Semaphore's private `_value`, so it broke when the per-operation caps
became CrossLoopSemaphores.
Adds a public `capacity` property and asserts on that. Reading the cap is a
reasonable thing for a caller to want; making it public is better than swapping one
private attribute for another.
Caught by re-running the full 3.11 suite against the final tree — the earlier 3.11
control predated this commit, so it would otherwise have shipped as an unnoticed
regression on the supported interpreter.
* build(docker): add a free-threaded image target (tag -py3.14t)
Adds `api-builder-freethreaded` and `api-only-freethreaded`:
docker build --target api-only-freethreaded -t hindsight:py3.14t .
Built as its own pair of stages rather than by parameterising the existing ones.
Almost nothing is shared — the interpreter has to be installed rather than taken
from the base image, the dependency resolution differs, and local ML is excluded —
so parameterising would have complicated the supported 3.11 path to no benefit.
Nothing above these stages changes.
Notes on the shape, each of which cost a build to find:
* There is no official free-threaded python image; the library/python tags ship the
GIL build only. uv installs the interpreter into /opt/pythons and the runtime
stage copies it alongside the venv.
* Not `uv sync --locked`. The resolution must drop the dependencies with no cp3XXt
wheel (overrides-freethreaded.txt) and `uv sync` takes no --override, so it
resolves fresh against the same pyproject. That is precisely why this ships as
its own tag instead of being assumed equivalent to the pinned image.
* No `uv pip check` either: it fails by design here, because pyproject still
declares orjson and quicktok-v1 for every other interpreter and pip check cannot
know their absence is deliberate.
* libpq-dev/libpq5 are needed because psycopg2-binary has no cp3XXt wheel and
builds from source. Shipping it costs the image nothing: migrations.py runs
alembic in a subprocess on a free-threaded build, so psycopg2 is never imported
into the serving process.
* local-ml is refused outright with an explicit error rather than silently
producing a mis-tagged image, since importing sentence_transformers re-enables
the GIL. This tag defaults to remote embeddings and reranking.
The build asserts what the tag claims: it imports the whole API under
PYTHONWARNINGS="error:The global interpreter lock:RuntimeWarning" and checks
`sys._is_gil_enabled()` is False. A C extension that has not declared
Py_MOD_GIL_NOT_USED re-enables the GIL on import and says so only with a warning,
so without this an image could look free-threaded and run single-threaded. The
runtime also sets HINDSIGHT_API_FREE_THREADING=strict so the container refuses to
start in that state rather than being merely slow.
Verified: image builds (1.94GB), starts, runs its startup migrations with the GIL
still disabled, and serves a real recall (142 results, matching the 3.11 image on
the same corpus) with zero GIL warnings in its log.
* test(api): guard multi-loop safety on the ordinary 3.11 suite
Every cross-loop bug found while bringing up the free-threaded server — the bank
stats cache, the connection budget manager, the LLM concurrency caps, and
dateparser's locale dictionaries — was found by running the server with eight event
loops, and none of them needed a free-threaded interpreter to reproduce.
asyncio.Lock/Semaphore/Future bind to the loop that first waits on them regardless
of the GIL, so two loops in two threads reproduce the whole class on 3.11. This adds
that as a normal test: it runs everywhere, in seconds, with no special build.
That is what keeps the free-threaded CI job from having to be the only safety net.
The cheap guard catches loop-binding on every PR; the expensive job is left to cover
what genuinely needs the interpreter — a C extension silently re-enabling the GIL,
and races that only appear under true parallelism.
Each case corresponds to a bug that shipped. Verified as a negative control by
restoring the pre-fix bank_stats_cache: on 3.11 the cache test fails with
"got Future ... attached to a different loop", exactly as the eight-loop server did.
The final case pins the premise itself — that a module-level asyncio primitive still
breaks across loops — so if CPython ever changes that, the guards above get revisited
rather than quietly becoming theatre.
* docs(code-review): state that both 3.11 and free-threaded 3.14 are supported
The Concurrency section explained which lock to use and why, but never said why a
reviewer should care — so the rules read as advice about a hypothetical future
interpreter rather than a property of the two builds the project actually ships.
Adds a "Supported interpreters" section naming both: CPython 3.11 (the default image,
the `.python-version` pin, what `uv.lock` resolves for) and free-threaded CPython 3.14
(the `-py3.14t` image target). Neither can be deferred to a follow-up.
It calls out the two things that catch people. Anything process-wide is genuinely
concurrent on 3.14t, because the GIL is no longer making check-then-act accidentally
atomic. And free-threading is lost SILENTLY: importing a C extension without
`Py_MOD_GIL_NOT_USED` re-enables the GIL for the whole process with only a
RuntimeWarning, so the 3.14t image keeps working and merely performs like the 3.11
one. "It passed CI" is therefore weaker evidence than usual, which is why the
free-threaded job asserts the GIL is off before running any test and the image build
asserts it too.
Also notes the thing that makes this tractable to review: most of what breaks is not
free-threading-specific but multi-loop, and multi-loop reproduces on 3.11 as soon as
two event loops exist in one process. So new shared state is expected to be covered by
tests/test_multi_loop_conformance.py in the ordinary suite, not left to the
free-threaded job.
Extends review step 11d and the must-fix list with the interpreter-dropping cases —
chiefly a new C-extension dependency with no cp3XXt wheel on the API import path.
* ci(api): drop quicktok from the free-threaded overrides, record orjson's cost
quicktok left the project in #4022, so overriding it out is dead weight.
Records what the remaining override actually costs, measured on a 1024-dim
embedding rather than assumed:
with orjson 23 us/vector, literal 6,531 chars
fallback (_repr_literal) 325 us/vector, literal 20,504 chars
14x slower to render and 3.1x more bytes on the wire per vector. Both spellings
parse to the same float32 bytes, so this costs throughput and nothing else, and it
is on the retain path only (memories/pg/writes.py, retain/link_utils.py) — recall
never renders a vector this way.
That is the number worth having before anyone decides whether orjson is worth
chasing upstream: it makes the free-threaded image a poor fit for a retain-heavy
deployment and a fine one for a recall-heavy deployment, which is exactly the
workload the free-threading work was aimed at.
* fix(api): turn the free-threading guard off inside the migration child
The migration subprocess imports psycopg2 deliberately — that is the entire reason
it exists. But the guard this branch adds is inherited by the child, defaults to
strict on a free-threaded build, and therefore turns psycopg2's "the GIL has been
enabled" warning into an exception. Every migration failed:
RuntimeError: Migration subprocess failed (exit 1).
ERROR __main__: Failed to run database migrations: The global interpreter lock
(GIL) has been enabled to load module 'psycopg2._psycopg'...
which took out 1765 tests on 3.14t — every fixture that migrates a schema.
This is a stacking bug, not a bug in #4033: that PR sets ENV_MIGRATION_ISOLATION to
"never" in the child to stop it recursing, which is all it needs because the guard
does not exist there. The guard is this branch's, so disabling it in the child is
this branch's job too.
Also clears PYTHONWARNINGS for the child, for the same reason one step removed: the
free-threaded CI job runs the suite with that warning promoted to an error, and the
child must not inherit it.
Found by running the full suite on both interpreters after the rebase — the
free-threading-only failures went from 19 to 1765, which is what a broken shared
fixture looks like rather than a broken feature.
* test(api): patch the isolation seam directly, not the cached env var
The migration orchestration tests opted out of the subprocess by setting
HINDSIGHT_API_MIGRATION_ISOLATION=never. That never worked: the flag is read through
get_config(), whose result is cached in a module global, so setting the env var after
any earlier get_config() call has no effect.
It passed on 3.11 by accident — "auto" resolves to "never" there anyway, because the
interpreter is not free-threaded — and failed on 3.14t, where "auto" isolates and the
patched step functions were never reached.
Patches migrations._should_isolate_migrations instead, which says plainly what these
tests need: the fan-out has to happen in this process, because what they assert is the
call sequence and a subprocess would not see the patches.
Worth folding into #4033: its tests are green on 3.11 for the same accidental reason,
so the flag's opt-out is not actually exercised there.
* build(docker): give the free-threaded image its own Dockerfile
The free-threaded stages lived in docker/standalone/Dockerfile. They shared nothing
with it that mattered: a different interpreter (installed rather than taken from the
base image), a different dependency resolution, no control plane, no local ML. The
only thing genuinely in common is start-all.sh, and the runtime hardening was
duplicated rather than reused anyway — so "reuse" was buying nothing while making a
550-line file longer and threading a build arg through four stages of the supported
3.11 image.
Moves them to docker/standalone/Dockerfile.freethreaded. docker/standalone/Dockerfile
is now byte-for-byte what it was before this branch.
docker build -f docker/standalone/Dockerfile.freethreaded -t hindsight-api:py3.14t .
Also drops overrides-freethreaded.txt entirely. It existed to remove dependencies with
no cp3XXt wheel: quicktok left in #4022 and orjson in #4040, and everything remaining
publishes free-threaded wheels, so there is nothing left to override. The CI job
installs plainly now too.
DEPENDS ON #4040. Until that merges, orjson is still in the runtime closure and the
free-threaded install fails on it — verified, that is exactly what the build does
without the override this commit removes.
* ci(docker): build, smoke test and release the free-threaded image
The `-py3.14t` image existed but nothing built it outside my machine, nothing
exercised it, and the release never published it. Closes all three.
CI (test-api free-threaded job) now builds the image and runs
docker/freethreaded-smoke.sh against it. The build already asserts the GIL is off
after importing the whole API, so a mis-tagged image fails before the smoke test
starts; the smoke test then covers what a build cannot:
* the container reaches /health/ready — i.e. it ran its migrations, which on this
image means the subprocess path, since psycopg2 would otherwise take the GIL for
the life of the process;
* it serves a real retain and recall;
* the SERVER process still has the GIL disabled, asserted rather than inferred from
the container working, because losing it is silent.
The smoke test is separate from test-image.sh rather than a flag on it: this image
ships no local models, so embeddings must be remote and test-image.sh assumes a
provider API key. It carries a deterministic stub embedder inline so it needs no
secrets and no network, which also means it runs on fork PRs — the free-threaded job
is deliberately not gated on secrets.
Release: adds the tag to the docker matrix. Every entry now names its Dockerfile,
since the free-threaded image has its own. Two deliberate asymmetries:
* `latest` never points at `-py3.14t`. It is not a drop-in for the default tag —
no local models — so it must be asked for by name.
* linux/amd64 only. The build installs the interpreter and compiles psycopg2 from
source, so emulated arm64 is slow enough to be worth adding deliberately rather
than inheriting by default.
Also silences the GIL warning inside the migration child. The child is SUPPOSED to
take the GIL, and left visible the warning surfaces in a `-py3.14t` container's log
as "the global interpreter lock (GIL) has been enabled" — which reads exactly like
the image has silently lost its free-threading when it has not. The smoke test now
treats any occurrence in the log as a failure, which only works once the expected one
is gone.
Documents the tag and its constraints in installation.md (regenerated docs skill).
Verified locally end to end: image builds, and the smoke test passes — starts,
migrates, retains, recalls, `free-threaded: 3.14.7`, no GIL warnings.
* ci(api): fix what the free-threaded CI job actually caught
The job failed with 64 failures on its first real run. My local container runs had
missed all of them, for two reasons worth recording: I never ran the suite with
PYTHONWARNINGS set, and my local venvs were not built with --all-extras the way CI's
3.11 job is.
51 of the 64 were the job's own configuration. It ran the whole suite with the GIL
re-enable warning promoted to an error, but the suite imports LiteLLM on purpose to
test that provider, and LiteLLM pulls in fastuuid, which has no free-threaded build —
so ~50 tests failed for doing exactly what they are meant to do
("NameError: name 'fastuuid' is not defined"). The filter now applies only to the
step that imports the whole API to assert the GIL is off. That is where it belongs,
and the property is still asserted three more times: in
tests/test_free_threading.py (in a subprocess), in the image build, and in the image
smoke test.
5 were a real 3.14 incompatibility in test code: `asyncio.get_event_loop()` no longer
auto-creates a loop, so `get_event_loop().run_until_complete(...)` raises
"There is no current event loop in thread 'MainThread'". Replaced with `asyncio.run`,
which is the supported spelling and behaves identically on 3.11.
2 were tests that need the local-ml extra, which a free-threaded install cannot have
(sentence-transformers re-enables the GIL) and CI's 3.11 job does have via
--all-extras. Both now `importorskip` the thing they actually need — torch's global
default dtype has nothing to assert without torch, and the reflect test's
MemoryEngine construction reaches the local embeddings provider.
The rest are pre-existing or already attributed: the xai_oauth cleanup test fails
with PYTHON_GIL=1 on the same binary, so it is a 3.14 asyncio change rather than a
free-threading one.
* ci(api): close the last four free-threaded CI failures
Down from 64 to 4 after the previous commit; these are the remainder.
Two were the local-ml pattern again. tests/test_jina_mlx_import_error.py stubs mlx,
but the path under test still reaches transformers for a tokenizer — so without the
extra the assertion sees "No module named 'transformers'" instead of the message it
checks. It now importorskips transformers, which is what it actually needs.
One was a missing environment variable rather than a code problem: the
github-copilot provider looks for CLI account metadata that no runner has, and falls
back to a token. The 3.11 test-api job passes GITHUB_TOKEN and this job did not.
Added. It is not a repository secret — Actions provides it to every run, forks
included — so the job stays runnable on fork PRs, which is deliberate given the
secret-gated test-api jobs skip there entirely.
The last is tests/test_xai_oauth_llm.py::test_cleanup_closes_a_client_still_draining
_from_a_recycle, which fails with PYTHON_GIL=1 on the same binary. It is a 3.14
asyncio scheduling change that a plain 3.14 upgrade would hit identically, not
something free-threading introduces, and it is left unfixed and attributed rather
than worked around.
* fix(xai-oauth): close a retired client whose drain task was cancelled
`cleanup()` cancels each in-flight drain task so shutdown does not block on a request
that may never land, and left the actual close to `_close_when_drained`'s `finally`.
From Python 3.12 that no longer works: the cancellation is delivered at the task's
next await — which IS the `await stale.aclose()` in that `finally` — so the close
never runs, and CancelledError is a BaseException, so the `suppress(Exception)`
around it does not catch it either.
The client was therefore never closed and leaked its connections on every shutdown
that happened while a recycle was still draining.
`cleanup()` now closes the retired clients itself, after the drain tasks are done.
The list is captured before cancelling, because the `finally` pops each entry out of
`_drained` on its way through whether or not the close happened. aclose() is
idempotent, so a drain that completed normally costs nothing.
Found by the free-threaded CI job, but it is NOT a free-threading bug: it reproduces
identically with PYTHON_GIL=1 on the same interpreter, so a plain 3.14 upgrade would
hit it too. 3.11 is unaffected — the older cancellation semantics let that `finally`
await run.
Two earlier theories were wrong and are recorded so nobody retries them: swallowing
the CancelledError is not enough (the next await is cancelled again), and
`Task.uncancel()` does not help either (`_must_cancel` still fires at the next await).
The close has to happen outside the cancelled task.
Adds a regression test asserting the drain task is gone AND the client is closed.
Verified as a negative control: with the fix reverted it fails on 3.14t with
"a retired client was left open after cleanup", and passes on 3.11 either way, which
is exactly the interpreter split the bug has.
* test(retain): stop asserting batch dispatch ORDER in the coalescer test
test_batches_never_exceed_the_backend_batch_size compared the flattened backend
calls to the input list, which pins the order in which batches reach the backend.
The coalescer never promised that: it runs up to `max_concurrent_requests` calls at
a time (`self._slots`), so which batch lands first is a scheduling detail.
Under the GIL the interleaving happened to be stable, so the assertion held. On a
free-threaded interpreter the batches genuinely race and it failed on order alone —
every text present, every batch within budget, every caller's vectors correct:
At index 8 diff: 'chunk-16' != 'chunk-8'
Compares a Counter instead, which keeps the property that actually matters — every
text dispatched exactly once, nothing dropped or duplicated — and leaves the
per-caller assertion below it untouched, since that is what proves each caller gets
its own vectors in its own position.
This is a test that was over-specified, not an implementation that regressed. The
batch-size assertion the test is named for is unchanged.
Verified 8/8 under xdist on 3.14t, where it was failing intermittently, and on 3.11.
* build(docker): install Rust in the free-threaded builder, for litellm
The image build failed on amd64:
Failed to build `litellm==1.99.0`
Error: command ['maturin', 'pep517', 'build-wheel', ...]
Caused by: No such file or directory (os error 2)
litellm publishes only abi3 wheels (cp310-abi3), and abi3 — the stable ABI — does not
apply to a free-threaded interpreter, so uv cannot use them and falls back to the
sdist, which builds litellm's Rust extension. litellm is a hard runtime dependency
(pyproject pins it per-platform), so this is not optional.
It is the same constraint that kept toktok from working free-threaded before #4022:
an abi3 wheel is invisible to a cp3XXt interpreter.
Rust is confined to the builder stage — the runtime image copies only the venv and
carries no toolchain. Drop this once litellm publishes cp3XXt wheels.
Only amd64 hit it: my local builds were linux/arm64, where the resolution differed.
That is a good argument for the CI job building the image at all, which is what
caught this.
20 KiB
Installation
Hindsight can be deployed in several ways depending on your infrastructure and requirements.
:::tip Don't want to manage infrastructure? Hindsight Cloud is a fully managed service that handles all infrastructure, scaling, and maintenance — sign up here. :::
Supported Platforms
Hindsight runs on Linux, macOS, and Windows:
| Platform | Docker | Bare Metal (pip) | Embedded DB (pg0) | Notes |
|---|---|---|---|---|
| Linux (x86_64, ARM64) | ✅ | ✅ | ✅ | Fully supported, recommended for production |
| macOS (Apple Silicon / arm64) | ✅ | ✅ | ✅ | Fully supported |
| macOS (Intel / x86_64) | ✅ | ⚠️ slim only | ✅ | Use hindsight-all-slim / hindsight-api-slim. The full bundle's local ML models (PyTorch, MLX) publish no Intel-Mac wheels, so pip install hindsight-all silently backtracks to a months-old release. Pair the slim bundle with a hosted embeddings/reranker provider or the in-process ONNX backend (hindsight-api-slim[local-onnx]). |
| Windows (x86_64) | ✅ | ✅ | ✅ | Fully supported — see Windows setup for external PostgreSQL option |
All platforms support the embedded database (pg0) for development. On Windows, you can also use an external PostgreSQL installation — see the Windows section for a step-by-step guide.
Prerequisites
PostgreSQL
Hindsight requires PostgreSQL 14+ with a vector extension for similarity search. The supported extensions are:
- pgvector (default)
- pgvectorscale
- vchord
- scann (AlloyDB)
Configure which one to use with HINDSIGHT_API_VECTOR_EXTENSION. See Configuration for details.
By default, Hindsight uses pg0 — an embedded PostgreSQL that runs locally on your machine. This is convenient for development but not recommended for production.
For production, use an external PostgreSQL with one of the supported vector extensions:
- Supabase — Managed PostgreSQL with pgvector built-in
- Neon — Serverless PostgreSQL with pgvector
- Azure Database for PostgreSQL — With pgvector and pgvectorscale support
- Google AlloyDB / AlloyDB Omni — With pgvector and ScaNN support
- AWS RDS / Cloud SQL — With pgvector extension enabled
- Self-hosted — PostgreSQL 14+ with your preferred vector extension
LLM Provider
You need an LLM API key for fact extraction, entity resolution, and answer generation. See Models for supported providers, model recommendations, and configuration.
Hardware
Hindsight is designed to run on commodity hardware. The footprint depends mainly on whether the full image (which bundles local embedding and reranker models) or the slim image (which delegates those to external providers) is used.
| Component | Minimum RAM | Recommended RAM | Notes |
|---|---|---|---|
| API — Full image | 1.5 GB | 2 GB | Loads local BGE embedder (~130 MB) and MiniLM cross-encoder (~90 MB) into memory, plus PyTorch/ONNX runtime arenas. Idle RSS settles around 0.8–1.0 GB; expect 1.2–1.5 GB under load. |
| API — Slim image | 512 MB | 1 GB | No local models. Steady-state RSS is dominated by Python runtime and DB connections. Requires external embedding and reranker providers (e.g. TEI, OpenAI, Cohere). |
| Control Plane (UI) | 128 MB | 256 MB | Next.js process, lightweight. |
| Worker (if separated) | Same as API image variant | Same as API image variant | Workers load the same models as the API server. |
| PostgreSQL | 512 MB | 1 GB+ | Scales with the number of memories and indexes. |
:::tip Reducing the footprint The bulk of the full image's memory comes from the bundled embedding and reranker models and their PyTorch/ONNX runtimes. To shrink the deployment to a few hundred MB of RAM, switch to the slim image and configure external embedding and reranker providers. :::
CPU vs GPU: 2 vCPUs on CPU-only is fine for development and basic workloads. For production traffic, the local reranker (cross-encoder) is the main bottleneck and typically benefits from a GPU to keep recall latency reasonable; alternatively, offload reranking to an external reranker provider (e.g. TEI, Cohere) on dedicated GPU hardware.
Docker
Best for: Quick start, development, small deployments
Run everything in one container with embedded PostgreSQL:
export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight --restart unless-stopped --shm-size=1g -p 8888:8888 -p 9999:9999 \
-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
-v hindsight-data:/home/hindsight/.pg0 \
ghcr.io/vectorize-io/hindsight:latest
- API Server: http://localhost:8888
- Control Plane (Web UI): http://localhost:9999
:::note Persisting data: named volume vs. host bind mount
The container runs as a non-root user (UID 1000). The hindsight-data named volume above is recommended — Docker creates it owned by the container user, so it works with no extra setup.
If you instead bind-mount a host directory (-v $HOME/.hindsight-docker:/home/hindsight/.pg0), that directory must be owned by UID 1000, or the embedded database fails to start with Permission denied:
sudo chown -R 1000:1000 $HOME/.hindsight-docker
Do not run the image under a different UID with --user to work around this. The image only defines the hindsight user (UID 1000), so any other UID has no /etc/passwd entry, and startup crashes in a library that looks the running user up by ID:
KeyError: 'getpwuid(): uid not found: 1042'
Running as UID 1000 (the default) with the directory chowned to match is the supported way to use a host bind mount. If you cannot chown the directory — for example on a NAS share mounted with a fixed uid= — use a named volume instead.
:::
All published images are signed with Cosign — verification is optional.
:::tip Set a stable HINDSIGHT_API_WORKER_ID in production
The worker uses the container hostname as its identity, which Docker sets to the container ID by default. That value changes on every restart, so any task that was being processed when the container went down stays parked under the old ID with no way for the new container to recognize it as its own.
Set HINDSIGHT_API_WORKER_ID to a stable value (e.g., -e HINDSIGHT_API_WORKER_ID=hindsight-prod) so the worker keeps the same identity across restarts. This is recommended even for single-container deployments. For diagnosis and recovery commands, see Admin CLI - Recovering stuck operations.
:::
Docker Image Variants
| Variant | Size (AMD64) | Size (ARM64) | When to use |
|---|---|---|---|
Full (latest) |
~9 GB | ~3.7 GB | Default. Embeddings and reranking run in the image; the LLM is always external, including for local inference. |
Slim (slim) |
~500 MB | ~500 MB | Use when you already rely on external services for embeddings and reranking (OpenAI, Cohere, TEI). Significantly smaller image, faster deploys. Requires external providers. |
The slim image corresponds to the hindsight-api-slim pip package. See Configuration for external provider options.
Neither image bundles llama.cpp, so the built-in llamacpp provider is not available in Docker. To run inference locally, start llama.cpp (or Ollama, LM Studio, vLLM) alongside Hindsight and point HINDSIGHT_API_LLM_BASE_URL at it — see docker/docker-compose/local-llm/ for a working compose file.
NVIDIA GPU Acceleration (CUDA)
To run in-process local embedding and reranker models on an NVIDIA GPU, build a custom CUDA-enabled image using the recipe in docker/docker-compose/cuda/:
docker compose -f docker/docker-compose/cuda/docker-compose.yaml up --build
The recipe upgrades in-process PyTorch to CUDA 12.6 and configures GPU passthrough via the NVIDIA Container Toolkit. Build it on the machine that will run it — an emulated cross-architecture image cannot reach the GPU.
The CUDA runtime is additive: the base image's CPU PyTorch wheel stays in the lower layers, so the result is roughly 11 GB on disk against the ~9 GB Full image. This is why no CUDA image is published — the cost is only worth paying when you actually have a GPU to use.
See docker/docker-compose/cuda/README.md for prerequisites, manual build steps, and verification.
Bundling Custom Models in a Custom Image
:::tip Production deployments with non-default local models
If you use a non-default local embedder or reranker, bake the models into a custom image at build time rather than enabling the Helm modelCache PVC. See docker/docker-compose/custom-models/ for a runnable example.
:::
Available Tags
# Standalone (API + Control Plane)
ghcr.io/vectorize-io/hindsight:latest # Full, latest release
ghcr.io/vectorize-io/hindsight:latest-slim # Slim, latest release
ghcr.io/vectorize-io/hindsight:0.4.9 # Full, specific version
ghcr.io/vectorize-io/hindsight:0.4.9-slim # Slim, specific version
# API only
ghcr.io/vectorize-io/hindsight-api:latest
ghcr.io/vectorize-io/hindsight-api:latest-slim
# Control Plane only
ghcr.io/vectorize-io/hindsight-control-plane:latest
# API only, on free-threaded CPython 3.14 (see below)
ghcr.io/vectorize-io/hindsight-api:latest-py3.14t
The -py3.14t tag
Built on free-threaded CPython 3.14, where the process can execute Python bytecode in parallel instead of serialising it on one interpreter lock. It is aimed at recall-heavy deployments; the win comes from running several event loops in one process rather than from the interpreter alone.
It is not a drop-in replacement for the default tag:
- No local models. Importing
sentence-transformersre-enables the GIL, so this image cannot ship it. Embeddings and reranking must be remote — setHINDSIGHT_API_EMBEDDINGS_PROVIDERandHINDSIGHT_API_RERANKER_PROVIDERto a remote provider (tei,openai,cohere, …). The image defaults both totei. linux/amd64only, where the other tags are also published forarm64.- Migrations run in a subprocess (
HINDSIGHT_API_MIGRATION_ISOLATION=auto), because alembic's psycopg2 would otherwise take the GIL for the life of the process. - The container refuses to start if free-threading has been lost
(
HINDSIGHT_API_FREE_THREADING=strict). That is deliberate: the failure is otherwise silent, and the only symptom would be performing like the default tag.
If you use local embedding or reranking models, stay on the default tags.
Verifying image signatures
Images are signed with Cosign keyless OIDC. To verify any tag:
cosign verify ghcr.io/vectorize-io/hindsight:<tag> \
--certificate-identity-regexp '^https://github\.com/vectorize-io/hindsight/\.github/workflows/(sign-images|release)\.yml@.*' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
Helm / Kubernetes
Best for: Production deployments, auto-scaling, cloud environments
# Install with built-in PostgreSQL
helm install hindsight oci://ghcr.io/vectorize-io/charts/hindsight \
--set api.llm.provider=groq \
--set api.llm.apiKey=gsk_xxxxxxxxxxxx \
--set postgresql.enabled=true
# Or use external PostgreSQL
helm install hindsight oci://ghcr.io/vectorize-io/charts/hindsight \
--set api.llm.provider=groq \
--set api.llm.apiKey=gsk_xxxxxxxxxxxx \
--set postgresql.enabled=false \
--set api.database.url=postgresql://user:pass@postgres.example.com:5432/hindsight
# Install a specific version
helm install hindsight oci://ghcr.io/vectorize-io/charts/hindsight --version 0.1.3
# Upgrade to latest
helm upgrade hindsight oci://ghcr.io/vectorize-io/charts/hindsight
Requirements:
- Kubernetes cluster (GKE, EKS, AKS, or self-hosted)
- Helm 3.8+
Distributed Workers
For high-throughput deployments, enable dedicated worker pods to scale task processing independently:
helm install hindsight oci://ghcr.io/vectorize-io/charts/hindsight \
--set worker.enabled=true \
--set worker.replicaCount=3
The chart deploys workers as a StatefulSet, so each pod gets a stable name (e.g. hindsight-worker-0) that the worker uses as its HINDSIGHT_API_WORKER_ID. Tasks claimed by a pod are recognized as its own across restarts. If you swap the chart for a plain Deployment, set HINDSIGHT_API_WORKER_ID explicitly per replica — otherwise hostnames are randomized and previously-claimed tasks become orphaned. See Admin CLI - Recovering stuck operations for diagnosis.
See Services - Worker Service for configuration details and architecture.
See the Helm chart values.yaml for all chart options.
Bare Metal (pip)
Best for: Running Hindsight as a standalone service on a host machine.
Install
pip install hindsight-api # Full — works out of the box
pip install hindsight-api-slim # Slim — requires external services for embeddings, reranking, and the database
When using hindsight-api-slim, you must configure external providers for all model operations. See Configuration for details.
Run with Embedded Database
For development and testing, Hindsight can run with an embedded PostgreSQL (pg0):
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx
hindsight-api
This creates a database in ~/.hindsight/data/ and starts the API on http://localhost:8888.
Run with External PostgreSQL
For production, connect to your own PostgreSQL instance:
export HINDSIGHT_API_DATABASE_URL=postgresql://user:pass@localhost:5432/hindsight
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx
hindsight-api
Note: The database must exist and have pgvector enabled (CREATE EXTENSION vector;).
CLI Options
hindsight-api --port 9000 # Custom port (default: 8888)
hindsight-api --host 127.0.0.1 # Bind to localhost only
hindsight-api --workers 4 # Multiple worker processes
hindsight-api --log-level debug # Verbose logging
Control Plane
The Control Plane (Web UI) can be run standalone using npx:
npx @vectorize-io/hindsight-control-plane --api-url http://localhost:8888
This connects to your running API server and provides a visual interface for managing memory banks, exploring entities, and testing queries.
Options
| Option | Environment Variable | Default | Description |
|---|---|---|---|
-p, --port |
PORT |
9999 | Port to listen on |
-H, --hostname |
HOSTNAME |
0.0.0.0 | Hostname to bind to |
-a, --api-url |
HINDSIGHT_CP_DATAPLANE_API_URL |
http://localhost:8888 | Hindsight API URL |
HINDSIGHT_CP_ACCESS_KEY |
(none) | Access key to protect the Control Plane UI. When set, users must enter this key to log in. |
Examples
# Run on custom port
npx @vectorize-io/hindsight-control-plane --port 9999 --api-url http://localhost:8888
# Using environment variables
export HINDSIGHT_CP_DATAPLANE_API_URL=http://api.example.com
npx @vectorize-io/hindsight-control-plane
# Production deployment
PORT=80 HINDSIGHT_CP_DATAPLANE_API_URL=https://api.hindsight.io npx @vectorize-io/hindsight-control-plane
Windows
Best for: Running Hindsight natively on Windows without Docker
Hindsight works on Windows with the embedded database (pg0) out of the box — just install and run:
pip install hindsight-api
set HINDSIGHT_API_LLM_PROVIDER=openai
set HINDSIGHT_API_LLM_API_KEY=sk-xxx
set HINDSIGHT_API_LLM_MODEL=gpt-4o-mini
hindsight-api
Using External PostgreSQL (optional)
If you prefer to use your own PostgreSQL instance instead of the embedded database:
# Install PostgreSQL
winget install PostgreSQL.PostgreSQL.17
# Build pgvector (requires Visual Studio Build Tools)
git clone https://github.com/pgvector/pgvector.git
cd pgvector
# Open "x64 Native Tools Command Prompt for VS" and run:
set PGROOT=C:\Program Files\PostgreSQL\17
nmake /F Makefile.win
nmake /F Makefile.win install
# Create the database and enable the vector extension
psql -U postgres -c "CREATE DATABASE hindsight;"
psql -U postgres -d hindsight -c "CREATE EXTENSION vector;"
Then run Hindsight pointing to your database:
pip install hindsight-api
set HINDSIGHT_API_DATABASE_URL=postgresql://postgres@localhost:5432/hindsight
set HINDSIGHT_API_LLM_PROVIDER=openai
set HINDSIGHT_API_LLM_API_KEY=sk-xxx
set HINDSIGHT_API_LLM_MODEL=gpt-4o-mini
hindsight-api
- API Server: http://localhost:8888
:::tip
You can also use the slim package (pip install hindsight-api-slim) if you configure external providers for embeddings and reranking. See Configuration for details.
:::
Windows + China Network Notes
If you are running on Windows behind China network restrictions:
- DeepSeek works well for
HINDSIGHT_API_LLM_PROVIDER, but DeepSeek does not provide an embeddings endpoint. - Use local embeddings (recommended for privacy and reliability in restricted networks).
- Set
HF_ENDPOINT=https://hf-mirror.combefore starting Hindsight so Hugging Face model downloads use a China-accessible mirror.
set HF_ENDPOINT=https://hf-mirror.com
set HINDSIGHT_API_LLM_PROVIDER=deepseek
set HINDSIGHT_API_LLM_API_KEY=sk-your-deepseek-key
set HINDSIGHT_API_LLM_MODEL=deepseek-v4-flash
set HINDSIGHT_API_LLM_BASE_URL=https://api.deepseek.com
set HINDSIGHT_API_EMBEDDINGS_PROVIDER=local
set HINDSIGHT_API_EMBEDDINGS_LOCAL_MODEL=BAAI/bge-small-en-v1.5
set HINDSIGHT_API_RERANKER_PROVIDER=flashrank
hindsight-api
The HF_ENDPOINT variable is used by Hugging Face tooling (huggingface_hub), not by Hindsight itself.
Embedded in a Python Application
Best for: Using Hindsight programmatically from Python without running a separate server process.
pip install hindsight-all # Full — works out of the box (Linux, Windows, Apple Silicon Macs)
pip install hindsight-all-slim # Slim — requires external services for embeddings, reranking, and the database
On Intel (x86_64) Macs, install hindsight-all-slim — see Supported Platforms.
hindsight-all supports two modes of embedding:
In-process (HindsightServer): the server runs in a background thread inside your application. Best when you want the tightest integration and are already managing your own process lifecycle.
from hindsight import HindsightServer, HindsightClient
with HindsightServer(llm_provider="openai", llm_api_key="sk-xxx") as server:
client = HindsightClient(base_url=server.url)
client.retain(bank_id="alice", content="Alice prefers concise answers.")
results = client.recall(bank_id="alice", query="How should I respond to Alice?")
Managed subprocess (HindsightEmbedded): the server runs as a background daemon process, shared across multiple Python processes or sessions. The daemon starts on first use and runs until it is stopped.
from hindsight import HindsightEmbedded
client = HindsightEmbedded(llm_provider="openai", llm_api_key="sk-xxx")
client.retain(bank_id="alice", content="Alice prefers concise answers.")
results = client.recall(bank_id="alice", query="How should I respond to Alice?")
See the Python SDK for the full API reference.
Next Steps
- Configuration — Environment variables and settings
- Models — ML models and providers
- Monitoring — Metrics and observability