mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
perf/profiling-env
36 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
614bfc96df |
feat: run Hindsight on free-threaded CPython 3.14 (-py3.14t image) (#4037)
* test(api): guard against silently losing free-threading
A free-threaded CPython re-enables the GIL the moment it imports a C extension
that has not declared `Py_MOD_GIL_NOT_USED`, and says so only with a
RuntimeWarning. One new module-scope import therefore reverts the whole server
to single-threaded execution while every existing test still passes and the
process still serves traffic -- which is exactly how the three imports fixed in
the previous commit went unnoticed.
Adds a subprocess probe that imports the full API surface and asserts the GIL is
still disabled, plus a second case pinning the diagnostic form:
PYTHONWARNINGS="error:The global interpreter lock:RuntimeWarning"
which turns the warning fatal at the offending import (naming the module and the
import chain) while leaving unrelated RuntimeWarnings alone.
The probe runs in a subprocess because the GIL can only be re-enabled once per
interpreter, so an import already done in the pytest parent would mask a
regression.
Skipped unless `Py_GIL_DISABLED`, so it is inert on the 3.11 matrix and only
bites on a python3.14t job.
Verified both ways on python3.14t: passes on the fixed tree, and a negative
control that adds `import psycopg2` to the probe fails with the module named in
the assertion message.
* feat(api): fail loudly when a free-threaded build loses the GIL
A free-threaded CPython re-enables the GIL for the whole process the moment it
imports a C extension that has not declared `Py_MOD_GIL_NOT_USED`, and says so only
with a RuntimeWarning. Nothing crashes and nothing degrades visibly: the server
starts, serves traffic, and passes its tests, having quietly reverted to
single-threaded execution. One new module-scope import is enough.
Adds `hindsight_api/_free_threading.py` and calls it eagerly from the package
`__init__`, next to `apply_default_thread_limits()` and for the same class of reason:
both configure how the process executes, and both are worthless once the libraries
they govern have loaded.
`HINDSIGHT_API_FREE_THREADING` selects the mode:
strict (default on a free-threaded build) — the GIL re-enable warning becomes an
exception, so the offending import raises with the module and full import
chain in the traceback. A GIL already on at startup raises.
warn — log and continue, for bringing up a deployment whose dependencies are not
all ready.
off — no guard; used by the migration subprocess, which imports psycopg2 on
purpose.
Every mode is a no-op on a normal build, so this is inert on 3.11.
* ci(api): run the test suite on free-threaded CPython 3.14
Adds a `test-api (free-threaded 3.14)` job plus the two pieces of packaging it
needs, so a regression that silently re-enables the GIL fails CI instead of
quietly costing the deployment its parallelism.
The job asserts free-threading before running anything -- a bad build or an
already-taken GIL shows up as its own red step rather than as a mass of confusing
downstream failures -- and runs pytest under
PYTHONWARNINGS="error:The global interpreter lock:RuntimeWarning"
so a regression fails at the offending import with the module named.
It is deliberately NOT gated on `has_secrets`. It uses the mock LLM, so it needs
no provider credentials and therefore also runs on fork PRs, which skip every
secret-gated test-api job today.
Packaging:
* `overrides-freethreaded.txt` drops orjson. PEP 508 has no marker for
"free-threaded build", so a never-true marker is the mechanism; the alternative
is forking pyproject.toml for one interpreter.
* `scripts/ci/build-freethreaded-quicktok.sh` builds quicktok with
pybind11>=2.13 and `py::mod_gil_not_used()`. Building the published sdist
unmodified is not enough: the extension would not declare free-threading
support and would re-enable the GIL on import. Delete this step once the
change is upstream and released.
local-ml is excluded from the job. Importing `sentence_transformers` re-enables
the GIL, so a process that loads the local models cannot stay free-threaded
(torch, tokenizers, safetensors and transformers are each fine on their own --
measured, not assumed). `tests/conftest.py` now skips the `embeddings` and
`cross_encoder` fixtures when that stack is absent, instead of collapsing every
DB-backed test into a misleading "sentence-transformers is required for
LocalSTEmbeddings" ImportError. On 3.11, where local-ml is installed, nothing
changes.
Measured on python3.14t (Linux/aarch64, Postgres 18 + pgvector, mock LLM): the
non-LLM suite is 5954 passed / 1719 skipped. Of the residue, the migration-test
errors are this harness missing the `embedded-db` extra (pg0), which the CI job
installs; the remaining ~28 failures are not yet attributed to the interpreter
and need a same-container 3.11 control run before any are called real.
NOT yet validated: the job's `uv pip install --group dev` and quicktok script
invocation exactly as written. The quicktok patch-and-build was verified by hand
(producing a cp314t wheel that imports GIL-free and tokenizes correctly), but the
container host died mid-run before the scripted forms were exercised end to end.
Refs: initiative kp-ac23cdb434a144bf947371b6e6e4f5e8
* ci(api): drop the quicktok wheel build from the free-threaded job
quicktok had no free-threaded wheel, so the free-threaded CI job patched and built
one (pybind11>=2.13 plus py::mod_gil_not_used()) before it could run anything.
#4022 moved token counting to toktok-rs, which publishes cp3XXt wheels and declares
gil_used = false, so none of that is needed: a plain `uv pip install` from PyPI now
yields a working free-threaded install. Removes the build step and
scripts/ci/build-freethreaded-quicktok.sh.
overrides-freethreaded.txt keeps only orjson, which still publishes no cp3XXt wheel
and whose build script refuses free-threading outright.
* fix(api): make the shared caches and budget manager loop-agnostic
Two process-wide singletons held an `asyncio.Lock`. An asyncio.Lock binds to the
loop that first *waits* on it, so with several event loops in one process
(free-threaded uvicorn) the first contended acquire claims it and every other loop
then fails with
RuntimeError: <asyncio.locks.Lock object ...> is bound to a different event loop
It passes a single request and collapses under load, which is the worst possible
shape: recall returned HTTP 500 from every loop but one.
* engine/bank_stats_cache.py — the TTL cache behind bank config resolution.
* engine/db_budget.py — ConnectionBudgetManager, a `_default_manager` singleton.
Both now use a threading.Lock. That is not a workaround: every critical section
either guards is await-free dict work, so the lock is never held across a
suspension point and cannot block a loop, and unlike asyncio.Lock it is
loop-agnostic.
The cache needed one thing more. Its in-flight map coalesces concurrent loads
behind an asyncio.Future, and a Future belongs to the loop that created it, so a
caller on another loop must never await it. In-flight slots are now keyed by
(running loop, cache key): coalescing happens within a loop, while the cached DATA
stays shared across all of them, which is the part worth having. `invalidate()`
detaches the key on every loop, since the slots are per-loop.
Measured on python3.14t, real recall against Postgres 18 + pgvector, 8 event loops
in one process, 64 concurrent:
before 0 successful requests, 57 cross-loop errors
after 124-165 rps at 5.6-5.9 cores, p50 354-438ms, 0 cross-loop errors
For reference the same workload on 3.11 (one loop, as uvicorn runs today) does
~40 rps at 0.96 cores with p50 ~1390ms.
Verified on 3.11 and python3.14t: the bank stats/info cache suites (12 passed,
13 skipped) pass identically on both.
Refs: initiative kp-ac23cdb434a144bf947371b6e6e4f5e8
* fix(api): replace the remaining process-wide asyncio primitives
Four module- or class-level `asyncio` primitives were left after the cache and
budget-manager fixes. Each binds to the loop that first waits on it, so with
several event loops in one process the first contended acquire claims it and every
other loop fails with "is bound to a different event loop" — under load only, which
is why none of them showed up in tests:
* llm_wrapper `_global_llm_semaphore` and `_per_op_llm_semaphores` (module scope,
built at import before any loop exists)
* cross_encoder `RemoteTEICrossEncoder._global_semaphore` (class attribute)
* llamacpp `_shared_server_lock` (module scope)
Adds `_cross_loop.py` with `CrossLoopSemaphore` / `CrossLoopLock`. The counter lives
in a threading primitive, which is loop-agnostic, and waiting is a short async
backoff so a loop is never blocked while it queues — unlike a bare threading.Lock,
these are safe to hold across `await`, which llamacpp's server start/stop needs.
The caps stay PROCESS-wide rather than becoming per-loop. That preserves the
existing contract: `--workers N` has always meant N independent caps, one per
process, so `HINDSIGHT_API_LLM_MAX_CONCURRENT=8` keeps meaning 8 in flight per
process instead of silently becoming 8 x loops against the provider.
Polling rather than a cross-loop future handoff is deliberate and documented: it
only runs while a cap is saturated, costs at most 20ms of extra latency acquiring a
slot, and carries none of the per-loop waiter-registry state that
`call_soon_threadsafe` would need. These gate LLM calls and subprocess startup, so
that is not measurable. The uncontended path does not yield at all — a test pins
that, since it is on every LLM call.
Also documents the whole class of bug in the code-review skill: a "Concurrency"
standards section (which lock, and why the choice is ownership rather than style)
and review step 11d with the greps to catch it.
tests/test_cross_loop_primitives.py covers cross-loop use, that the cap really is
process-wide, exclusivity held across an await, and the uncontended fast path. It
also pins the failure mode of a plain asyncio.Semaphore, so the reason this module
exists cannot quietly stop applying. These tests need no free-threaded build — two
event loops in one process reproduce it on 3.11.
* fix(api): expose CrossLoopSemaphore's cap instead of a private counter
test_llm_per_op_concurrency asserted the configured cap by reading
asyncio.Semaphore's private `_value`, so it broke when the per-operation caps
became CrossLoopSemaphores.
Adds a public `capacity` property and asserts on that. Reading the cap is a
reasonable thing for a caller to want; making it public is better than swapping one
private attribute for another.
Caught by re-running the full 3.11 suite against the final tree — the earlier 3.11
control predated this commit, so it would otherwise have shipped as an unnoticed
regression on the supported interpreter.
* build(docker): add a free-threaded image target (tag -py3.14t)
Adds `api-builder-freethreaded` and `api-only-freethreaded`:
docker build --target api-only-freethreaded -t hindsight:py3.14t .
Built as its own pair of stages rather than by parameterising the existing ones.
Almost nothing is shared — the interpreter has to be installed rather than taken
from the base image, the dependency resolution differs, and local ML is excluded —
so parameterising would have complicated the supported 3.11 path to no benefit.
Nothing above these stages changes.
Notes on the shape, each of which cost a build to find:
* There is no official free-threaded python image; the library/python tags ship the
GIL build only. uv installs the interpreter into /opt/pythons and the runtime
stage copies it alongside the venv.
* Not `uv sync --locked`. The resolution must drop the dependencies with no cp3XXt
wheel (overrides-freethreaded.txt) and `uv sync` takes no --override, so it
resolves fresh against the same pyproject. That is precisely why this ships as
its own tag instead of being assumed equivalent to the pinned image.
* No `uv pip check` either: it fails by design here, because pyproject still
declares orjson and quicktok-v1 for every other interpreter and pip check cannot
know their absence is deliberate.
* libpq-dev/libpq5 are needed because psycopg2-binary has no cp3XXt wheel and
builds from source. Shipping it costs the image nothing: migrations.py runs
alembic in a subprocess on a free-threaded build, so psycopg2 is never imported
into the serving process.
* local-ml is refused outright with an explicit error rather than silently
producing a mis-tagged image, since importing sentence_transformers re-enables
the GIL. This tag defaults to remote embeddings and reranking.
The build asserts what the tag claims: it imports the whole API under
PYTHONWARNINGS="error:The global interpreter lock:RuntimeWarning" and checks
`sys._is_gil_enabled()` is False. A C extension that has not declared
Py_MOD_GIL_NOT_USED re-enables the GIL on import and says so only with a warning,
so without this an image could look free-threaded and run single-threaded. The
runtime also sets HINDSIGHT_API_FREE_THREADING=strict so the container refuses to
start in that state rather than being merely slow.
Verified: image builds (1.94GB), starts, runs its startup migrations with the GIL
still disabled, and serves a real recall (142 results, matching the 3.11 image on
the same corpus) with zero GIL warnings in its log.
* test(api): guard multi-loop safety on the ordinary 3.11 suite
Every cross-loop bug found while bringing up the free-threaded server — the bank
stats cache, the connection budget manager, the LLM concurrency caps, and
dateparser's locale dictionaries — was found by running the server with eight event
loops, and none of them needed a free-threaded interpreter to reproduce.
asyncio.Lock/Semaphore/Future bind to the loop that first waits on them regardless
of the GIL, so two loops in two threads reproduce the whole class on 3.11. This adds
that as a normal test: it runs everywhere, in seconds, with no special build.
That is what keeps the free-threaded CI job from having to be the only safety net.
The cheap guard catches loop-binding on every PR; the expensive job is left to cover
what genuinely needs the interpreter — a C extension silently re-enabling the GIL,
and races that only appear under true parallelism.
Each case corresponds to a bug that shipped. Verified as a negative control by
restoring the pre-fix bank_stats_cache: on 3.11 the cache test fails with
"got Future ... attached to a different loop", exactly as the eight-loop server did.
The final case pins the premise itself — that a module-level asyncio primitive still
breaks across loops — so if CPython ever changes that, the guards above get revisited
rather than quietly becoming theatre.
* docs(code-review): state that both 3.11 and free-threaded 3.14 are supported
The Concurrency section explained which lock to use and why, but never said why a
reviewer should care — so the rules read as advice about a hypothetical future
interpreter rather than a property of the two builds the project actually ships.
Adds a "Supported interpreters" section naming both: CPython 3.11 (the default image,
the `.python-version` pin, what `uv.lock` resolves for) and free-threaded CPython 3.14
(the `-py3.14t` image target). Neither can be deferred to a follow-up.
It calls out the two things that catch people. Anything process-wide is genuinely
concurrent on 3.14t, because the GIL is no longer making check-then-act accidentally
atomic. And free-threading is lost SILENTLY: importing a C extension without
`Py_MOD_GIL_NOT_USED` re-enables the GIL for the whole process with only a
RuntimeWarning, so the 3.14t image keeps working and merely performs like the 3.11
one. "It passed CI" is therefore weaker evidence than usual, which is why the
free-threaded job asserts the GIL is off before running any test and the image build
asserts it too.
Also notes the thing that makes this tractable to review: most of what breaks is not
free-threading-specific but multi-loop, and multi-loop reproduces on 3.11 as soon as
two event loops exist in one process. So new shared state is expected to be covered by
tests/test_multi_loop_conformance.py in the ordinary suite, not left to the
free-threaded job.
Extends review step 11d and the must-fix list with the interpreter-dropping cases —
chiefly a new C-extension dependency with no cp3XXt wheel on the API import path.
* ci(api): drop quicktok from the free-threaded overrides, record orjson's cost
quicktok left the project in #4022, so overriding it out is dead weight.
Records what the remaining override actually costs, measured on a 1024-dim
embedding rather than assumed:
with orjson 23 us/vector, literal 6,531 chars
fallback (_repr_literal) 325 us/vector, literal 20,504 chars
14x slower to render and 3.1x more bytes on the wire per vector. Both spellings
parse to the same float32 bytes, so this costs throughput and nothing else, and it
is on the retain path only (memories/pg/writes.py, retain/link_utils.py) — recall
never renders a vector this way.
That is the number worth having before anyone decides whether orjson is worth
chasing upstream: it makes the free-threaded image a poor fit for a retain-heavy
deployment and a fine one for a recall-heavy deployment, which is exactly the
workload the free-threading work was aimed at.
* fix(api): turn the free-threading guard off inside the migration child
The migration subprocess imports psycopg2 deliberately — that is the entire reason
it exists. But the guard this branch adds is inherited by the child, defaults to
strict on a free-threaded build, and therefore turns psycopg2's "the GIL has been
enabled" warning into an exception. Every migration failed:
RuntimeError: Migration subprocess failed (exit 1).
ERROR __main__: Failed to run database migrations: The global interpreter lock
(GIL) has been enabled to load module 'psycopg2._psycopg'...
which took out 1765 tests on 3.14t — every fixture that migrates a schema.
This is a stacking bug, not a bug in #4033: that PR sets ENV_MIGRATION_ISOLATION to
"never" in the child to stop it recursing, which is all it needs because the guard
does not exist there. The guard is this branch's, so disabling it in the child is
this branch's job too.
Also clears PYTHONWARNINGS for the child, for the same reason one step removed: the
free-threaded CI job runs the suite with that warning promoted to an error, and the
child must not inherit it.
Found by running the full suite on both interpreters after the rebase — the
free-threading-only failures went from 19 to 1765, which is what a broken shared
fixture looks like rather than a broken feature.
* test(api): patch the isolation seam directly, not the cached env var
The migration orchestration tests opted out of the subprocess by setting
HINDSIGHT_API_MIGRATION_ISOLATION=never. That never worked: the flag is read through
get_config(), whose result is cached in a module global, so setting the env var after
any earlier get_config() call has no effect.
It passed on 3.11 by accident — "auto" resolves to "never" there anyway, because the
interpreter is not free-threaded — and failed on 3.14t, where "auto" isolates and the
patched step functions were never reached.
Patches migrations._should_isolate_migrations instead, which says plainly what these
tests need: the fan-out has to happen in this process, because what they assert is the
call sequence and a subprocess would not see the patches.
Worth folding into #4033: its tests are green on 3.11 for the same accidental reason,
so the flag's opt-out is not actually exercised there.
* build(docker): give the free-threaded image its own Dockerfile
The free-threaded stages lived in docker/standalone/Dockerfile. They shared nothing
with it that mattered: a different interpreter (installed rather than taken from the
base image), a different dependency resolution, no control plane, no local ML. The
only thing genuinely in common is start-all.sh, and the runtime hardening was
duplicated rather than reused anyway — so "reuse" was buying nothing while making a
550-line file longer and threading a build arg through four stages of the supported
3.11 image.
Moves them to docker/standalone/Dockerfile.freethreaded. docker/standalone/Dockerfile
is now byte-for-byte what it was before this branch.
docker build -f docker/standalone/Dockerfile.freethreaded -t hindsight-api:py3.14t .
Also drops overrides-freethreaded.txt entirely. It existed to remove dependencies with
no cp3XXt wheel: quicktok left in #4022 and orjson in #4040, and everything remaining
publishes free-threaded wheels, so there is nothing left to override. The CI job
installs plainly now too.
DEPENDS ON #4040. Until that merges, orjson is still in the runtime closure and the
free-threaded install fails on it — verified, that is exactly what the build does
without the override this commit removes.
* ci(docker): build, smoke test and release the free-threaded image
The `-py3.14t` image existed but nothing built it outside my machine, nothing
exercised it, and the release never published it. Closes all three.
CI (test-api free-threaded job) now builds the image and runs
docker/freethreaded-smoke.sh against it. The build already asserts the GIL is off
after importing the whole API, so a mis-tagged image fails before the smoke test
starts; the smoke test then covers what a build cannot:
* the container reaches /health/ready — i.e. it ran its migrations, which on this
image means the subprocess path, since psycopg2 would otherwise take the GIL for
the life of the process;
* it serves a real retain and recall;
* the SERVER process still has the GIL disabled, asserted rather than inferred from
the container working, because losing it is silent.
The smoke test is separate from test-image.sh rather than a flag on it: this image
ships no local models, so embeddings must be remote and test-image.sh assumes a
provider API key. It carries a deterministic stub embedder inline so it needs no
secrets and no network, which also means it runs on fork PRs — the free-threaded job
is deliberately not gated on secrets.
Release: adds the tag to the docker matrix. Every entry now names its Dockerfile,
since the free-threaded image has its own. Two deliberate asymmetries:
* `latest` never points at `-py3.14t`. It is not a drop-in for the default tag —
no local models — so it must be asked for by name.
* linux/amd64 only. The build installs the interpreter and compiles psycopg2 from
source, so emulated arm64 is slow enough to be worth adding deliberately rather
than inheriting by default.
Also silences the GIL warning inside the migration child. The child is SUPPOSED to
take the GIL, and left visible the warning surfaces in a `-py3.14t` container's log
as "the global interpreter lock (GIL) has been enabled" — which reads exactly like
the image has silently lost its free-threading when it has not. The smoke test now
treats any occurrence in the log as a failure, which only works once the expected one
is gone.
Documents the tag and its constraints in installation.md (regenerated docs skill).
Verified locally end to end: image builds, and the smoke test passes — starts,
migrates, retains, recalls, `free-threaded: 3.14.7`, no GIL warnings.
* ci(api): fix what the free-threaded CI job actually caught
The job failed with 64 failures on its first real run. My local container runs had
missed all of them, for two reasons worth recording: I never ran the suite with
PYTHONWARNINGS set, and my local venvs were not built with --all-extras the way CI's
3.11 job is.
51 of the 64 were the job's own configuration. It ran the whole suite with the GIL
re-enable warning promoted to an error, but the suite imports LiteLLM on purpose to
test that provider, and LiteLLM pulls in fastuuid, which has no free-threaded build —
so ~50 tests failed for doing exactly what they are meant to do
("NameError: name 'fastuuid' is not defined"). The filter now applies only to the
step that imports the whole API to assert the GIL is off. That is where it belongs,
and the property is still asserted three more times: in
tests/test_free_threading.py (in a subprocess), in the image build, and in the image
smoke test.
5 were a real 3.14 incompatibility in test code: `asyncio.get_event_loop()` no longer
auto-creates a loop, so `get_event_loop().run_until_complete(...)` raises
"There is no current event loop in thread 'MainThread'". Replaced with `asyncio.run`,
which is the supported spelling and behaves identically on 3.11.
2 were tests that need the local-ml extra, which a free-threaded install cannot have
(sentence-transformers re-enables the GIL) and CI's 3.11 job does have via
--all-extras. Both now `importorskip` the thing they actually need — torch's global
default dtype has nothing to assert without torch, and the reflect test's
MemoryEngine construction reaches the local embeddings provider.
The rest are pre-existing or already attributed: the xai_oauth cleanup test fails
with PYTHON_GIL=1 on the same binary, so it is a 3.14 asyncio change rather than a
free-threading one.
* ci(api): close the last four free-threaded CI failures
Down from 64 to 4 after the previous commit; these are the remainder.
Two were the local-ml pattern again. tests/test_jina_mlx_import_error.py stubs mlx,
but the path under test still reaches transformers for a tokenizer — so without the
extra the assertion sees "No module named 'transformers'" instead of the message it
checks. It now importorskips transformers, which is what it actually needs.
One was a missing environment variable rather than a code problem: the
github-copilot provider looks for CLI account metadata that no runner has, and falls
back to a token. The 3.11 test-api job passes GITHUB_TOKEN and this job did not.
Added. It is not a repository secret — Actions provides it to every run, forks
included — so the job stays runnable on fork PRs, which is deliberate given the
secret-gated test-api jobs skip there entirely.
The last is tests/test_xai_oauth_llm.py::test_cleanup_closes_a_client_still_draining
_from_a_recycle, which fails with PYTHON_GIL=1 on the same binary. It is a 3.14
asyncio scheduling change that a plain 3.14 upgrade would hit identically, not
something free-threading introduces, and it is left unfixed and attributed rather
than worked around.
* fix(xai-oauth): close a retired client whose drain task was cancelled
`cleanup()` cancels each in-flight drain task so shutdown does not block on a request
that may never land, and left the actual close to `_close_when_drained`'s `finally`.
From Python 3.12 that no longer works: the cancellation is delivered at the task's
next await — which IS the `await stale.aclose()` in that `finally` — so the close
never runs, and CancelledError is a BaseException, so the `suppress(Exception)`
around it does not catch it either.
The client was therefore never closed and leaked its connections on every shutdown
that happened while a recycle was still draining.
`cleanup()` now closes the retired clients itself, after the drain tasks are done.
The list is captured before cancelling, because the `finally` pops each entry out of
`_drained` on its way through whether or not the close happened. aclose() is
idempotent, so a drain that completed normally costs nothing.
Found by the free-threaded CI job, but it is NOT a free-threading bug: it reproduces
identically with PYTHON_GIL=1 on the same interpreter, so a plain 3.14 upgrade would
hit it too. 3.11 is unaffected — the older cancellation semantics let that `finally`
await run.
Two earlier theories were wrong and are recorded so nobody retries them: swallowing
the CancelledError is not enough (the next await is cancelled again), and
`Task.uncancel()` does not help either (`_must_cancel` still fires at the next await).
The close has to happen outside the cancelled task.
Adds a regression test asserting the drain task is gone AND the client is closed.
Verified as a negative control: with the fix reverted it fails on 3.14t with
"a retired client was left open after cleanup", and passes on 3.11 either way, which
is exactly the interpreter split the bug has.
* test(retain): stop asserting batch dispatch ORDER in the coalescer test
test_batches_never_exceed_the_backend_batch_size compared the flattened backend
calls to the input list, which pins the order in which batches reach the backend.
The coalescer never promised that: it runs up to `max_concurrent_requests` calls at
a time (`self._slots`), so which batch lands first is a scheduling detail.
Under the GIL the interleaving happened to be stable, so the assertion held. On a
free-threaded interpreter the batches genuinely race and it failed on order alone —
every text present, every batch within budget, every caller's vectors correct:
At index 8 diff: 'chunk-16' != 'chunk-8'
Compares a Counter instead, which keeps the property that actually matters — every
text dispatched exactly once, nothing dropped or duplicated — and leaves the
per-caller assertion below it untouched, since that is what proves each caller gets
its own vectors in its own position.
This is a test that was over-specified, not an implementation that regressed. The
batch-size assertion the test is named for is unchanged.
Verified 8/8 under xdist on 3.14t, where it was failing intermittently, and on 3.11.
* build(docker): install Rust in the free-threaded builder, for litellm
The image build failed on amd64:
Failed to build `litellm==1.99.0`
Error: command ['maturin', 'pep517', 'build-wheel', ...]
Caused by: No such file or directory (os error 2)
litellm publishes only abi3 wheels (cp310-abi3), and abi3 — the stable ABI — does not
apply to a free-threaded interpreter, so uv cannot use them and falls back to the
sdist, which builds litellm's Rust extension. litellm is a hard runtime dependency
(pyproject pins it per-platform), so this is not optional.
It is the same constraint that kept toktok from working free-threaded before #4022:
an abi3 wheel is invisible to a cp3XXt interpreter.
Rust is confined to the builder stage — the runtime image copies only the venv and
carries no toolchain. Drop this once litellm publishes cp3XXt wheels.
Only amd64 hit it: my local builds were linux/arm64, where the resolution differed.
That is a good argument for the CI job building the image at all, which is what
caught this.
|
||
|
|
88f1472e52 |
docs(docker): reserve shared memory for embedded PostgreSQL (#3896)
* docs(docker): reserve shared memory for embedded postgres * docs(docker): keep shared memory guidance concise |
||
|
|
d7d137fa1e |
fix(daemon): drop the idle timeout that could kill an in-flight request (#3930)
* fix(daemon): drop the idle timeout that could kill an in-flight request The daemon's idle checker measured idleness as "time since the last request *started*" (IdleTimeoutMiddleware stamped last_activity on entry, never on completion), so a retain, reflect or consolidation that outlived the configured timeout was SIGTERM'd mid-flight — the client saw a connection error and the work was lost. Raising the timeout only lowered the odds; 0 (the default everywhere but the legacy cursor hook package) was the only safe value. Rather than teach the middleware to count in-flight requests, remove the feature: a shared local daemon that quietly exits under a long operation is not worth the resource reclamation it buys. The daemon now runs until it is stopped. - Delete IdleTimeoutMiddleware, the idle-checker thread and DEFAULT_IDLE_TIMEOUT. - Keep `--idle-timeout` parseable — hindsight-embed and every coding-agent integration still pass it — but ignore it, printing a note when it is non-zero. The integration packages therefore need no change and no release. - Drop it from the current docs, the SDK/integration READMEs and the generated docs skill. The coding-agents README keeps the row marked deprecated because docs-freshness.test.ts requires every readable RawConfig field to be documented; versioned docs, blog posts and changelogs are left as history. Closes #3903 Claude-Session: https://claude.ai/code/session_012gXUp1i7YrWmLJVUYki53g * review: drop the now-dead logger and refresh the comments the removal invalidated - daemon.py: the module logger only served the deleted idle checker. - coding-agents daemon.ts / runtime.ts / runtime.test.ts and openclaw: comments, a debug line and the plugin-schema description still described an auto-exit that can no longer happen. The config fields stay (inert) as agreed. Claude-Session: https://claude.ai/code/session_012gXUp1i7YrWmLJVUYki53g |
||
|
|
89f4d2e34d |
docs: stop recommending --user for Docker bind mounts (#3863) (#3934)
The bind-mount note told users to run the image as their host user with
`--user $(id -u):$(id -g) -e HOME=/home/hindsight` when the host directory
was not owned by UID 1000. That command crashes the container for every
host UID that is not 1000.
The image only creates the `hindsight` user (UID 1000), so any other UID
has no /etc/passwd entry. torch calls `getpass.getuser()` unconditionally
while resolving its inductor cache directory, which falls through to
`pwd.getpwuid(os.getuid())` and raises:
KeyError: 'getpwuid(): uid not found: 1042'
The reporter's directory was mode 0777 and owned by their UID, so the
pg0 writability pre-check passed and the failure surfaced later as an
opaque torch traceback rather than a permission error.
Replace the advice with the one supported option — chown the host
directory to 1000:1000 and run as the default user — and say explicitly
that --user with another UID is not supported, so the next person
recognises the getpwuid error. Point at a named volume for hosts where
chowning is not possible.
The same advice was printed by the pg0 writability failure message in
start-all.sh, so correct it there too.
Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk
|
||
|
|
a6a6c8e5a1 |
docs(docker): add guide and recipe for building CUDA standalone image (#3749)
* docs(docker): add guide and recipe for building CUDA standalone image Provide a Docker Compose example and standalone Dockerfile recipe for building a CUDA-enabled PyTorch image for NVIDIA GPU-accelerated local embedding and reranker models. Document prerequisites, build steps, and NVIDIA Container Toolkit configuration. * docs(docker): correct CUDA recipe verification, drop x86-only pin, state size Review follow-ups on the CUDA recipe: - The README told users to verify GPU placement by grepping the logs for both the embedding and cross-encoder device, but only the embedder logged one. LocalSTCrossEncoder resolved _device_type and never reported it, so the documented check showed a single line and looked like the reranker was still on CPU. Log the device on both reranker init paths and show the real log output in the README. - Dropped the hardcoded `--platform=linux/amd64`. PyTorch ships cu126 wheels for aarch64 too, and on an arm64 host the pin silently produced an emulated amd64 image that cannot reach the GPU at all. Documented instead that the build must be native. - Dropped `--index-strategy unsafe-best-match`. It relaxed the index isolation that hindsight-api-slim/pyproject.toml deliberately sets up, and it was not needed: resolving without it succeeds and yields the same package set. - Documented the image size (~11 GB vs ~9 GB for the base) in installation.md and the recipe README, since that cost is the reason no CUDA image is published. - Added a HINDSIGHT_VERSION build arg so the base tag can be pinned, fixed the manual `docker build` context, and enabled RERANKER_LOCAL_FP16 in the compose file (faster on GPU, quality-identical). Verified: `uv pip install` inside the base image replaces only torch (2.10.0+cpu -> 2.10.0+cu126) and adds the nvidia/cuda runtime wheels; compose config validates; docs skill regenerates clean; test_local_cross_encoder.py passes (21). Claude-Session: https://claude.ai/code/session_017ufCz6qrNxn36Stug7ek8A --------- Co-authored-by: Nicolò Boschi <boschi1997@gmail.com> |
||
|
|
bdee2bc88d |
fix(llamacpp): report a missing llama.cpp instead of endless connection errors (#3758)
Running the published Docker image with HINDSIGHT_API_LLM_PROVIDER=llamacpp downloaded a 3.5 GB model, crash-looped, and — once a model was supplied by hand — failed every retain and reflect with an unexplained APIConnectionError against 127.0.0.1 (#3733). Three separate defects stacked up: * A failed start still installed the shared server: it was assigned to the module global before start() was awaited, so every later call took the "already started" branch and built an OpenAI client against a port nothing listens on. The real error (no module named 'llama_cpp') surfaced once, at boot, where LLM verification only warns — and was masked from then on. The server is now published only once it is serving, and a subprocess left behind by a timed-out start is reaped so a retry cannot stack another one. * The model downloaded before anything checked whether a server could run at all. The `local-llm` extra is now required up front, with a message naming both ways out: the extra for local installs, the llama.cpp sidecar for the published image, which deliberately omits it. * HINDSIGHT_API_MODEL_INIT_TIMEOUT — the documented knob for a slow first-time download — had no effect in Docker, because the container entrypoint waits for /health on its own undocumented timer and killed the container at 300s regardless. That wait now follows the API's cap when it is the longer of the two, plus a grace period so the API reports its own timeout first. Docs: the configuration page advertised the built-in provider with no Docker caveat and the image variants table read as "works out of the box except the LLM", so the setup looked supported. Both now point at the sidecar compose file, HINDSIGHT_API_STARTUP_WAIT_SECONDS is documented, and auto-downloaded models are noted as needing persistent storage. |
||
|
|
e0221cae6e |
docs: flag Intel (x86_64) macOS as slim-only in supported-platforms grid (#2115) (#2129)
* docs: flag Intel (x86_64) macOS as slim-only in supported-platforms grid (#2115) `pip install hindsight-all` on Intel Macs silently backtracks to a months-old release: every release since 0.4.18 pulls hindsight-api-slim[all], whose local-ML extra requires torch>=2.6.0 and mlx, neither of which ships x86_64 macOS wheels. The docs' supported-platforms grid claimed Intel macOS bare-metal pip was "fully supported", which is false. - Split the macOS grid row into Apple Silicon (fully supported) and Intel/x86_64 (Docker + pg0 ✅, bare-metal pip ⚠️ slim only). - Mirror the grid into README.md. - Point Intel-Mac users to hindsight-all-slim / hindsight-api-slim plus a hosted embeddings/reranker provider or the in-process ONNX backend (which has x86_64 macOS wheels). - Replace the ad-hoc warnings with one-line pointers to the grid. Refs #2115 * docs: move Supported Platforms grid to bottom of README * docs: simplify README platform table to icons, link docs for details |
||
|
|
75a7c19d6a |
fix(docker): clear diagnostic for pg0 bind-mount permission failure (#1483) (#2010)
* fix(docker): clear diagnostic for pg0 bind-mount permission failure (#1483) The standalone image runs rootless (UID 1000). A host bind mount whose directory isn't owned by UID 1000 — the default on macOS Docker Desktop and most non-1000 Linux hosts — makes embedded pg0 fail with the opaque "Permission denied (os error 13)". Auto-chowning the volume would require running as root, which we deliberately avoid. Instead: - Recommend a Docker named volume in the README/installation docs; named volumes are seeded with the image's UID-1000 ownership, so they work with zero setup and stay rootless. - Add a pg0 writability pre-check in start-all.sh that prints an actionable message (named volume, or --user) and exits cleanly instead of letting pg0 emit os-error-13. Skipped when an external database is configured. - Add regression tests for the new check in test-start-all.sh. * docs(readme): drop bind-mount explanation, keep named-volume fix |
||
|
|
8a1f0461cf |
docs(docker): drop --rm, add --name + restart policy in run examples (#1927)
A single child segfault under load propagates through start-all.sh and exits the whole container; with the documented --rm run there was no recovery. Replace --rm with --name hindsight --restart unless-stopped in the documented server-run commands so a transient crash self-heals. Leaves the throwaway --rm --entrypoint sh model-inspection command in custom-models/README.md untouched. Refs #1918. |
||
|
|
771922cd70 | docs: add Windows/China deployment guidance for embeddings config (#1549) | ||
|
|
b628716f15 |
fix(cp): improve access-key auth UX and harden middleware (#1533)
* fix(cp): improve access-key auth UX and harden middleware - Move logout button from sidebar to header bar (next to GitHub icon), shown only when access-key auth is configured - Remove redundant status bar from dashboard page - Return 401 JSON for unauthenticated API requests instead of HTML redirect - Redirect to /login on 401 in the API client (skip if already on /login) - Allow /logo.png through middleware for the login page - Replace brain emoji with Hindsight logo on login page - Fix error message visibility in dark mode - Add loading spinner for bank selector while banks are fetching - Expose access_key_auth as a feature flag via version endpoint - Document HINDSIGHT_CP_ACCESS_KEY in configuration and installation docs * fix(cp): spread default features to handle unknown fields from API * fix(cp): wrap login page in Suspense for useSearchParams |
||
|
|
312bde1b4d |
docs: surface stable worker_id guidance and zombie-operation recovery (#1522)
* docs: surface stable worker_id guidance and zombie-operation recovery Worker identity defaults to the container hostname, which Docker rotates on every restart. That stranded several real deployments' consolidation queues (issue #1470 and the related closed tickets #991 / #696 / #624). Move the guidance from the configuration reference table — where it only gets read after the bug bites — into the install path and add a recovery section next to the decommission commands. * docs(faq): add zombie-operations entry |
||
|
|
b20bcb62c1 |
follow-up to #1459: AlloyDB ScaNN docs + review-nit cleanups (#1506)
* docs: document AlloyDB ScaNN vector extension Follow-up to #1459. Adds `scann` to the supported vector-extension list in installation.md and configuration.md, with installation hints, the 10k-row deferred-build caveat, AlloyDB Omni compose pointer, and the relaxed switching rules (switching *to* scann is allowed with existing data). * refactor(_vector_index): address review nits from #1459 - Lift `from sqlalchemy import text` (and add `Connection`) to module top in `_vector_index.py`; both helpers now have proper type hints. - Make `pg_diskann` a first-class entry in a new `RESOLVED_EXTENSIONS` tuple via `_normalize_resolved`. The configurable boundary stays strict (`validate_extension` rejects `pg_diskann`); the resolved helpers (`index_using_clause`, `index_type_keyword`, `minimum_rows_for_index`, `uses_per_bank_vector_indexes`) accept it without per-call special-case branches. Behavior is identical. - Harden `test_alembic_vector_migrations_freeze_vector_sql_locally` to resolve the migrations dir from `__file__` so the test no longer depends on cwd. - Add a one-liner explaining why `_drop_per_bank_vector_indexes` inlines identifiers instead of using bound parameters (DDL). Tests: tests/test_vector_index.py (10), tests/test_migration_shape.py + tests/test_migrations_thread_safety.py (64). Lint and ty clean. |
||
|
|
095f397770 |
docs(installation): bake custom models into image instead of PVC (#1504)
* docs(installation): bake custom models into image instead of PVC Add a runnable example under `docker/docker-compose/custom-models/` that extends the slim image and pre-downloads non-default embedder/reranker models at build time. Document this as the recommended pattern for production over enabling the Helm `modelCache` PVC: image layers cache per node for free, while a PVC adds storage cost, pins pods to a node, and needs lifecycle management on uninstall/upgrade. Add pointers from the api/worker `modelCache` values in the chart to the new section. Refs vectorize-io/hindsight#1383 * fix(docker/custom-models): install local-ml deps via uv into the venv The slim image's venv at /app/api/.venv was created by uv sync and does not ship its own pip, so a bare `pip install` falls through to the system pip and lands the packages in /home/hindsight/.local — invisible to the venv python that runs hindsight-api at runtime. Use `uv pip install --python /app/api/.venv/bin/python` to install into the venv directly. Verified the resulting image loads both baked-in models with HF_HUB_OFFLINE=1. * docs(installation): trim custom-models section to a tip and pointer The Dockerfile/compose example in docker/docker-compose/custom-models/ already has its own README explaining when to use it and why it beats the modelCache PVC. The installation page only needs to point readers there. |
||
|
|
e63100b6a2 |
ci: cosign-sign release images + document verification (#1502)
* ci: cosign-sign release images + document verification Folds the now-proven keyless cosign signing flow into the release workflow so future releases sign automatically alongside the build, and adds a "Verifying image signatures" subsection to the Docker installation docs so downstream consumers know how to verify. The verification regex accepts signatures from both sign-images.yml (used to backfill 0.6.0) and release.yml (future releases) so a single documented command covers all signed tags. Closes #1484 * docs: tighten cosign verification section |
||
|
|
ae0e3cec8d |
docs(installation): document memory footprint and hardware requirements (#1282)
* docs(installation): document memory footprint and hardware requirements Add a Hardware subsection under Prerequisites with per-component RAM guidance (full vs slim image, control plane, worker, postgres) and extend the Docker Image Variants table with an Idle RAM column so users know what to provision before deploying. * docs(installation): leave Docker Image Variants table alone, soften GPU note - Revert the Idle RAM column on the Docker Image Variants table; the Hardware subsection already carries that detail. - Reword the CPU/GPU line: CPU is fine for dev and basic workloads, but the local cross-encoder reranker typically benefits from a GPU under production traffic — or offload reranking to an external provider. * docs(skill): regenerate hindsight-docs skill mirror |
||
|
|
ca561aca9e |
feat: add consolidation_max_memories_per_round config (#1123)
* feat: add consolidation_max_memories_per_round config Prevents a single bank with a large backlog from monopolizing a worker slot. When the limit is reached, the consolidation job yields its slot and re-queues itself so other banks get fair scheduling. Mental model refreshes only run on the final round (when all memories are processed). Default: 100 memories per round. Set to 0 for unlimited (previous behavior). Configurable per bank via the config API. * fix(docs): fix broken anchors in blog post and installation pages - Blog post linked to non-existent #embeddings--reranker-providers anchor - Installation pages linked to removed #package-variants heading * fix: update configurable fields count and add openai-agents frontmatter - Bump expected configurable field count from 34 to 35 (new consolidation_max_memories_per_round) - Add missing title/description frontmatter to openai-agents integration doc * chore: regenerate docs skill references * chore: fix openai-agents formatting (pre-existing lint drift) |
||
|
|
349c112c61 |
docs: add supported platforms and Windows installation guide (#700)
* docs: add supported platforms section and Windows installation guide Adds a platform compatibility table (Linux, macOS, Windows) and a dedicated Windows setup section with step-by-step instructions for installing PostgreSQL + pgvector and running Hindsight natively. Follows up on #699 which added Windows native support. Also fixes a ty type-check error in metrics.py for the conditional resource module import. * chore: sync generated clients and lock file after #699 Regenerate client SDKs to pick up ValidationError model changes and update uv.lock with platform-specific uvloop/winloop deps. * docs: update Windows section — pg0 now supports Windows pg0 v0.12.0 added Windows support, so embedded DB works everywhere. Restructure Windows section to show simple install-and-run first, with external PostgreSQL as an optional alternative. * chore: sync generated docs skill and openapi references |
||
|
|
4a69a422a0 | doc: fix build | ||
|
|
15ea23d5d6 |
feat: introduce hindsight-api-slim and hindsight-all-slim packages (#560)
* feat: introduce hindsight-api-slim and hindsight-all-slim packages Closes #552 - Move all source code from hindsight-api/ to new hindsight-api-slim/ - hindsight-api-slim has heavy ML deps (torch, sentence-transformers, transformers, einops, flashrank, mlx, mlx-lm, safetensors) and pg0-embedded as optional extras: [local-ml], [embedded-db], [all] - hindsight-api becomes a zero-code meta-package depending on hindsight-api-slim[all] for full backward compatibility - Add hindsight-all-slim meta-package: hindsight-api-slim + client + embed - hindsight-all updated to depend on hindsight-api-slim[all] - pg0.py: lazy-import pg0 with clear ImportError pointing to [embedded-db] - Dockerfile: replace sed hack with proper uv sync --extra flags - Update release.yml, test.yml, lint.sh, release.sh, CLAUDE.md and all path references throughout the repo * refactor: rename hindsight/ directory to hindsight-all/ * docs: document hindsight-api-slim and hindsight-all-slim package variants Add package variants table and extras explanation to installation.md * docs: remove emojis from installation.md, use professional tone * docs: link Docker slim variant to pip package variants section * docs: consolidate Docker image variants into single table * ci: fix working-directory paths after package restructure - Replace all hindsight-api → hindsight-api-slim in test.yml - Replace hindsight → hindsight-all in test.yml - Add --extra embedded-db to test-embed API install step * ci: add local-ml and embedded-db extras to API sync steps These extras were previously implicit in the old hindsight-api package (which bundled everything). Now that hindsight-api-slim uses optional extras, we must explicitly request local-ml and embedded-db in CI. * ci: add API install step with embedded-db to test-embed smoke test The smoke test starts hindsight-api as a daemon, which requires pg0-embedded. Add a dedicated install step for hindsight-api-slim with embedded-db extra so the daemon can start successfully. * ci: remove --no-install-project when using optional extras When --no-install-project is combined with --extra, the optional deps are not installed because extras require the project to be active. Remove --no-install-project from steps that need local-ml or embedded-db. * ci: fix ordering of uv sync steps to preserve optional extras When uv sync runs for a different workspace member, it removes optional extras installed for other members. Fix by always running extra-requiring API sync last, after other workspace member syncs. Also remove --no-install-project from embedded-db sync in test-embed, as --no-install-project prevents optional extras from being active. * ci: add local-ml extra to test-embed API install for smoke test The smoke test starts the full API server which needs sentence-transformers for local embeddings (default provider). Add local-ml extra to the install. * ci: simplify extras with --all-extras and add slim pip smoke test - Replace explicit --extra local-ml --extra embedded-db with --all-extras for cleaner, more maintainable sync steps - Add test-pip-slim job: tests hindsight-api-slim[embedded-db] without local ML models, using Cohere for embeddings/reranking (mirrors Docker slim smoke test approach) * ci: simplify slim smoke test to health check only (mirrors Docker test) |
||
|
|
b813bd2728 | doc: fix build | ||
|
|
278344b3b3 |
doc: improve api explanation (#415)
* doc: improve api explanation * doc: improve api explanation * doc: improve api explanation * fix: add include_facts to reflect client, fix retain.sh temp files, fix main-methods based_on access * fix: create report.pdf in working directory for retain.sh file upload examples |
||
|
|
476726c2a2 | feat: support azure pg_diskann (#381) | ||
|
|
a1f22dabd2 |
Replace waitlist links with direct Hindsight Cloud signup URL (#349)
The waitlist is no longer needed. Update all references from vectorize.io/hindsight/cloud to ui.hindsight.vectorize.io/signup and change "request early access" language to "sign up". |
||
|
|
f64817814a |
feat: slim docker distro (#314)
* feat: slim docker distro * feat: slim docker distro * push |
||
|
|
4c792400c1 |
feat: new 'worker' service (#176)
* feat: new 'worker' service * doc * docs * tests |
||
|
|
eb2702bcba |
misc: performance improvements (#140)
* misc: performance improvements * misc: performance improvements * misc: performance improvements |
||
|
|
1c6acc3ba0 | feat: simplify mcp installation + ui standalone (#41) | ||
|
|
476a62da47 |
Add Hindsight Cloud links to README and docs (#42)
- Add Hindsight Cloud link to README header - Add Hindsight Cloud navbar item in docs - Add callout in installation docs for managed alternative |
||
|
|
47be07f97f |
bump pg0 0.11.x and improve documentation (#33)
* bump pg0 0.11.x and improve documentation * bump pg0 0.11.x and improve documentation * bump pg0 0.11.x and improve documentation * ci: test notebooks on ci * ci: test notebooks on ci * rm llms-full from repo * formatting * formatting |
||
|
|
99db7b26c3 | fix docs on clients | ||
|
|
fa554b8980 | brandind and misc fixes | ||
|
|
6073ac4ffd | docs, packages and quick start | ||
|
|
4b8fccb5e8 | fix readme github images | ||
|
|
58592d4abc |
fix docker cp image build on ci (#10)
* fix docker cp image build on ci * fix docker * fix docker again |
||
|
|
f7cf33c610 | fix regressions and bunch of issues |