Files
Nicolò Boschi 5f9bff9056 fix(docker): drop curl from the API runtime images (#4202)
* fix(docker): drop curl from the API runtime images

curl's only in-container consumer was the readiness loop in start-all.sh;
there is no HEALTHCHECK instruction anywhere. It is also the sole
reverse-dependency of libcurl4t64, which brings libssh2-1t64, so one line in
each install list accounted for nine HIGH findings - all status=affected with
no Debian fix published, so `apt-get upgrade` could not clear them and not
shipping the package was the only remediation.

Trivy 0.74.0 HIGH+CRITICAL, on locally built slim images:

  api-only     3C / 60H -> 3C / 51H
  standalone   3C / 60H -> 3C / 51H

with exactly the curl, libcurl4t64 and libssh2-1t64 findings removed and
nothing new.

Replace it with http_probe, which reproduces `curl -sf` WITHOUT -L rather than
approximating it. The distinction matters: curl does not follow redirects
unless asked, so a 302 is a completed transfer and succeeds regardless of what
it points at. urllib.request.urlopen follows it and raises on a 404 behind it,
which would report a healthy service that redirects as "not ready". http_probe
uses http.client and tests `status < 400` itself, bypassing urllib's redirect
handler, and carries userinfo through as Basic auth the way curl does.

Verified equivalent to `curl -sf` on 2xx, 3xx-to-good, 3xx-to-bad, 4xx, 5xx,
query strings, userinfo auth and connection refused. Exit codes are not
reproduced (curl's 22 and 7 become 1); every call site tests zero/non-zero.

One deliberate difference, since it is a change and not a translation: curl was
called with --connect-timeout, which caps only the connection phase, and the
API health loop passed no timeout at all - so a server that accepted a
connection and never answered hung the probe forever. The timeout now covers
the whole request.

There is no wget fallback. BusyBox wget cannot reproduce these semantics (no
--max-redirect, so it always follows), and it is not needed: every image that
probes anything is Python-based. cp-only, the one image with neither, performs
no probe at all and dropped curl in #4197. Missing python3 now fails loudly at
startup instead of degrading into a readiness loop that can never succeed.

Closes #4198

* refactor(docker): move the readiness probe into hindsight_api.http_probe

The first version of this probe was Python embedded in a shell string inside
start-all.sh. That was a bad shape for code encoding rules this fiddly: every
quote had to survive two levels of escaping, ruff and ty never saw it, and it
could only be exercised through the shell.

Move it to hindsight_api/http_probe.py, shipped with the code and covered by
tests/test_http_probe.py, which pins each case to what `curl -sf` does for the
same response. start-all.sh keeps a three-line wrapper that shells out to
`python3 -m hindsight_api.http_probe`.

hindsight-admin was the obvious home and is the wrong one: it takes 5.1s to
start in the built image, against 0.028s for bare stdlib, because it pulls in
the CLI and everything behind it. The readiness loop polls once per second, so
importing the API to ask whether the API is up would break the loop it drives.
This module imports stdlib only; measured 0.035s per probe in the image.
`hindsight_api/__init__` is cheap by design and has to stay that way for this
to hold - its docstring already says so.

The shell test drops to checking the wiring, since the semantics now have a
real home, and skips when the package is not importable: test-start-all.sh also
runs in CI from a bare checkout with no virtualenv.

Reformatting by `ruff format` on first contact is the point - the embedded
version could never have received it.

* refactor(probe): make the readiness probe its own package, isolated from the API

hindsight_api.http_probe was the wrong home. The probe answers "is an API
process up?", and living inside the package it probes invited exactly the
coupling that would break it: an import of the engine or the config would put
API startup cost - and API startup side effects - on a loop that runs once a
second.

Move it to hindsight_probe, a sibling top-level package in the same
distribution. Its dependencies are now explicit by construction: none. It
imports the standard library and nothing else.

Packaging alone does not enforce that. Both packages install into the same
virtualenv, so `import hindsight_api` from the probe would still resolve at
runtime. So the rule is a test, not a convention:
test_imports_nothing_but_the_standard_library imports the package in a clean
subprocess and asserts that nothing outside sys.stdlib_module_names was pulled
in. Adding `import hindsight_api` to the probe fails it with the offending name.

The audit ignores _sysconfigdata_*, a platform-specific stdlib internal whose
name embeds the build triple and so is absent from stdlib_module_names
everywhere.

Both Dockerfiles now copy the package; the api-builder previously copied only
hindsight_api, so the first build without this shipped an image whose probe
could not import. That surfaced as require_http_probe_runtime failing at
startup with a clear message rather than a readiness loop that could never
succeed, which is what that guard is for.

Verified in the built Linux image: no non-stdlib imports, hindsight_api never
loaded, 0.036s per probe, and the full end-to-end boot still reaches
"Hindsight is running" with /health answering 200.

* refactor(probe): keep the readiness probe inside hindsight_api

Reverts the separate hindsight_probe package. It was justified on a bad
measurement: an earlier cold-cache timing suggested `import hindsight_api` cost
~0.12s against ~0.03s for a standalone package. Measured properly, warm, in the
built image, they are the same - ~0.03s each - and importing hindsight_api
pulls in zero third-party modules. Its PEP 562 lazy-attribute design already
does the work the split was meant to do, so the split bought nothing and cost a
second top-level package, four pyproject entries and a COPY in each Dockerfile.

What was worth keeping is the enforcement, which is orthogonal to where the
module lives. test_imports_nothing_heavy imports the probe in a clean
subprocess and asserts it pulled in no third-party package and nothing from
hindsight_api.engine, .api or .config. Adding `from hindsight_api.engine import
memory_engine` to the probe fails it with 43 packages named, numpy, sqlalchemy
and asyncpg among them - which is the failure mode the rule exists to prevent.

Verified in the built image: no third-party or engine imports, 0.034s per
import, probe wiring works, curl absent.
2026-09-08 11:11:04 +02:00
..