mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
10573afd0a
Answers "what is burning the CPU?" for a process you cannot attach a debugger to.
HINDSIGHT_API_PROFILE holds JSON -- the shape config.py already uses for structured env
config -- and unset, nothing starts:
HINDSIGHT_API_PROFILE='{"every": 60, "top": 20}'
It reports to the log stream rather than a file or an endpoint, because the case it
exists for is a process dying without explanation: a file inside the container dies with
the container unless a volume was mounted in advance, and an endpoint needs a live
process and a route to it. Container runtimes keep the previous container's stdout, so
the last report before a crash is still readable afterwards. Each report is flushed as it
is written, since a fatal signal takes buffered output with it.
Three things the implementation had to work around, all verified on a running API:
* The profiler is process-wide and single-instance. Since 3.12 it is a global
monitoring tool, so enable() covers every thread whatever thread calls it, and a
second concurrent profiler raises `tool 2 is already in use`. Per-thread profilers
are not possible; this arms one for the process.
* Snapshots must use getstats(), which reads the accumulated entries without stopping
the profiler. Snapshot-and-clear through pstats disables the global tool, and every
report after the first then silently contains nothing. Reports are deltas between
snapshots, and a test asserts a second window still has data.
* Sampling via sys._current_frames() is not an alternative on a free-threaded build: it
stops the world, so it catches threads parked at safe points, which are I/O waits. It
reported event-loop threads idle in selectors.select while /proc showed those same
threads at 50-65% of a core. Every report therefore carries per-thread CPU from
/proc, which profiler overhead cannot distort, as the arbiter.
py-spy remains the better tool on a GIL build. It cannot read a Py_GIL_DISABLED process
at all -- it locates threads through the GIL -- which is what left free-threaded
deployments with nothing, and is why this exists.