Commit Graph

99 Commits

Author SHA1 Message Date
Nicolò Boschi b1de1b9418 fix(operations): allow cancelling in-flight operations (#4131)
`DELETE /v1/default/banks/{bank}/operations/{id}` only accepted `pending`
operations, so an operation stranded in `processing` — orphaned when a worker
was killed before it could write a terminal status — could only be cleared by
hand-editing `async_operations` and restarting the container.

Cancel now accepts `processing` too. It stays cooperative and is never
immediate: the row is flipped to `cancelled` and the worker running it stops at
its next `_check_op_alive` checkpoint (between retain sub-batches/documents,
between consolidation LLM batches). For the orphaned case nothing is running, so
the flip is the whole fix. No heartbeat and no per-batch bookkeeping is added.

Making the flip stick required guarding the worker writes that had none, and
would otherwise overwrite it:

- `_schedule_retry` — the one that actually resurrected cancelled work: a task
  failing after cancellation went back to `pending` and was re-claimed.
- `_mark_failed` (poller and engine), `_defer_operation`.

`_mark_completed` already guarded on `status='processing'`; tests now pin it.

The sibling rollup counted only `completed`/`failed` as done, so a cancelled
child stranded its `batch_retain` parent in `processing` forever — the same
wedge one level up. Both rollup copies now treat `cancelled` as done and settle
the parent on `cancelled` (a real failure still outranks it), cancel performs
the rollup itself so cancelling the last outstanding child terminalizes the
parent, and a cancelled parent is never flipped back by a child finishing later.

Control plane: the Cancel button was gated on `pending`, hiding the fix from the
UI an operator would reach for. It now shows for `processing` rows too.
2026-09-07 11:21:06 +02:00
2anoubis a72f9de15a fix(cli): restore default SIGPIPE handling to prevent abort on broken pipe (#3925)
* fix(cli): restore default SIGPIPE handling to prevent abort on broken pipe

Release builds set panic = "abort", so when stdout is piped to a
consumer that closes early (e.g. `hindsight ... | head`), Rust's
println! panics on EPIPE and the process dies with SIGABRT instead of
terminating cleanly. Restore the default SIGPIPE disposition so the CLI
exits normally on a broken pipe, matching standard Unix CLI behavior.

* fix(cli): gate SIGPIPE reset to Unix and add regression test

libc::SIGPIPE does not exist on Windows (libc only defines SIGINT, SIGILL,
SIGFPE, SIGSEGV, SIGTERM, SIGABRT there), so gate the signal() call with
#[cfg(unix)] to keep the Windows build compiling.

Also add a deterministic regression test that spawns the binary with stdout
on a pipe whose read end is already closed, asserting the process exits with
SIGPIPE (signal 13) rather than aborting.

---------

Co-authored-by: Mike <mike@email.nonexist>
2026-09-07 10:50:32 +02:00
Nicolò Boschi 163fbb0ede feat(api)!: retire the bank profile and background endpoints (#4127)
* feat(api)!: retire the bank profile and background endpoints

GET/PUT /v1/default/banks/{bank_id}/profile and
POST /v1/default/banks/{bank_id}/background have been deprecated for
several releases. They now answer 410 Gone with the replacement call in
the detail, joining the two endpoints (entity regenerate, synchronous
document export) that already do.

The routes stay in the OpenAPI spec with unchanged signatures, so no
generated SDK method disappears from under a caller — only the behaviour
changes.

Disposition traits and the reflect mission are bank configuration, and
already were: _get_bank_profile_authenticated overlaid config on top of
the legacy DB columns. The `name` these endpoints also returned is a
display-only label available on the bank list.

To make the config API a complete replacement, GET .../config is no
longer gated on HINDSIGHT_API_ENABLE_BANK_CONFIG_API — that flag now
gates only the writes (PATCH/DELETE). A bank must always be able to read
its own resolved settings.

Clients migrated in the same change:
- control plane: bank-profile-view and bank-config-view read disposition
  and mission from the config API, the display name comes from the
  filtered bank list, and the dead /api/profile proxy route is gone.
- hindsight-cli: `bank disposition`, `bank set-disposition` and the
  hidden `bank background` move to the config API. `background` warns
  that it now replaces the mission rather than LLM-merging into it —
  nothing replaces that merge — and its `--no-update-disposition` flag
  is accepted but ignored, as the server stopped inferring disposition
  from the mission long ago.
- TS wrapper: getBankProfile carries a @deprecated pointer.

* test(control-plane): cover the composed bank profile, and drop its extra fetch

bank-context only needs the display name, so it reads the id-filtered bank
list directly instead of going through getBankProfile, which would also
fetch the bank config it has no use for.

* feat(cli)!: drop the deprecated `bank background` command

The server endpoint is gone, and the LLM merge it performed has no
replacement — `bank mission` sets the mission outright. Keeping the
command as an alias would have silently turned a merge into an
overwrite, so it is removed rather than repointed.

* fix(ci): update the CLI coverage manifest, doc example and TS client test

- .openapi-coverage.toml: the three retired operations move to [skip]
  alongside export_documents_sync_removed, and the stale add_bank_background
  / update_bank_disposition field sections are dropped. The CLI helper is
  renamed set_bank_disposition so it no longer satisfies the coverage grep
  by name while calling update_bank_config underneath.
- cli-reference.sh: the two `bank background` snippets become one
  `bank mission`, the command that replaces them.
- main_operations.test.ts: TestBankProfile asserts the 410 and reads the
  same data back from the bank config. try/catch rather than .rejects,
  since this file runs under both jest and Deno's @std/expect shim.

* style: rustfmt the CLI edits, and say why the 410 handlers keep unused params
2026-09-04 18:09:05 +02:00
Nicolò Boschi 280f098202 feat(retain): inline images and files as first-class content (#4077)
Makes images and files first-class raw content in `retain`. `content` accepts an
ordered list of text/image/file blocks, the extractor reads each attachment in
the position it occupies, and every read surface hands back the attachments
behind what it returns. A plain string behaves exactly as before — text-only
retain is byte-identical, because everything new sits behind an ATTACHMENTS
block that is empty when a chunk carries none.

Blocks are flattened at the API boundary into one canonical body with atomic
placeholders, so `documents.original_text` stays plain text and content_hash
idempotency, `update_mode=append`, chunk-delta re-extraction and
`reprocess_document` keep working untouched. Bytes live in the existing
FileStorage abstraction, content-addressed by sha256.

Schema (one migration, both dialects): `attachments` for the blob,
`document_attachments` for which documents reference it, and
`memory_units.attachment_ids` for which attachments a *fact* came from — a
column rather than a third table, because those ids behave exactly like `tags`.

Provenance is per fact, not per chunk. Extraction runs one call per chunk, and a
chunk holding a screenshot also holds the prose around it, so a chunk-level edge
cited the diagram as evidence for the paragraph that never mentioned it. The
extractor is asked instead, and a fact stated in the prose carries nothing.

Extraction quality was measured against a real image-QA dataset with a raw-VLM
ceiling arm before merging: transcribing structured attachments rather than
summarizing them, and recording how each value is drawn, took the gap between
"the model can read this off the image" and "memory can answer it" from 31.3% to
10.0% on the same 40 charts. The prose-article benchmark went 75% -> 100% over
the same change, so it is not chart-specific tuning.

Also here:

* A vision slot (`HINDSIGHT_API_VLM_*`) so attachment-bearing chunks alone use a
  vision model and text-only chunks stay on a cheaper retain LLM. A vision call
  deliberately does not fail over to the retain chain's text models — that would
  reintroduce the silent omission the 422 gate exists to prevent.
* The extension retain hook can now see each attachment (media type, size, kind,
  filename) and refusing a retain reclaims its bytes, which previously stayed
  fetchable forever.
* A filename lives on the document edge, not the blob: the same PDF can be
  attached under a different name elsewhere, and content-addressing made the
  first name win for both.

Known limitations, documented rather than hidden: store-owned memory backends
get nothing (that retain path is Postgres-free and pre-dates this work), very
dense pages are sampled rather than exhausted, and the Python client's
ContentBlock is a plain dict where TypeScript gets the real union.

Breaking for Go and Rust callers: `content` is now a union, so a bare string no
longer satisfies it. Go gains a `TextContent()` helper; Rust uses
`Content::Variant0(...)`.
2026-09-04 12:48:08 +02:00
Nicolò Boschi 737e5bf420 fix(cli): print the server's response body on API errors (#4049) (#4113)
`hindsight document get <bank> <missing-id>` reported "API endpoint not
found (404)" with endpoint-path/version guidance, while the server had
answered 404 `{"detail":"Document not found"}` — pointing the operator at
an API breakage instead of an absent document.

Two things were dropping the body:

- `api.rs` only ran three of ~70 generated-client calls through
  `humanize_client_error`; the rest used a bare `.await?`, so the error
  rendered via progenitor's `Display` as a body-less "Unexpected
  Response: Response { .. }" and the body was never read at all. Every
  call now goes through a `Humanized` extension trait on the call future
  (`call(..).humanized().await?`), which attaches status and body.
- `errors.rs` discarded the body in the 404/401/403/5xx branches.
  `server_detail` now pulls the JSON `detail` (or the raw body) out of
  the error and every branch surfaces it: a 404 that explains itself
  prints "Not found (404): Document not found", and only a bodyless 404
  — a genuine unknown route — keeps the path/version guidance.

Nothing in the type system stops the next call site from writing the
bare form, and no unit test exercises a real error response, so
`tests/api_error_surface.rs` checks the whole family: every statement in
`api.rs` that awaits a generated-client call must end in `.humanized()`.
2026-09-04 11:17:27 +02:00
Nicolò Boschi d20893a83f fix(api): paginate GET /observations/scopes and GET /webhooks (#4075)
* fix(api): paginate GET /observations/scopes and GET /webhooks

Both endpoints returned the whole collection: the scopes endpoint grouped
every observation in a bank and shipped one row per distinct tag set, and
the webhooks endpoint selected every webhook row. Neither is bounded by
construction — a bank has as many scopes as it has distinct tag sets — so
the payload and the per-row work grew with the data.

Both now take `limit` (default 100, max 1000) and `offset` and return
`total` / `limit` / `offset` alongside the page, the same shape
`list_tags` and the other paged list endpoints use. The bound reaches the
SQL: the scope histogram is grouped, ordered and paged in one query with
a separate `COUNT(DISTINCT scope)` for `total`, and the webhook listing
gets `LIMIT/OFFSET` plus a count query. Webhook rows now order by
`created_at, id` so a page boundary cannot fall inside a group of rows
sharing a timestamp.

Consumers page with them: the control-plane webhooks view walks every
page (it renders the whole list and its count), the two proxy routes
forward `limit`/`offset`, and the CLI prints "Showing N of M webhooks"
rather than presenting the first page as the whole list.

Clients, OpenAPI spec and the docs skill regenerated.

* fix(tests): page the in-memory store's observation_scope_counts stub

The MemoriesExtension conformance test asserts every stub method accepts
the interface's params, so widening the interface with limit/offset left
the InMemoryMemories stub behind (test-api shards 2/3 and free-threaded).
It now groups, orders and pages its own observation rows and returns the
same {scopes, total, limit, offset} shape as its list_tags neighbour,
rather than the bare list it used to hand back.

* fix(control-plane): page the webhooks view instead of loading every page

The first cut walked every page in loadWebhooks and rendered the lot, so
the bound existed but nothing in the UI showed it — a bank with hundreds
of webhooks still built one unbounded table, and a reader had no way to
tell there was more than a screenful.

It now fetches one 50-row page and exposes the pager entities-view uses:
first/prev/next/last, "1 / 3", and a "51-100 of 120" range line, with the
header count coming from `total` rather than the loaded page. The controls
hide when everything fits on one page.

Two paging hazards handled: a bank switch resets to page 1 (the page
number belongs to the bank being left, and carrying it over lands on an
offset the new bank may not reach), and creating a webhook jumps to the
page it lands on — rows are oldest-first, so reloading the current page
would close the dialog and appear to do nothing. Deleting the last row of
the final page steps back a page rather than stranding an empty one.
2026-09-04 10:34:54 +02:00
2anoubis ddf57744c1 fix(cli): preserve base64 '=' padding when parsing config-file api_key (#3909)
The config parser used line.split('=').nth(1), which splits on every
'=' in the line and silently strips the trailing padding characters
from base64 API keys read from ~/.hindsight/config. Keys ending in '='
were corrupted, so the CLI failed authentication when relying on the
config file instead of HINDSIGHT_API_KEY.

Use split_once('=') so the value keeps any literal '=' it contains, and
delegate the config-file loop to parse_config_value to share the logic
with profile loading. Add regression tests for base64 padding and
values containing multiple '='.

Co-authored-by: 2anoubis <2anoubis@users.noreply.github.com>
2026-08-31 12:30:04 +02:00
Miguel de Benito Delgado 1cc2f9c10b [CLI] Add a --trigger-mode full|delta flag to mental-models update (#3513)
Adds --trigger-mode full|delta to `hindsight mental-model` create and update, and
stops the update resetting the trigger settings the user did not name.

The server patches a supplied trigger over the stored one (#3506/#3549), but the
generated Rust client serializes its four non-Option fields unconditionally —
mode, refresh_after_consolidation, exclude_mental_models, keep_trace — so a
freshly-built trigger looked like an explicit request for this type's defaults.
Switching a model to delta mode also turned off its trace capture and its
mental-model exclusion. The update now reads the model's current trigger and
applies only the passed flags on top of it.

Also exposes --trigger-refresh-cron, --trigger-min-refresh-interval-seconds,
--trigger-tags-match, --trigger-keep-trace and --trigger-exclude-mental-models.

Co-authored-by: Miguel de Benito Delgado <m.debenito.d@gmail.com>
2026-08-31 11:31:41 +02:00
Nicolò Boschi 3de41af867 feat(recall): let callers supply the temporal window instead of parsing it (#3678)
* feat(recall): let callers supply the temporal window instead of parsing it

Recall derives the temporal arm's window by parsing dates out of the query
text. A caller that already knows the range it means — a date picker, an agent
that resolved "last quarter" itself — had no way to say so, and had to phrase
it in English and hope dateparser agreed.

Add `temporal_window: {start, end}` to RecallRequest. When set it is used
verbatim and the extraction is skipped entirely, which is also the point: that
work is pure CPU serialised through a single worker and costs up to ~1.3s on
document-sized query text, which is exactly what consolidation and reflect
recall with.

Naming and wording carry weight here, because the obvious reading of a date
range on a search API is "restrict results to this period" and that is not
what this does. The temporal arm is one of four retrieval arms: it surfaces
memories whose own dates fall in the window so fusion ranks them higher, and
the other three arms are untouched, so memories outside the window are still
returned. Every description — model docstring, OpenAPI field, MCP tool, both
wrappers, control plane, docs — says so explicitly.

It does not override `enable_temporal_retrieval`. That per-bank flag gates the
arm itself and stays the single switch for it, so a supplied window cannot
re-enable an arm a bank turned off.

Bounds are inclusive and naive datetimes are read as UTC at parse time, so
both ends are unambiguous before they reach a query; a reversed window is
rejected at the boundary rather than silently returning nothing.

The Rust struct literals in the CLI and the client's doctest have to name the
new field or progenitor's generated RecallRequest stops compiling.

* feat(control-plane): add the temporal window to the Recall Analyzer

The recall UI could not reach the window it now proxies. Adds a Time window
row to the Recall Analyzer: two datetime inputs, a Clear button, and a hint
that states what the window actually does — ranks memories dated in the range
higher, does not hide the ones outside it — since "date range on a search
form" reads as a filter otherwise.

The two rules live in lib/temporal-window.ts rather than the component so they
are testable: a window needs both ends (one alone is an incomplete range, not
a half-open filter), and a reversed range is withheld and flagged inline with
the Recall button disabled, instead of being sent for the API to reject with a
422.

Comparing the raw `datetime-local` strings is exact — they are already
YYYY-MM-DDTHH:mm, which sorts chronologically — so there is no Date parsing
and no local-timezone reinterpretation between the input and the request. The
value is sent with no offset, which the API reads as UTC, matching what
query_timestamp already does from this same form; the hint says so.

Strings added to all ten locales.

* fix(control-plane): reject a reversed window on Enter, not just on the button

Disabling the Recall button left the Enter-key handler on the query input
calling runSearch() directly. With a reversed range that ran the search anyway
and silently dropped the window, which is the failure the inline warning
exists to prevent. Guard in runSearch so every entry point agrees, and toast
the same message rather than doing nothing visible.

* feat(cli): expose the recall temporal window as --window-start/--window-end

check-cli-coverage caught that recall_memories gained a request-body field the
CLI neither exposes nor exempts. The exemption list is for genuinely complex
nested bodies — min_scores' four calibrated floats, tag_groups' boolean tree —
and two datetimes is not that, so expose it rather than write it off.

Flattened into two flags and recorded as such in the coverage manifest, the
same shape `include` already uses.

Both ends are required: one alone is an incomplete range, not a half-open
filter, and running a recall without the window the caller asked for is worse
than refusing. A reversed range is rejected before the request rather than
sent for the API to 422, and a datetime with no offset is read as UTC, which
is how the API reads one.
2026-08-21 13:02:46 +02:00
Alan Li 71029fb81d fix(cli): exit with non-zero status when configure receives an invalid API URL (#3356)
Previously, `hindsight configure --api-url <invalid>` printed an error
but returned exit code 0 because the validation block called
`print_error!` followed by `return Ok(())`. Scripts and CI pipelines
relying on the exit code could not detect the failure.

Replace the silent success with `anyhow::bail!` so the top-level error
handler prints the message and exits with status 1, matching the
behavior of every other CLI error path.

Add two integration tests:
- configure rejects URLs missing http(s):// with non-zero exit
- configure accepts both http:// and https:// URLs
2026-08-20 10:48:47 +02:00
Nicolò Boschi 2b54a1d3cd feat(mental-models): a minimum interval between automatic refreshes (#3480) (#3621)
* feat(mental-models): a minimum interval between automatic refreshes (#3480)

A bank with several models on refresh_after_consolidation paid for a full
agentic rebuild of every stale model on every retain, however small: a
three-fact retain cost ~250k tokens, and ordinary conversational traffic
reached ~11.6M tokens a day, ~90% of it refresh work.

min_refresh_interval_seconds puts a floor on how often that may happen,
configurable at all three levels and defaulting to 0 everywhere, so nothing
changes for anyone who does not set it:

  HINDSIGHT_API_MENTAL_MODEL_MIN_REFRESH_INTERVAL_SECONDS   server
  mental_model_min_refresh_interval_seconds                 tenant / bank
  trigger.min_refresh_interval_seconds                      one model

The per-model value wins, including an explicit 0 — that is how one model
that does need to stay current opts out of a floor its bank imposes.

The refresh is parked, not skipped. The handler raises DeferOperation before
any recall or LLM work, which leaves the operation queued with next_retry_at
set to the end of the window; every trigger that fires while it waits folds
into it (#3487), and the eventual run reads everything that accumulated.
Skipping the submit instead would drop the work: nothing re-checks the model
afterwards, so memories that arrived during a burst would stay unsynthesised
until some later, unrelated trigger happened to land outside the window.

Only the triggers nobody asked for individually are rate-limited — the
after-consolidation flush and the cron scan. An explicit refresh asked for one
now and gets one; when it folds into a parked refresh it also releases the
park, since otherwise "refresh now" would silently inherit somebody else's
wait. Releasing means dropping the automatic marker as well as next_retry_at:
left in place, the next claim would park the operation all over again.

Built on existing machinery rather than new scheduling: next_retry_at is
already indexed and honoured by every claim query on both dialects, and
DeferOperation already means "requeue later, not a failure". No new table, no
new sweep, and nothing dialect-specific — the payload edit is read-modify-write
rather than jsonb `-`, which Oracle's SQL translation does not rewrite.

The control-plane trigger form carries the field because update_mental_model
replaces the trigger wholesale: without it, editing a model in the UI would
silently wipe the interval (the #3549 bug class). The TypeScript wrapper client
needed it explicitly too — unlike the Python wrapper, which splats the whole
trigger dict, it maps a hand-picked list of trigger fields.

* feat(control-plane): a Mental Models section, and show when an operation is parked

Two gaps found while testing the refresh floor on a real bank.

The bank-level setting had no home in the UI at all — it was reachable only
through the config API. It now has its own "Mental Models & Knowledge Pages"
section after Reflect: it governs the models and the knowledge pages backed by
them rather than consolidation, and the section is where the rest of that
surface will go.

The operations list showed a parked refresh as plain "pending", which reads as
stuck. The dataplane has always returned next_retry_at, but no control-plane
type declared it and nothing rendered it, so the one piece of information that
explains the wait was invisible. A pending operation held until a known time now
renders as "Waiting" with that time, in the list and the detail panel; the badge
keys off status AND a future timestamp, because next_retry_at is not cleared
when the task finally runs.

That last point also corrects the API docs, which claimed next_retry_at is
"always null for completed tasks" — a completed operation that was deferred
still carries the timestamp of its last wait.
2026-08-19 16:24:38 +02:00
Yossi Ovadia b7c65fbb2c fix(cli): suppress ANSI color and spinner when output is not a TTY (#3598)
* fix(cli): suppress ANSI color and spinner when output is not a TTY

The CLI emitted truecolor gradient escapes and an 80ms animated spinner
unconditionally, with no terminal detection, NO_COLOR, or TERM=dumb
handling. Piped or redirected invocations were flooded with escape codes:
a real recall piped to a file is 2,254 escape sequences at HEAD, and a
long-running call sprays ~22KB of spinner frames. Hindsight's documented
primary consumers are coding agents, which read piped output by definition.

Add a central `ansi_enabled()` predicate (stdout().is_terminal() &&
NO_COLOR unset && TERM != "dumb") and apply it to the gradient/dim color
helpers, the spinner (now a no-op with no thread and no output when
disabled), and the --help logo. The decision layer is split into pure,
injectable functions so it is unit tested without depending on the test
harness's terminal. Interactive TTY output is unchanged.

* fix(cli): route the ANSI gate through colored's existing policy

Review follow-up on the non-TTY suppression.

The binary already resolves a color policy: `print_error` and everything
in `errors.rs` go through `colored`, which honours CLICOLOR_FORCE, then
NO_COLOR, then CLICOLOR plus a stdout tty check. The new predicate
implemented a second, narrower policy, so the two halves of one command's
output could disagree -- `CLICOLOR_FORCE=1 hindsight ... | less -R` kept
the error color but lost the gradients, with no way to get them back, and
`CLICOLOR=0` in a terminal did the reverse. Defer to `colored` and keep
`TERM=dumb` as the one addition on top. It is a `lazy_static`, so the
environment is now resolved once rather than on every styled string.

Also cover the wiring: every existing test injects the enabled flag, so a
helper that forgot to consult the predicate would still pass. Adds a unit
test that drives the real `gradient`/`dim`/`before_help_logo`/spinner
entry points with the override forced off, and an integration test that
asserts the actual binary emits no escapes when stdout is a pipe (which is
how `Command::output()` runs it).

`get_logo()` was left unreachable outside `ui.rs` once main.rs moved to
`before_help_logo()`; folded it in.

---------

Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
2026-08-19 14:36:27 +02:00
Nicolò Boschi 6692e38c80 fix(api): paginate the bank list (#3586)
* fix(api): paginate the bank list

GET /v1/default/banks returned every bank in the system: no limit, no
offset, and a query with no LIMIT clause. Beyond the unbounded payload, the
per-bank work — config resolution and a live store count for banks whose
memories live outside SQL — ran for every bank rather than the ones being
shown.

The endpoint now takes limit/offset (defaults 100/0, matching
list_documents) plus a `q` substring filter on bank id and name, and
returns total/limit/offset alongside `banks`. Paging happens after
filter_bank_list rather than in SQL: that extension hook can drop any bank,
so a SQL page would hand back short pages and a total counting banks the
caller can't see.

Consumers page instead of taking the first 100: the control-plane bank
selector scrolls infinitely and searches server-side, the CLI walks every
page, and the Zapier bank dropdown became canPaginate.

* fix(api): bound the bank-list probes and keep the selected bank's name

Follow-ups from reviewing the pagination change:
- the control-plane health probe and the Zapier credential test only need to
  know the endpoint answers, so they ask for limit=1 instead of a default page
- the header showed the raw bank id whenever the selected bank sat past the
  first page, so its name is fetched directly
- limit/offset are clamped in the engine: the page is a Python slice, and the
  MCP tool takes both straight from a model with no HTTP-layer validation

* docs(mcp): document list_banks query/limit/offset

* fix(control-plane): make the bank selector actually page and report empty searches

Verified against a 130-bank instance: the infinite scroll never fired. The
observer effect read listRef.current/sentinelRef.current on the commit that
flips the popover open, but Radix mounts the content in a portal afterwards, so
both refs were null and nothing re-ran the effect — the selector sat on its
first 50 banks forever. Tracking the nodes as state through callback refs
re-runs the effect when they attach; paging now walks offset 0/50/100 and stops
at the total.

An empty result also read "No memory banks yet." after a search that simply
matched nothing, so searches get their own message.

* feat(control-plane): smooth the bank selector as pages land and searches narrow

The list is paged and searched server-side now, so rows appear and vanish in
batches — every page landed as a hard 50-row pop, and a search that narrowed to
one bank snapped the popover shut from 300px.

- rows fade and lift in, staggered within their page and capped so the tail of a
  50-row page doesn't crawl; only rows that actually mount animate, so appending
  page 2 leaves page 1 still
- the list height follows cmdk's --cmdk-list-height, easing down to the filtered
  set instead of jumping
- the previous results hold their place and dim while the next set is in flight,
  rather than blanking on every keystroke

The animations are defined in globals.css next to the existing logo keyframes:
tailwindcss-animate is a Tailwind v3 plugin declared in tailwind.config.ts, but
this app runs Tailwind v4 with the CSS-first config, so `animate-in` and friends
compile to nothing here.

* refactor(control-plane): tidy the bank row className and import
2026-08-18 17:27:27 +02:00
Nicolò Boschi df8ac42b52 fix(memory): add resolve_entities flag to update_memory and retain (#3576)
* fix(curation): resolve edited entity names exactly, not fuzzily (#3479)

update_memory ran the caller's entity names through the same fuzzy resolver
retain uses, so an entity name was a *guess* to be reconciled against the graph
rather than an instruction. Name identity is worth at most 0.5 of the 0.6 match
threshold, while co-occurrence (0.3) and recency (0.2) make up the rest, so a
similar-but-wrong entity that is well connected to the other names in the same
edit outscores the one the caller actually named — with a 200 and no warning.

Curation now resolves exactly: an existing entity is reused only when its
canonical name matches case-insensitively, any other name creates its own
entity, and same-batch names are never merged with each other. Retain keeps
fuzzy resolution, which is right for names that came out of extraction.

The exact path skips the trigram/UTL_MATCH probe and the co-occurrence fetch
entirely and reuses the existing find-or-create pass, so it is dialect-agnostic
and strictly less work than the fuzzy one.

* fix(curation): add entity_resolution_mode to update_memory, default fuzzy

Make the exact/fuzzy choice the caller's, rather than changing what an edit
does. `entity_resolution_mode` defaults to "fuzzy" — retain's behaviour, so
every existing caller is unaffected — and "exact" opts into literal matching for
hand-authored corrections.

Plumbed through the HTTP request model, the MCP tool, the control-plane proxy
route and its client, with the engine validating the value for direct callers.
The control-plane memory editor sends "exact": a person typing an entity list
into the admin UI is naming the entity they mean.

* refactor(curation): make the flag a boolean, resolve_entities

Replaces the entity_resolution_mode enum with a plain boolean on update_memory.
`resolve_entities` defaults to True — retain's behaviour, so existing callers are
unaffected — and False takes the submitted names literally.

Carries the same change through the resolver, which now takes `fuzzy_matching`
rather than a mode string. The engine's value guard goes away with the enum: a
bool needs no validation, so the invalid-value test goes too (the HTTP boundary
still 422s a non-boolean, which the HTTP test covers).

* feat(retain): honour resolve_entities for caller-supplied entities too

Same flag, same default, on the retain item — a caller passing explicit entity
names there has the same exposure as one correcting a memory: a name close to an
existing entity can be matched onto it and quietly replaced.

Retain resolves caller-supplied and LLM-extracted names in ONE batch, so a
per-batch flag would have turned resolution off for the extractor's names too and
filled the bank with near-duplicate entities. The flag is therefore carried per
mention: extracted names always resolve, supplied names follow the item's flag,
and a supplied name the extractor also produced keeps the caller's intent. The
in-batch dedup pass (#3107) skips the literal names for the same reason.

Also renames the resolver's `fuzzy_matching` parameter to the per-mention
`resolve` key, so one name is used end to end.

* fix(clients): carry resolve_entities through the maintained wrappers

Code review found two gaps the generated SDKs hide.

The TypeScript and Python convenience wrappers rebuild each retain item field by
field, so `resolve_entities` was silently dropped for every wrapper caller — the
same class of gap #2975/#3042 closed for the mental-model methods, and here it
would have quietly restored the substitution the flag prevents. Both wrappers now
forward it, with mapping tests on each side.

Intake also lost the flag when normalization collapsed two spellings into one:
entity_processing dedups caller-supplied against extracted names on the RAW text,
so a caller's literal "Acme Corp" and the extractor's "Acme\nCorp" both reach
_prepare_entities_for_resolution and only merge there. Keeping the first entry
verbatim dropped the caller's resolve=False with it; the merge now keeps the
stricter flag.

* fix: build the Rust CLI, and keep pg_trgm detection on empty batches

Two CI breaks from the retain change.

MemoryItem gained a field, and the CLI builds it with a struct literal, so every
Rust job failed to compile. The CLI supplies no entities, so `true` (the server
default) is the right value there.

The skip-the-probe guard also fired on an *empty* batch — `any([])` is False — so
_resolve_entities_batch_impl returned before the pg_trgm auto-detection that
hangs off the strategy dispatch. Only shortcut when there is data and none of it
resolves.

* fix(rust): add resolve_entities to the remaining MemoryItem literals

The first pass only fixed the CLI's src/ literal — `cargo build` does not compile
test targets, so the ones in hindsight-cli/tests/integration_test.rs and
hindsight-clients/rust/src/lib.rs went unnoticed until CI. Verified with
`cargo check --all-targets` in both crates this time.
2026-08-18 16:29:33 +02:00
Nicolò Boschi b64943d195 fix(knowledge-base): patch a page's trigger instead of replacing it, on create and update (#3549)
* fix(knowledge-base): merge a client's page trigger over the page defaults

Creating a knowledge page with ANY trigger silently discarded every
knowledge-page default. Two things combined to do it: the endpoint dumped
the whole request model (`model_dump()` fills every unset field with
MentalModelTrigger's own defaults -- mode="full",
exclude_mental_models=False), and the engine then replaced
KNOWLEDGE_PAGE_DEFAULT_TRIGGER with that dict wholesale.

So a client that wanted one setting -- different fact types, a cron
schedule -- also gave up `mode: "delta"` and `exclude_mental_models`, and
its page became a from-scratch rebuild that reflected over its sibling
pages on every refresh. That is what the coding-agents plugin had been
doing to every page it created (#3506), and there was no way for it not
to: the API offered no partial override.

The endpoint now forwards only the fields the client actually set
(`exclude_unset=True`) and `_merge_page_trigger` layers them over the
defaults -- which is what the engine's docstring already promised.

The two refresh triggers stay mutually exclusive: a client asking for a
cron schedule drops the default's `refresh_after_consolidation` rather
than inheriting a pair `MentalModelTrigger` would have rejected outright.

* feat(knowledge-base): let a page's refresh trigger be updated, as a patch

The trigger was write-once through the knowledge-base API: `UpdateNodeRequest`
carried no `trigger` field and the handler filtered to
`{source_query, tags, max_tokens}`, so a page created with one policy was
stuck with it -- including every page created before the create-path fix
above, which is still doing full from-scratch rebuilds. The engine already
supported it end to end; only the HTTP surface didn't expose it.

Both endpoints now behave the same way: send the fields you want changed,
keep the rest. Create patches over KNOWLEDGE_PAGE_DEFAULT_TRIGGER, update
patches over the page's CURRENT trigger -- which matters, because
`update_mental_model` overwrites the whole trigger column, so forwarding a
partial one straight through would have reintroduced the create-path defect
one endpoint over.

Exclusivity holds in both directions on update: moving a page onto a cron
schedule clears the auto-refresh it was created with, and moving it back
clears the cron. Neither pair is expressible in a request, so neither
should be reachable by merging.

The hand-written TS and Python wrappers both take the new parameter (they
are the surface most consumers actually call), each with a mapping test.
OpenAPI spec, generated clients and the docs skill regenerated.

* fix(cli): carry the new page trigger field through the Rust CLI

`types::UpdateNodeRequest` is generated from the OpenAPI spec, so adding
`trigger` to it broke the CLI's struct literal (and with it test-rust-cli
and the Windows embed build, which builds the CLI).

The field is passed as None -- omitted, it leaves the page's current
trigger alone -- and recorded in .openapi-coverage.toml with the reason,
alongside the same exemption the mental-model commands carry.

* fix(control-plane): expose the page trigger on the typed node-update client

The proxy route forwards the PATCH body verbatim, so the field already
reaches the dataplane -- but lib/api.ts enumerates the body fields, so
`trigger` was unreachable from any typed caller in the control plane.
2026-08-17 22:05:24 +02:00
Nicolò Boschi 4278f0989d feat(reflect,mental-models): surface structured output in the control plane (#3113)
* feat(reflect,mental-models): surface structured output in control plane

Reflect's response_schema -> structured_output was already implemented and
tested in the engine but never exposed in the UI. Surface it in the reflect
(think) view, and extend the same structured-output extraction to mental
models via a per-model response_schema stored in the trigger config.

- engine: refresh_mental_model reads trigger.response_schema, forwards it to
  the internal reflect call, and persists the parsed structured_output onto
  the stored reflect_response payload; fix stale 'not yet supported' docstrings
- api: add response_schema to MentalModelTrigger
- control-plane: reflect route + api.ts forward response_schema; think-view
  gets a JSON-schema input and renders structured_output; create/update mental
  model dialogs get a schema editor; detail modal renders structured_output
- tests: mental model structured-output plumbing (schema forwarded + persisted)
- regenerate OpenAPI spec + client SDKs; add i18n keys for all locales

* feat(control-plane): show configured response_schema in mental model config tab

Adds a read-only JSON card for the mental model's trigger.response_schema in
the detail modal's Configuration tab (mirrors the tag_groups card), plus the
regenerated go openapi.yaml.

* style: ruff format test_mental_model_structured_output

* fix(control-plane): don't route the JSON schema example through next-intl

The response_schema placeholder was a t() message whose value is literal JSON.
next-intl parses messages as ICU, so the '{' in the example was read as an
argument placeholder, the parse failed, and the field rendered the raw message
key instead of the example. Inline the JSON example directly on the placeholder
prop (i18n:check skips JSON-shaped placeholders) and drop the now-unused
*Placeholder message keys. Caught by running the control plane.

* feat(structured-output): validate response_schema + add a no-code schema builder

Validation (both reflect and mental models): a schema that is valid JSON but
not a usable object-with-properties silently produced empty structured_output
or blew up inside the LLM extraction call later. Now:
- engine: validate_response_schema() enforces the usable-shape contract
  (object schema, non-empty properties, well-formed required); wired as Pydantic
  field_validators on ReflectRequest.response_schema and
  MentalModelTrigger.response_schema (invalid -> HTTP 422).
- control-plane: the reflect and mental-model forms validate the schema shape on
  submit (not just JSON.parse) and surface the specific error.

No-code schema builder: a 'Build schema' button on both the reflect view and the
mental-model dialogs opens a dialog with Visual and Code modes. Visual mode edits
a flat field list (name, type, array item-type, description, required); Code mode
edits raw JSON. The two stay in sync and Apply is gated on a usable schema. Shared
frontend lib (response-schema.ts) mirrors the backend contract.

tests: test_response_schema_validation.py (16 cases: validator + model integration).

* refactor(control-plane): schema only via the builder, show set/unset status

Removes the inline response_schema JSON textarea from the reflect view and the
mental-model dialogs. Editing now happens exclusively in the schema builder; the
page shows only whether a schema is set (field count + names, with Edit/Remove)
or a Build schema button when none. Extracts the shared ResponseSchemaField
component used identically by reflect and both mental-model dialogs.

* fix(mental-models): derive structured_output from final content, not reflect's answer

In delta mode reflect only sees facts created since the last refresh, so its
answer (and any structured_output it derived) reflects just the delta — while the
stored content is the delta-merged document. Persisting the reflect-derived value
made structured_output inconsistent with the markdown.

Now the mental-model refresh no longer passes response_schema to reflect; instead
it extracts structured_output from the FINAL stored content (correct for both full
and delta), and carries the previous value forward untouched when a delta refresh
preserves content (no new facts). Adds a delta test asserting extraction runs
against the merged document, not reflect's partial answer.

* fix(schema-builder): allow switching an empty schema from Code back to Visual

An empty schema serialises to properties:{}, which schemaToFields mapped to an
empty array — and the Code->Visual guard treated 'empty' the same as 'not
representable', blocking the switch. schemaToFields now returns [] (representable)
for a missing/empty properties map and null only for genuinely unrepresentable
schemas; the switch seeds a blank field when empty.

* docs(reflect): document structured output (response_schema) + schema builder

Adds a Structured Output section to the reflect docs: how response_schema returns
both text and a structured_output projection of the same answer, the schema rules,
mental-model structured output (extracted from the final/merged document), and the
no-code Build schema editor. Regenerates the docs-skill mirror.

* feat(schema-builder): recursive visual editor for nested objects & arrays

The visual editor was flat — object/array fields had no way to define their inner
shape. Reworks the field model into a recursive tree (each field has a node; an
object node nests fields, an array node nests an item node) so you can build
nested objects and arrays-of-objects entirely in the visual editor. Code<->Visual
round-trips losslessly; schemas using features the editor can't represent (enum,
oneOf, $ref, tuple items, …) stay in code mode rather than being silently
flattened.

* fix(structured-output): recursive model for nested schemas + fail refresh loudly

Two problems surfaced by nested schemas on Gemini:

1. _generate_structured_output mapped object/array properties to bare dict/list,
   which serialize with additionalProperties — rejected by the Gemini API. So any
   schema with a nested object/array silently failed extraction. Now it builds a
   proper recursive Pydantic model (nested objects -> nested models, arrays ->
   typed lists), matching how retain's structured output already works on Gemini.

2. On extraction failure the mental-model refresh silently persisted content with
   no structured_output, clobbering the previously-stored value. Now, when a
   response_schema is configured and extraction yields nothing, the refresh raises
   MentalModelRefreshError — prior content and structured_output are preserved and
   the refresh can be retried.

Verified live on Gemini: a nested {location:object, people:array} schema now
extracts (structured_output present) instead of failing on additionalProperties.
Adds a fail-loud regression test.

* fix(schema-builder): readable error text in dark mode

text-destructive resolves to a dark red (#C0183A) in dark mode, which is
low-contrast on the dark dialog background. Use the codebase's standard
readable pattern (text-red-600 dark:text-red-400) for the builder's validation
error and the invalid-schema notice.

* fix(cli): set response_schema on MentalModelTriggerInput literals

Adding response_schema to MentalModelTrigger regenerated the Rust
MentalModelTriggerInput struct with a new field; the hand-written CLI struct
literals must initialize it (E0063). Sets response_schema: None in the three
construction sites (create/update mental model, knowledge-base pin).

* docs(api): document mental-model response_schema; fix stale reflect text-empty claim; test schema lib

- api/mental-models: document the trigger.response_schema flag + a Structured
  Output section (extraction from final content, fail-loud, validation).
- api/reflect: correct the stale claim that text is empty with response_schema —
  reflect returns both text and structured_output.
- control-plane: vitest unit tests for the response-schema lib (validation +
  recursive fields<->schema round-trip).
- regenerate docs-skill mirror.

* chore: regenerate bank-template-schema for MentalModelTrigger.response_schema

The bank template schema embeds MentalModelTrigger; adding response_schema to
the trigger changed the generated schema. Regenerated so verify-generated-files
passes.
2026-08-03 16:51:38 +02:00
Nicolò Boschi 06e9c7054e feat(mental-models): dry-run refresh and keep_trace for troubleshooting (#3119)
When a refresh produced an unexpected document, nothing said why. The mode
decision, resolved scope, snapshot window, retrieved-versus-used fact counts
and dropped delta operations only ever reached a log line — and cron- or
consolidation-driven refreshes run with nobody watching.

Two ways to see that reasoning, from opposite directions.

POST /mental-models/{id}/dry-run-refresh runs the production refresh
pipeline and reports what it would do, skipping exactly two writes: the
content (with its structured document and history entry) and the watermark
that moves last_refreshed_at. It takes no parameters, on purpose — a dry run
you can configure stops predicting the refresh it exists to predict. Because
nothing is persisted, a delta dry run reads exactly the window the next real
refresh will.

trigger.keep_trace records the same reasoning on every refresh of a model,
scheduled ones included, under reflect_response.trace. It is written even
when a refresh fails, which is when it matters most. The trace is shaped
like reflect's — the calls the agent made plus the refresh decision — and
holds nothing derivable from elsewhere: evidence stays in based_on, and the
resolved scope and window are reported by the dry run. Each tool call
records the window bound it was given, named `updated_at` for what the
predicate actually filters; null means the tool applies no time bound at
all, which is what explains results older than the window would suggest.

refresh_mental_model is split into a shared _execute_mental_model_refresh
that computes a result and writes nothing, plus a thin persistence step, so
the preview and the real refresh run the same body. Existing refresh
behaviour is unchanged.

In the control plane the dry run is an action on the mental model, and its
result opens in a dialog built from the History tab's own diff components.
History shows each version's own trace: the history snapshot now carries
`trace` alongside `based_on` so it survives being superseded.

Surfaced but deliberately not fixed here: when delta operations fail, the
fallback writes a candidate built from a delta-scoped recall over the whole
document, dropping content grounded in older memories (#3112).
2026-08-03 15:56:28 +02:00
Nicolò Boschi 4d22a882f5 docs(knowledge-pages): document Knowledge Pages and Mental Models, and manage them from the CLI (#3151)
* docs(knowledge-pages): document Knowledge Pages and Mental Models

Knowledge Pages shipped in #2455 with no documentation at all — no
architecture page, no API page, no mention in the sidebar. Mental models
had an API page but nothing explaining what they are or why they are
fast. Add both, as top-level entries under Architecture and API.

- Architecture: how pages are mental models with a simplified,
  document-shaped configuration; the folder hierarchy; the `hindsight fs`
  filesystem projection; page-level search; and why a projected view over
  reconciled memory is not the same thing as a folder of raw files.
- Architecture: mental models as standing answers built in the background,
  so an application reads the current version instead of paying for
  synthesis on the request path.
- API: the full knowledge-base endpoint surface, the page defaults and
  what each one buys, staleness gating, what a refresh reads, and how
  delta mode edits a structured document instead of regenerating prose.
  The mental-model trigger table gains the seven settings it was missing.
- FAQ: mental model vs knowledge page. Also corrects the neighbouring
  answer, which described mental models as built automatically during
  retain — that is observations.

The API examples use the maintained clients like every other API page, so
this adds the knowledge-base surface to the Python and TypeScript wrappers
(kept at parity, with request-mapping tests on both sides) and runnable
Python/Node/Go examples.

* feat(cli): manage knowledge pages from the CLI

The knowledge base was reachable from every client except the CLI, where
the eight endpoints were listed as deliberate coverage skips ("managed in
the control plane UI"). That left `hindsight fs` able to mirror pages
read-only but nothing able to create, edit, search, or delete them — and
it meant the API docs could not show a CLI tab alongside Python/Node/Go.

Adds `hindsight knowledge-base` with tree, create-folder, create-page,
get-page, search, update, delete, and export, removing the skips so
cli-coverage-check enforces the surface from here on.

`create-page` sends no trigger unless --mode or --fact-types is passed, so
the server's page defaults stand; when either is given the whole trigger
has to be restated, because a supplied trigger replaces the defaults
rather than merging with them.

Also adds the CLI tab to the Knowledge Pages API page and a Knowledge Base
section to the CLI reference.
2026-08-03 14:53:57 +02:00
Nicolò Boschi 218e6d34b1 feat(knowledge-base): client-managed knowledge pages, control-plane UI + hindsight fs CLI (#2455)
* feat(knowledge-base): self-curating knowledge base (OKF pages + folder missions)

Server-side knowledge base: a hierarchy of folders and pages over mental
models, projected to the Open Knowledge Format, with a mission-driven curator
that maintains pages automatically after each consolidation.

- knowledge_pages table (PG + Oracle): parent_id tree, kind folder/page,
  mission, managed, last_curated_at; partial unique index on (folder, name)
  for concurrency-safe dedup; added to BACKUP_TABLES.
- api/okf.py: OKF serializer (frontmatter + body, index/log, constellation graph).
- engine/knowledge_curator.py: folder curator (LLM op plan + safe apply); reads
  new memories since last curation (delta, not recall); ops create/merge/delete
  page + spawn sub-folder (bounded depth<=3, <=8). Runs as an async curate_folder
  task on folder/mission create and after consolidation. Curator pages use an
  observation-only delta trigger with exclude_mental_models.
- MemoryEngine: folder/page CRUD, tree, curate, async submit + worker handler.
- /v1/default/banks/{bank}/knowledge-base/* endpoints.
- Control plane: knowledge-base tree view + constellation toggle, missions,
  OKF page panel + bundle export; proxies, client, sidebar, i18n.
- Tests: okf unit, knowledge-base HTTP, curator apply + dedup guard, hs_llm_core e2e.
- Regenerated OpenAPI + SDK clients + docs-skill.

* feat(hindsight-fs): mirror a bank's mental models as a live local folder

Add @vectorize-io/hindsight-fs, a CLI under hindsight-tools/ that mirrors a
Hindsight bank's mental models as real markdown files (YAML frontmatter + body)
in a local directory, refreshed from the API on an interval. Once mounted,
ordinary shell tools (ls, cat, grep, find, ...) work against current memory.

- Pull-based sync engine: full list each tick, write changed/new/tampered
  files, skip unchanged (content-hashed), prune deleted models. Atomic writes;
  a transient API error never wipes the mirror.
- One-way mirror enforced two ways: files are read-only (0444) so agent edits
  fail with EACCES, plus a tamper-revert backstop that compares on-disk bytes
  and overwrites drift on the next pass. --writable opts out.
- Commands: mount/start/stop/restart/sync/status/list/logs/unmount. Background
  daemon via detached process + pidfile; per-mount config is remembered.
- status doubles as a healthcheck: --json report and a non-zero exit when the
  mount is dead/failed/stale (--stale-after overrides the threshold).
- Tests: unit (sync engine, frontmatter, health) + e2e that spawns the real
  CLI against a mock API and exercises real bash commands. 26 tests.

* refactor(hindsight-fs): mirror the knowledge-base tree, not mental models

Re-point hindsight-fs at the knowledge base so it projects a bank's folder/page
hierarchy as nested directories + .md files, instead of a flat list of mental
models.

- client: fetch GET /knowledge-base/tree + /export (two calls, any bank size)
  and join by page id; replaces the paginated mental-models list.
- format: planMirror() walks the tree into folder dirs + page files at nested
  paths (slug per segment, collision-safe); pages render the page's OKF doc.
- sync: create folder dirs, write pages at nested paths, prune removed pages and
  emptied folders; state keyed by relative path + tracked dirs.
- config/cli: drop the mental-model `detail` flag; `list` prints folders+pages;
  help/README updated. Tests rewritten for the tree/export model.

Verified live against a bank's knowledge base: the `people` folder mirrors to
people/anna.md + people/marco.md with OKF frontmatter.

* refactor(knowledge-base): drop server-side curation + folder missions

The knowledge base is now purely client-managed (CRUD over folders/pages); the
server no longer auto-curates. Removes the folder curator entirely and the
folder `mission` concept, and leads the sidebar with Knowledge Base.

- Remove engine/knowledge_curator.py, the curate_folder task (handler + dispatch
  + submit_async_curate_folder / _bank_folders), the post-consolidation curation
  hook, and the folder-create / mission-update curation triggers.
- Remove folder `mission` and `last_curated_at` (columns + engine + API + UI);
  keep `managed` as a client-set flag. Migration a5b6 now adds `managed` only;
  the last_curated_at migration is dropped and the unique-index migration
  repointed. Single alembic head preserved.
- API: KnowledgeNode/CreateFolderRequest/UpdateNodeRequest lose `mission`;
  PATCH node handles name/parent_id only.
- Control plane: sidebar leads with Knowledge Base (before Memories); remove the
  mission field, edit-mission dialog, and mission display from the KB view.
- Delete the curator tests; regenerate OpenAPI + SDK clients.

* feat(knowledge-base): default pages to living-document trigger + 4096 tokens

Client-created pages had no server curation applying a trigger, so they fell back
to the plain mental-model default (no refresh, full mode, all fact types). Make a
knowledge page a living document by default: when the client omits `trigger`, use
observation-only + delta + exclude_mental_models + refresh_after_consolidation;
when it omits `max_tokens`, default to 4096 (vs the mental-model 2048). Clients
can still override either.

* feat(control-plane): bank Home dashboard, Knowledge tabs, Notion editor + memory Euler graph

A large control-plane pass on the knowledge base UX:

- Home dashboard (home-view): memory constellation + read-only knowledge-page
  TOC (reuses the Pages tree) + recent documents + the bank-profile "Memory store"
  card and "Memories by ingested time" chart (extracted as reusable exports).
  Fixed-height top row so the constellation fills and the side cards scroll.
- Sidebar: add Home (first); order Home → Memories → Knowledge.
- Knowledge view: Pages / Mental Models sub-tabs (Mental Models moved out of the
  Memories view). Pages tab is an Obsidian-style workspace — file-tree sidebar +
  inline editor with open-page tabs; the generation prompt is tucked behind a
  "How this page is derived" expander; a "Backed by N memories" line opens the
  backing model's based_on via the existing mental-model detail modal. First page
  auto-opens; deep-link via ?page=.
- Knowledge graph: reframed as an Euler/Venn of the source memories — nodes are
  based_on memories, one translucent circle per page (overlaps = shared memories),
  plus the memory graph's own edges. New Constellation venn mode (nodeGroupsFn /
  groupColorFn / groupLabelFn) drawing overlapping per-group circles + pill labels.
  Backend: knowledge_page_memory_graph endpoint (pages' based_on → memory nodes).
- Pages default to the living-document trigger (observation-only, delta, exclude
  mental models, auto-refresh) + 4096 max_tokens.
- Documents: metadata badges are expandable (show all keys, not just 3).
- Regenerated OpenAPI + docs-skill.

* feat(cli): port hindsight-fs into the Rust CLI as `hindsight fs`

Rewrite the standalone TypeScript hindsight-fs tool as a native subcommand
of hindsight-cli. Mirrors a bank's knowledge base (folders + pages) to a
local folder of markdown files, one-way (API -> disk) with read-only files
and drift-revert, plus a detached background refresh daemon.

- new src/commands/fs/ module (client, format, sync, state, daemon,
  health, config, paths) with unit tests for the pure logic
- subcommands: mount/start/stop/restart/sync/status/list/logs/unmount
- deps: reqwest blocking + sha2 + libc
- remove the TS package and its npm workspace entry

* feat(knowledge-base): drop the pages Graph view + polish the Pages UX

Client:
- Remove the Tree/Graph toggle and the whole graph branch (the toggle row
  was the dead band between the tabs and the content).
- Tree rows: full-width page name + compact status dot (was a pill that
  crushed the name to a few chars); float the hover actions so they no
  longer reserve width; widen the sidebar (w-64 -> w-72).
- Borderless workspace card; solid sticky editor-tab bar (was bg-muted/20,
  so scrolled body text bled through); tighter sub-tab spacing.
- Drop getKnowledgeBaseGraph + the /api/knowledge-base/graph proxy route
  and the now-unused i18n keys (viewTree/viewGraph/graphEmpty).

Server:
- Remove GET /knowledge-base/graph, KnowledgePageGraphResponse, the
  knowledge_page_memory_graph engine method + its helper, and the orphaned
  KnowledgeGraph/palette code in okf.py.
- Regenerate OpenAPI + SDK clients + docs-skill.

* feat(knowledge-base): add hybrid GET /knowledge-base/search (BM25 + vector)

Doc-level hybrid search over a bank's knowledge pages, fused with Reciprocal
Rank Fusion in a single query — no reranker, tuned for latency (~sub-100ms
query on top of the embed).

- engine.search_knowledge_pages: vector arm (mm.embedding ANN) + BM25 arm
  (mm.search_vector, the generated tsvector over page name+content) via
  websearch_to_tsquery('english'), RRF-fused (k=60) in SQL. BM25-only
  fallback when the query embedding is unavailable. Folders excluded.
- GET /v1/default/banks/{bank}/knowledge-base/search?q=&limit= with
  KnowledgePageSearchResult/Response models (id, name, mental_model_id,
  snippet, score, updated_at).
- tests: ranks the relevant page first, excludes folders, respects limit,
  requires q.
- Regenerate OpenAPI + Python/TS/Go clients + docs-skill (Rust client has
  no KB ops; the CLI's fs port uses raw reqwest there).

* feat(control-plane): wire knowledge-page hybrid search into the Pages sidebar

A debounced search box at the top of the Knowledge sidebar queries the new
/knowledge-base/search endpoint (BM25 + vector, RRF-fused). A non-empty query
swaps the folder tree for a ranked result list (name + snippet); clicking a hit
opens the page in a tab. Clear (×) restores the tree.

- new /api/knowledge-base/search proxy route
- client.searchKnowledgePages(bankId, q, limit) in lib/api.ts
- search box + results list in knowledge-base-view.tsx (reuses openPage)
- i18n: searchPlaceholder / clearSearch / searchEmpty + api.errors.knowledgeBase.search across all locales

* feat(control-plane): widen the Knowledge tree to 1/3 + show page tags inline

- file tree pane w-72 -> w-1/3 (content gets the other 2/3)
- render each page's tags as small chips under its timestamp in the tree

* feat(control-plane): let new pages set tags in the create dialog

The create-page dialog gains a comma-separated Tags field (wired to the
existing createKnowledgePage tags param); a type:<x> tag sets the page's
OKF type. i18n added across all locales.

* perf(control-plane): stop the home dashboard blocking on a 22MB graph

The memory constellation fetched limit=1000 (≈67k edges, ~22MB) and the whole
dashboard awaited it before painting. Now the light panels (stats/pages/docs)
render immediately and the constellation loads on its own (with a spinner) at a
200-node cap (~3.7MB). The full graph stays in the Memories view.

* feat(control-plane): add a 'View all' button to the home memory card

The memory-constellation card header gets a 'View all →' link to the Memories
(data) view, matching the Knowledge pages / Recent documents cards.

* feat(knowledge-base): edit a page's source query, tags, and token budget

PATCH /knowledge-base/nodes/{id} now also updates a page's options on its
backing mental model — source_query, tags, max_tokens — each applied only when
present. Changing source_query schedules an async refresh so the page rebuilds
against the new question.

- engine.update_knowledge_page + UpdateNodeRequest fields + endpoint wiring
- control plane: an "Edit" button on the open page opens a dialog (name /
  source query / tags); tags pre-fill from the raw tree tags so the type:<x>
  tag isn't dropped on save. i18n across all locales.
- tests: update options persist; empty PATCH is 400.
- regenerated OpenAPI + clients + docs-skill.

* refactor(knowledge-base): squash page migrations into one + drop OKF naming

- Fold the three knowledge_pages migrations (table / managed column / unique
  page-name index) into a single a9b8c7d6e5f4 migration.
- Rename api/okf.py -> api/page_markdown.py and replace the "OKF" / "Open
  Knowledge Format" terminology throughout (server, control plane, CLI, i18n)
  with plain "markdown" / "page" wording — pages just render to markdown.
- Add the knowledge-base endpoints to the CLI OpenAPI-coverage skip list (they
  live in the control plane UI / are mirrored by `hindsight fs`, no CLI cmd).
- Regenerate OpenAPI + clients + docs-skill.

* test(knowledge-base): rename test_okf -> test_page_markdown, drop dead graph tests

Follow-up to the okf.py rename + Graph-view removal: the test module still
imported the old `okf` module (breaking collection for the whole api test
suite) and still tested the removed tag-based knowledge_graph() builder.

* fix(transfer): classify knowledge_pages in export-bank (skip for now)

test_export_bank_covers_schema requires every BACKUP_TABLES entry to be
classified by export-bank. knowledge_pages is skipped: carrying its self-
referential parent_id tree needs a parents-first restore order that the generic
per-row _restore_rows doesn't provide (follow-up). Mental models are carried, so
the target can recreate the tree.
2026-07-30 16:02:41 +02:00
Nicolò Boschi 6a460d2c9a feat(reflect): add apply_all_directives to bypass directive tag scoping (#3031) (#3046)
* feat(reflect): add apply_all_directives to bypass directive tag scoping (#3031)

Directives are tag-scoped like memories: a reflect with no tags loads only
untagged directives, and tagged directives apply only when the request's tags
match. This is deliberate (isolation_mode), but it means an operator's
tag-organized directives silently never reach an untagged reflect — 45% of
standing rules in the deployment reported in #3031.

Add an opt-in `apply_all_directives` flag on the reflect request (default
false, preserving current behavior). When true, every active directive is
loaded regardless of tags, ignoring tag scope. Wired through the HTTP API,
both MCP reflect variants, and the engine.

Also correct the docs, which claimed directives are "always" enforced without
mentioning tag scoping.

Regenerated OpenAPI, clients (Go/Python/TS/Rust), and the docs skill mirror;
updated the control-plane reflect proxy + api.ts types.

* chore(cli): record apply_all_directives CLI-coverage exemption

The reflect field is intentionally not exposed as a CLI flag (available via
the REST API, SDKs, and control plane). Record the exemption so cli-coverage-check
passes, matching the existing tag_groups entry.
2026-07-29 15:39:50 +02:00
Sanderhoff-alt feac397324 chore(repo): remove unused code (#3007)
* chore(api): remove unused code

Remove confirmed unreferenced helpers from the API and engine.

Delete tests only where they cover superseded internal paths. Keep
active test helpers and public memory operations unchanged.

* chore(cli): remove unused code

Remove dead CLI configuration, client, and output helpers.

Drop the parser implementation and tests used only by the retired
output path.

* chore(control-plane): remove unused code

Remove unused ControlPlaneClient methods and unreachable directive
detail state from the think view.

* chore(dev): remove unused code

Remove unreferenced benchmark and repository maintenance helpers.

* chore(embed): remove unused code

Remove the unused daemon port lookup helper while preserving current
profile-based daemon discovery.

* chore(integrations): remove unused code

Remove unreferenced helpers across supported integrations.

Drop tests only for retired internal paths and retain active test and
lifecycle infrastructure.

* test(consolidation): port prompt regression tests to split builders

The dead-code cleanup removed build_batch_consolidation_prompt and its tests,
but those tests guarded behaviors that are still live in the current
build_consolidation_system_prompt / build_consolidation_input path:

- brace-safety of a mission / capacity note containing literal { } (a lone
  brace would raise KeyError in the internal str.format() and crash
  consolidation)
- output-language directive injection into the cached system prompt
- the built-in default mission when none is supplied

Re-add these as regression tests against the current builders instead of
dropping the coverage. Also fix a stale comment referencing the removed
utils.extract_facts module.

---------

Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
2026-07-29 11:18:30 +02:00
Nicolò Boschi c41ad9bd75 feat(api): filter memory list by linked entity + entity timeline UI (#2945)
* feat(api): filter memory list by linked entity + entity timeline UI

Add an `entity_id` query param to `GET /memories/list` — an exact reverse
lookup over stored entity links (not text/semantic match), backed by the
existing idx_unit_entities_entity_unit index. Because entity links reference
live memory units only, combining `entity_id` with `state=invalidated`
returns nothing.

Wire it through the control-plane list route + clients, and use it in the
entity detail panel to render an observation timeline (reuses the memories
TimelineView) — click an entity, see its linked observations over time.

Closes #2936.

* fix(control-plane): entity timeline shows all linked memories, not just observations

Verified against real data: observations are derived/consolidated summaries and
carry no entity links — entity links live on the source world/experience facts,
which are also the ones with occurred dates. Filtering the entity timeline to
type=observation therefore always rendered an empty panel. Drop the type filter
so the panel shows every memory linked to the entity (the actual dated timeline),
and relabel the section "Timeline" with dedicated i18n keys.

* chore(control-plane): drop now-unused observation i18n keys from entitiesView

* style(reflect): wrap over-length _generate_structured_output call

Ruff format wraps this >120-char call; committing the formatter output so the
verify-generated-files CI check (which runs the formatter and diffs) is clean.
2026-07-24 14:28:44 +02:00
Nicolò Boschi 31218127e0 fix(retain): make async retries idempotent via caller-supplied operation_id (#2937) (#2947)
* fix(retain): make async retries idempotent via caller-supplied operation_id

An async retain whose HTTP acknowledgement is lost or times out leaves the
caller unable to tell whether the operation was created; retrying enqueues a
second parent operation and repeats extraction, embeddings, and provider spend.

Add an optional caller-supplied operation_id (UUID) used directly as the parent
async_operations primary key. Re-submitting with the same id returns the
original operation and creates no new work; the existing primary key is the
concurrency authority, so no new columns, constraints, or migration are needed.
Reusing an id owned by a different bank or operation type returns HTTP 409.
Omitting operation_id keeps the current create-each-time behavior.

Fixes #2937

* docs(retain): explain why the idempotency read is not in the create txn

* fix(retain): sync generated docs-skill + Rust clients for operation_id

- Regenerate the two docs-skill artifacts derived from the retain doc /
  OpenAPI change (verify-generated-files).
- Add operation_id: None to the Rust client test and CLI RetainRequest
  literals so both crates compile against the regenerated struct.
2026-07-24 13:43:57 +02:00
Parafee41 1ff09ccf9c fix(cli): preserve HTTP 400 details (#2916)
* fix(cli): preserve HTTP 400 details

* sync generated OpenAPI version
2026-07-24 12:25:39 +02:00
Sanderhoff-alt 91ee2537e2 fix(cli): sync regenerated OpenAPI operation changes (#2867)
* fix(cli): pass tag filters to list memories

OpenAPI added tags and tags_match to list_memories in #2848, but the
CLI wrapper still passed the previous positional arguments. Generated
Rust client builds then failed with E0061.

Pass None for both filters to preserve existing CLI behavior and match
the generated method signature.

* feat(cli): expose terminal operation deletion

OpenAPI added delete_operation in #2777 without exposing it through
the Rust CLI or accounting for it in the coverage manifest. The CLI
coverage check therefore rejected branches rebased onto that change.

Add operation delete with confirmation and --yes support. Pass the
request through the generated client and cover command parsing. This
counts the endpoint as implemented without a coverage exception.
2026-07-21 14:53:15 +02:00
Nicolò Boschi 41d71a9818 fix(#2808): make mental model tags_match configurable on all creation surfaces (MCP, TS client, CLI) (#2858)
* feat(mcp): let create_mental_model configure tags_match (#2808)

A tagged mental model with no explicit tags_match in its trigger JSON
refreshes under all_strict (a memory must carry every one of the model's
tags), while the staleness check and every recall/reflect path default to
any. Broadly-tagged models reading narrowly-tagged memories therefore get
marked stale and then refresh to empty content.

The HTTP API, generated SDK clients, and Control Plane UI already let users
set trigger.tags_match; the MCP create_mental_model tool did not. Add a
tags_match argument (validated against TagsMatch) to both MCP variants. It
is only written into the trigger when explicitly passed, so the resolved
all_strict default is preserved for existing callers.

Document the all_strict footgun and the tags_match override in the MCP and
mental-models API docs (regen skills/hindsight-docs mirror).

* fix(ts-client): expose tags_match/tag_groups on createMentalModel

The ergonomic TypeScript wrapper's createMentalModel accepted only
{ refreshAfterConsolidation } in its trigger option and dropped every other
trigger field, so a wrapper user could not set tags_match — the exact knob
needed to avoid the empty-refresh footgun in #2808. The low-level generated
sdk already accepts the full MentalModelTriggerInput; thread tagsMatch and
tagGroups through, mirroring how recall/reflect already expose them.

The Python client needs no change: its wrapper takes a pass-through
trigger dict and the generated MentalModelTriggerInput already validates
tags_match.

* test(ts-client): cover createMentalModel trigger mapping

Mock the generated sdk layer (no server needed) and assert the ergonomic
camelCase trigger options map onto the snake_case body: tagsMatch ->
tags_match, tagGroups -> tag_groups, refreshAfterConsolidation still maps,
and omitting trigger sends none (preserving the all_strict default). Locks
in the #2808 wrapper fix.

* docs(mental-models): add tags_match code snippet

Replace the static JSON block in the tags_match override section with a
live CodeSnippet pulled from the Python example, showing how to create a
model with trigger.tags_match="any" so a broadly-tagged model reads
narrowly-tagged memories on refresh (#2808).

* feat(cli): add --tags-match to mental-model create + all-language docs

The Rust CLI's `mental-model create` was the last creation surface with no
way to set tags_match, so a tagged model created via the CLI hit the same
empty-refresh footgun (#2808). Add a `--tags-match` flag (any/all/any_strict/
all_strict/exact) that is only sent when passed, preserving the server's
all_strict default; invalid values are rejected before the request.

Expand the mental-models docs "tags_match override" example from a single
Python snippet to a full Tabs block (Python / Node.js / CLI / Go), each
pulled from the runnable example files, and regen the skills mirror.
2026-07-21 12:04:23 +02:00
Parafee41 e839c65537 fix(cli): set default user agent (#2564) 2026-07-07 15:12:18 -04:00
Parafee41 f8ce15b9bf show full memory details in explorer (#2490) 2026-07-01 13:37:08 +02:00
Nicolò Boschi d68f618969 feat(stats): distributed bank_stats cache + ?refresh param + stats perf suite (#2495)
* feat(stats): distributed (table-backed) bank_stats cache on PostgreSQL

get_bank_stats aggregates over memory_links/unit_entities — a multi-second scan
on large banks. It was cached per-process (in-memory), so every API worker
recomputed once per TTL and the first caller after expiry stalled.

Add a bank_stats_cache table and a DistributedBankStatsCache that shares one
worker's computation across all workers. Same get_or_load/invalidate contract as
the in-memory cache, so the hot path is a single PK SELECT on a hit; only a miss
runs the existing _compute_bank_stats loader and UPSERTs the row (ON CONFLICT,
no lock — concurrent misses recompute, last write wins). All DB touches are
best-effort: an unreachable/missing cache table degrades to computing uncached
rather than failing the endpoint. PostgreSQL only; Oracle keeps the in-memory
cache (selected by dialect at construction).

* feat(stats): add ?refresh query param to force fresh /stats (default off)

Adds force_refresh to get_bank_stats (and both cache backends): when set, the
cached value is bypassed and recomputed, and the fresh result refreshes the
cache for subsequent callers. Exposed on GET /stats as ?refresh=true (default
false). Regenerated OpenAPI spec + clients.

* test(perf): add stats benchmark suite + huge prod-sim scale

New 'stats' perf suite measures get_bank_stats: uncached aggregation latency
(node/link counts + entity rollup) vs cached, run with the result cache disabled
so the headline numbers are the real per-poll cost. Adds a 'huge' prod-simulation
scale that bulk-loads ~500k units / ~17.8M physical memory_links via COPY (entity
links derived from unit_entities, not stored).

* test(stats): exclude bank_stats_cache from backup guard + HTTP refresh test

- bank_stats_cache is a derived TTL cache (no FK to banks, repopulates on
  demand), so exclude it from test_backup_tables_covers_entire_schema rather
  than back up stale cache rows — a restore starts it cold.
- Add a ?refresh=true assertion to the /stats HTTP integration test.

* fix(cli): pass refresh arg to get_agent_stats after ?refresh param

The new /stats ?refresh query param adds a positional arg to the progenitor-
generated get_agent_stats; the CLI reads the cached value, so pass None.
2026-07-01 13:35:54 +02:00
Parafee41 82afa76182 fix(cli): explore header selection (#2489) 2026-07-01 11:49:25 +02:00
Evo b0038e9855 cli: show operation filenames (#2435) 2026-06-29 12:05:05 +02:00
Nicolò Boschi 758f346d30 feat(recall): structured per-stage scores and two-level min_scores filtering (#2422)
Replace the recall result's single `score` with a `scores` object exposing the
scores from each pipeline stage, and replace the `min_score` request param with
`min_scores`, a per-stage filter that operates at two levels.

Response — each result carries `scores`:
- final     : the value results are ranked by
- reranker  : cross-encoder normalized relevance (null for passthrough rerankers)
- semantic  : raw vector cosine similarity (null if not surfaced semantically)
- text      : raw keyword/BM25 score (null if not surfaced by keyword search)

Per-arm semantic/text scores are aggregated across retrieval arms during RRF /
interleave fusion (ArmScores on MergedCandidate), since fusion otherwise keeps
only the first-seen arm's score per doc.

Request — `min_scores` floors (inclusive, AND-ed, opt-in; default no filtering):
- semantic / text : retrieval-level cutoffs pushed into the SQL arms, overriding
  the global similarity / BM25 minimums for the request (prune before fusion)
- reranker / final: post-query filters on the scored results

There is deliberately no default threshold: the cross-encoder's absolute scores
are reliable for ordering but not calibrated across queries (a clearly-relevant
match can score ~0.001 on one query and ~1.0 on another), so a fixed cutoff would
silently drop good results.

Also surfaces proof_norm in the search trace and reworks the control-plane trace
view to render scores at full precision (no rounding) and show the per-stage
`scores` breakdown; relabels the trace's "CE" column to "reranker score".

Threaded through engine, HTTP, MCP (both recall tools), and the control-plane
proxy; OpenAPI spec, Python/TS/Go/Rust clients, and the docs-skill mirror
regenerated; docs updated.
2026-06-26 17:12:27 +02:00
Nicolò Boschi 63a92bef5f fix(cli): pass u64 limit/offset to regenerated client (#2370)
The Rust client was regenerated with limit/offset typed as Option<u64>
(unsigned, minimum 0 in the OpenAPI spec), but the api.rs wrappers still
passed Option<i64>, breaking `cargo build` (and the test-rust-cli /
test-doc-examples CI jobs). Cast the values to u64 at each call site
(list_documents, list_memories, list_entities, get_graph, list_tags),
matching the existing pattern already used for list_documents.
2026-06-23 14:11:54 +02:00
Nicolò Boschi 5f0b715517 feat(recall): add prefer_observations to dedupe raw facts superseded by observations (#2311)
Recalling `observation` alongside `world`/`experience` can return the same
information twice — once as a raw fact and once folded into an observation
consolidated from it. The opt-in `prefer_observations` flag drops any raw fact
that a returned observation lists in its `source_memory_ids`, so the observation
supersedes it. Dedup is by provenance (exact id membership), not semantics, and
runs before recall truncation so freed slots backfill — keeping the result count
at the requested budget.

Disabled by default (opt-in). Internal callers — notably consolidation, which
needs the raw facts it folds into observations — leave it off.

Exposed on the full client surface: the maintained Python (`recall`/`arecall`)
and TypeScript (`recall`) wrappers, the Rust CLI (`--prefer-observations`), the
regenerated OpenAPI + low-level Python/TS/Go/Rust SDKs, the control-plane proxy +
types, and the generated docs skill. Includes docs and deterministic
provenance-based tests (engine + both wrappers).
2026-06-23 11:15:30 +02:00
Sanderhoff-alt 59d5319b84 feat(retain): make structured chunk size configurable (#2139)
Add retain_structured_chunk_size as an explicit retain chunking knob
for structured inputs. When unset, structured inputs follow the
effective retain_chunk_size instead of the hidden 1.5x overflow factor.

Thread the setting through retain extraction, append/prepend chunking,
bank config resolution, templates, MCP docs, maintained clients,
generated OpenAPI artifacts, the control-plane retain strategy UI, and
the Rust CLI set-config command.

Validate retain_chunk_size and retain_structured_chunk_size as positive
integers while allowing either value to be smaller. Keep the existing
retain_max_completion_tokens check scoped to retain_chunk_size.

Preserve upstream validation details for client errors through the
control-plane proxy so UI alerts and toasts can show concrete
configuration errors without exposing server-side failure details.

Update chunking, config, hierarchical config, template, MCP, client
payload, control-plane serialization, SDK-response, API-client, and
retain UI validation tests for the new behavior.
2026-06-15 15:51:51 +02:00
Nicolò Boschi 72d9881a6a fix(cli): parse get-memory response correctly (#2111)
The `hindsight memory get` command deserialized the API response into a
local `MemoryUnitDetail` struct whose shape had drifted from what
`GET /memories/{memory_id}` (MemoryEngine.get_memory_unit) actually
returns:

- `entities` is a flat list of canonical-name strings, but the struct
  expected a list of `{id, name}` objects, so serde failed whenever a
  memory had entities — surfaced to users as the misleading
  "Invalid API response format".
- the fact type is exposed as `type`, but the struct renamed it to
  `fact_type`, so the Type line always printed UNKNOWN.

The endpoint returns an untyped JSON body in the OpenAPI spec, so the
generated client never validates it and the mismatch only blew up in the
CLI handler. These commands had no test coverage (docs use curl).

Fix the struct to match the response and add regression tests.
2026-06-10 18:07:24 +02:00
Nicolò Boschi de22b606e7 feat(memory): reversible curation — edit/invalidate/revert memory units (#1976)
Edit (text/context/dates/fact_type/entities), invalidate (move to a separate
invalidated_memory_units archive, reversible), and revert raw memory units via
PATCH /memories/{id}. Tracks user edits with edited_at. Control-plane UI, docs
(Memories API page), and multi-language examples included. RFC #1951.
2026-06-10 17:20:59 +02:00
Nicolò Boschi 1296e9fc12 feat(api): config flag to skip storing raw document text (#2061) (#2062)
* feat(api): add HINDSIGHT_API_STORE_DOCUMENT_TEXT flag to skip raw text storage

When set to false, the retain pipeline runs unchanged (chunking, fact
extraction, embedding, entity linking) but drops the raw source text:
documents.original_text is stored as NULL and chunks.chunk_text as empty.
content_hash is still computed from the real text so delta-retain dedup is
unaffected, and recall is unaffected because it reads from memory_units.

Closes #2061

* feat(api): reject append + drop source-text reads in privacy mode

Follow-up to the HINDSIGHT_API_STORE_DOCUMENT_TEXT flag, covering the
features that read raw document/chunk text back:

- retain update_mode='append' is now rejected when text storage is
  disabled (it rebuilds the document from the stored original_text, which
  is NULL, and would silently drop prior content).
- reflect no longer offers the 'expand' tool (get chunk/document source),
  gated via get_reflect_tools(include_expand=...); the hallucination guards
  no longer hardcode 'expand' as always-allowed.
- reflect's recall step no longer attaches empty source chunks.

Other read sites (get-document/list-chunks/get-chunk endpoints + MCP,
export/import, public recall include_chunks) already degrade gracefully to
empty/None and are documented.

* fix(api): get-document 200 with null text + 400 on append in privacy mode

Caught while testing the flag live against a running server:

- DocumentResponse.original_text was a non-optional str, so the GET
  document endpoint raised ResponseValidationError -> HTTP 500 when the
  text is NULL. Made it str | None.
- The retain handler mapped all exceptions (including the append-rejection
  and duplicate-document_id ValueErrors) to HTTP 500. Map ValueError to 400,
  matching the convention used by the other endpoints.

Adds an HTTP-level regression test (the engine-level test passed because
get_document returns a dict, bypassing response-model validation).

* feat(ui): warn when document text storage is disabled

Surface the store_document_text flag so the control plane can warn users
that raw source text isn't persisted:

- /version feature flags now include store_document_text (regenerated
  OpenAPI spec + SDK clients; also picks up the earlier DocumentResponse
  original_text optional change).
- features-context exposes it (defaults true, so the warning only shows
  when the server explicitly reports privacy mode).
- Document detail dialog: the Content tab shows a small notice instead of
  an empty body when text isn't stored.
- Add Document dialog: a small notice that raw text won't be kept.
- i18n strings added across all locales.

Adds an API test asserting /version reports the flag.

* refactor: drop "privacy mode" wording; reposition document-text warnings

- Remove the "privacy mode" phrasing I had introduced from comments,
  docstrings, test names, docs, and UI labels. The flag is described by
  what it does (skip storing raw document text) instead.
- Add Document dialog: move the warning to just above the action buttons.
- Document dialog: also show the warning on the Chunks tab.

* fix(cli): handle optional document original_text

original_text is now Option<String> in the generated client (it can be
null when document text storage is disabled), so the CLI can't print it
with {} directly. Show "(not stored)" when absent.
2026-06-09 14:46:09 +02:00
Anton Evseev df73c7924e fix(cli): hindsight memory retain --timestamp + correct fact-type values (#1881)
Two unrelated CLI bugs surfaced during sandbox testing on 2026-05-31.

1) `hindsight memory retain --timestamp <ISO 8601>` never worked.

   `MemoryItem.timestamp` is generated from the OpenAPI schema
   `anyOf: [{type: string, format: date-time}, {type: string}]`. Progenitor
   emits that as a struct with two `#[serde(flatten)]` Option subtypes —
   which serde refuses to serialize for primitives:

     "can only flatten structs and maps (got a string)"

   So even constructing the value manually fails at serialize time, before
   the request hits the wire. The CLI's `serde_json::from_value::<…>(String)`
   round-trip also fails (struct deserializer expects an object).

   Fixed at the codegen boundary by adding a pre-codegen spec-massage step
   `collapse_string_anyof_unions` in hindsight-clients/rust/build.rs that
   collapses any `anyOf` whose members are all `{type: string}` into a
   single `{type: string}`. The `format: date-time` distinction is lossless
   on the wire — both serialize to the same string — so this is safe.
   Result: `MemoryItem.timestamp: Option<String>`, no broken type generated.

   The CLI no longer needs to round-trip through a wrapper type; the user
   string is passed through directly.

2) `hindsight memory clear --fact-type` rejected the valid value
   `observation` and accepted stale values `agent` / `opinion` that the
   server silently treats as no-ops.

   Help text on `bank graph`, `memory list`, `memory recall`, and
   `memory clear` referred to a non-existent fact type `opinion`. The
   canonical fact types per the API are `world | experience | observation`
   (see hindsight_api.api.http.MemoryItem and the `Literal[…]` arm on
   fact_types in recall/reflect requests).

   Fixed: `opinion` → `observation` everywhere in CLI help / clap defaults,
   and `agent`/`opinion` → `experience`/`observation` in the clear
   command's value_parser allow-list.

Regression test:
  hindsight-cli/tests/integration_test.rs::
    test_memory_item_timestamp_serializes_as_plain_string

Verified:
  - cargo build → clean
  - cargo test --bin hindsight → 55/55 pass
  - cargo test --test integration_test test_memory_item_timestamp_… → pass
  - cargo clippy → no new warnings (171 pre-existing uninlined_format_args)
  - hindsight memory clear --help → [possible values: world, experience, observation]
  - hindsight memory recall --help → [default: world experience observation]
  - hindsight bank graph --help → (world, experience, observation)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 10:02:32 +02:00
Nicolò Boschi 28ec22c3dc fix(ci): align config field count and CLI consolidation call with #1746 (#1757)
PR #1746 added enable_auto_consolidation to _CONFIGURABLE_FIELDS and
introduced a ConsolidationRequest body on the /consolidate endpoint, but
didn't update test_hierarchical_fields_categorization (still expects 35
fields) or the CLI's trigger_consolidation wrapper (still calls the
generated client with 2 args), so CI on this branch breaks on test-api,
test-rust-cli, test-embed-windows, and test-doc-examples (cli).

Bump the expected count to 36, add enable_auto_consolidation to the
explicit assertions, and pass a default ConsolidationRequest to the
generated client so the no-scope CLI invocation keeps consolidating all
unconsolidated memories.
2026-05-26 15:12:05 +02:00
Ben 51ea9aa286 fix(cli, control-plane): make retain Event Date / timestamp actually reach the API (#1622)
* fix(cli, control-plane): make Event Date / timestamp actually reach the API

- CLI `hindsight memory retain` now accepts `-t/--timestamp <ISO>`. The
  internal MemoryItem.timestamp was hardcoded to None, so retains from the
  CLI lost any caller-supplied event date even though the Python/Node/Go
  SDKs accept one. Add a flag and pass it through; regression test asserts
  --help advertises the option.
- Control plane "Event Date" inputs in the new-document and per-file flows
  used `<input type="datetime-local">`, which only commits a value when the
  user enters both date AND time. Typing a date alone silently left the
  value empty, so `item.timestamp` was never sent and the resulting
  operation payload had no event_date. Switch to `type="date"` and pad
  with `T00:00:00` before sending, so date-only entries reach the API as
  valid ISO datetimes.

* fix(cli): decode --timestamp into MemoryItemTimestamp enum

MemoryItem.timestamp is generated as Option<MemoryItemTimestamp>
(progenitor's anyOf wrapper), not Option<String>. Round-trip the
flag value through serde_json so the right variant is selected for
both ISO datetimes and the 'unset' sentinel. Fixes CI build break.
2026-05-13 16:51:21 -04:00
Nicolò Boschi 375747f516 feat(cli): add --strategy flag to memory retain-files (#1494)
Allows callers to pick a named retain strategy when bulk-importing files,
overriding the bank's default. The API already accepts a per-file strategy
in FileRetainMetadata; this just wires a CLI flag through to the multipart
metadata.

Closes #1492
2026-05-06 17:52:41 +02:00
Nicolò Boschi 8fbe85f0ca feat: mental-models List view + /tags?source=mental_models (#1296)
* feat(api): list mental-model tags via /tags?source=mental_models

Adds a `source` query param to GET /v1/default/banks/{bank_id}/tags so the
same endpoint can list tags from either memory_units (default) or
mental_models. Mental-model tag suggestions previously had no API; the
alternative of a sibling /mental-models/tags route would have shadowed
GET /mental-models/{mental_model_id} for the literal id "tags".

Engine: new list_mental_model_tags method sharing a private
_list_tags_from_table helper with the existing list_tags.

Tests: covers the engine method (basic counts, wildcard) and an HTTP-level
check that source=mental_models reads from mental_models while default
remains memory_units.

* feat(control-plane): mental-models List view with tag filter

Adds a default split-pane "List" view to the Mental Models page (sidebar of
files + content on the right) and a reusable <TagFilterInput> with free-text
entry, debounced suggestions from the server, and chip selection.

Changes:
- Default Mental Models view is "List" (file/folder metaphor); the existing
  card "Dashboard" view stays as a secondary toggle. Old "Table" view removed.
- Sidebar entries show name, source query subtitle, and relative refresh time.
- Tag filtering is server-side via the existing tags/tags_match params on
  /mental-models; suggestions populate from /tags?source=mental_models.
- Memories (data-view) reuse the same TagFilterInput, gaining suggestions
  it didn't have before.
- Adds proxy route for GET /tags (forwards optional source query param).
- TagFilterInput holds the caller's fetchSuggestions in a ref to keep the
  debounce effect from refiring on every render when callers pass an inline
  closure (which would otherwise loop).
2026-04-28 12:31:21 +02:00
Nicolò Boschi 13b1d92297 fix(docs): escape curly braces in generated changelog entries (#1248)
* fix(docs): escape curly braces in generated changelog entries

LLM-generated changelog summaries occasionally contain literal
`{...}` (e.g. "{user_id}" template variable), which docusaurus MDX v3
tries to evaluate as a JSX expression, breaking SSG with
`ReferenceError: user_id is not defined`.

Escape `{`/`}` in `entry.summary` at render time in the generator, and
hand-fix the two already-landed claude-code changelog files so main's
Deploy Docs workflow goes green again.

* fix(cli): update get_graph call for new chunks API query params

#1236 added chunk_id/document_id/q/tags/tags_match query params to
/banks/{id}/graph but the CLI wrapper was not updated, so a fresh
cargo build fails with an E0061 arity mismatch against the regenerated
progenitor client. Surfaces here because this PR touches hindsight-docs,
which turns on the test-doc-examples (cli) matrix.

Pass None for the new params and keep the existing type_filter/limit
forwarding; argument order matches the alphabetised generated signature.
2026-04-24 11:17:17 +02:00
Nicolò Boschi 8f6e0e5bec feat(api): add exclude_parents filter to list operations (#1230)
* feat(api): add exclude_parents filter to list operations endpoint

Batch retain operations create parent + child rows, cluttering the
operations list. Add an `exclude_parents` query parameter that filters
out parent operations (is_parent=true in result_metadata). The control
plane UI now passes this by default so users only see leaf operations.

* test: add unit test for exclude_parents filter

* fix: update Rust CLI and docs skill for new exclude_parents param
2026-04-23 15:24:13 +02:00
Nicolò Boschi 80982da577 fix(ops): expose processing/cancelled statuses through API and UI (#1231)
* fix(ops): expose processing/cancelled statuses through API and UI

The API was collapsing 'processing' into 'pending' before returning
operation status to clients. Cancel was deleting the operation row
instead of preserving it with a 'cancelled' status.

- Stop mapping processing→pending in list/get operation responses
- Add 'processing' to OperationStatusResponse Literal type
- Change cancel_operation to set status='cancelled' instead of DELETE
- Guard cancel to only accept pending operations (409 otherwise)
- Extend retry to accept both failed and cancelled operations
- Add _check_op_alive support for cancelled status
- Add DB migration for 'cancelled' in status check constraint
- Add processing/cancelled badges and filters in operations UI
- Add cancel/retry buttons in operation detail dialog
- Align stats card status colors and labels with operations table
- Regenerate OpenAPI spec and all client SDKs

* chore: regenerate docs skill openapi reference

* chore: regenerate clients and openapi spec (full sync)

* fix(cli): handle processing/cancelled status variants in Rust CLI
2026-04-23 15:06:46 +02:00
Nicolò Boschi 8b80959bba feat(mental-models): structured-ops delta refresh + observation cleanup on upsert (#1101)
* feat(mental-models): structured-ops delta refresh + observation cleanup on upsert

Mental model delta mode (primary feature)
- Store mental models as a structured document (sections + typed blocks) in
  a new `structured_content` JSONB column. Markdown shown to users is a
  deterministic render of the structured doc, never an LLM output.
- Delta refresh emits typed operations (`append_block`, `replace_block`,
  `add_section`, `remove_section`, `replace_section_blocks`, …) against the
  structured doc. Sections not mentioned by any op are physically copied
  through unchanged, so prose drift is structurally impossible.
- Text-mode JSON for the LLM call (Gemini rejects the discriminated-union
  schema Pydantic emits); we parse + validate ourselves.
- Token budget for the delta call is 1.5× the doc cap with a 2048 floor and
  the budget is surfaced in the prompt so models can self-trim.
- New `mode: "full" | "delta"` enum on the trigger jsonb. First refresh on
  an empty document falls back to full; a source_query change forces full
  rebuild via `last_refreshed_source_query` tracking column.
- Worker handler `_handle_refresh_mental_model` now delegates to the public
  `refresh_mental_model` (single source of truth — previously had its own
  copy of the reflect+update pipeline that bypassed delta entirely).
- Refuse to overwrite existing content with an empty render — small models
  occasionally return empty answers from the reflect agent and the previous
  behaviour destroyed the working document on transient failures.

Observation cleanup on document upsert (production bug fix)
- `fact_storage.handle_document_tracking` (the retain/upsert path) used to
  delete the document row via FK cascade, removing the source memory_units
  but leaving observations whose source_memory_ids referenced now-deleted
  rows. Only the explicit `MemoryEngine.delete_document` API ran the
  cleanup.
- Extracted `delete_stale_observations_for_memories` to a free function in
  `fact_storage.py`; both code paths (retain upsert + delete API) now run
  the same SQL.
- Migration `c4x5y6z7a8b9` re-runs Pass 2 of `g7h8i9j0k1l2` to sweep the
  orphan observations that accumulated since the last cleanup.

UI
- Refresh-mode select in create/update mental model dialogs.
- Per-row actions dropdown (Edit / Refresh / Delete) on dashboard + table,
  matching the detail dialog's actions menu.
- History diff view: per-token whitespace-insensitive inline diff so only
  the actually-changed substrings light up red/green; runs of unchanged
  lines render as plain text.
- Mental-model dialogs widened to `sm:max-w-2xl` and the scroll wrapper
  inherits the global themed scrollbar (matches the detail modal layout).
- Auto-refresh badge colour unified to green across all surfaces.

Operational logging fixes
- Surface the actual provider response body on `APIStatusError` retries in
  `openai_compatible_llm` instead of only logging on final failure. New
  `_summarize_status_error` helper used in `call()` and `call_with_tools()`.
- Consolidator now logs the failing memory IDs in batch-LLM warnings, so
  `json_validate_failed` + similar errors can be traced to a specific
  memory without waiting for adaptive bisection to narrow it down.
- Worker `[WORKER_STATS]` pool metric was mis-labelled: `waiters` was
  reading `pool._queue.qsize()` (free holders), the opposite of what the
  name implied. Split into `free_holders` (idle holders in queue) and
  `pending_acquires` (`len(_queue._getters)` — actual coroutines blocked
  on `pool.acquire`).

Tests
- 39 unit tests in `test_structured_doc.py` covering schema, renderer,
  parser, op application, ID stability, byte-identical preservation.
- 6 plumbing tests in `test_mental_model_delta.py::TestDeltaRefreshPlumbing`
  covering full/delta branching, source-query change → full rewrite,
  per-row LLM-failure fallback, etc.
- 3 real-LLM eval tests in `TestDeltaRefreshGeminiEval` (gated on
  `HINDSIGHT_RUN_GEMINI_EVALS=1`, prefers Gemini, falls back to OpenAI).

Migrations
- `a2v3w4x5y6z7` — `last_refreshed_source_query TEXT`
- `b3w4x5y6z7a8` — `structured_content JSONB`
- `c4x5y6z7a8b9` — backsweep orphan observations v2

* chore: regenerate clients + add regression tests + lint fixups

- Regenerate OpenAPI spec and Python/TypeScript/Go client SDKs to surface
  the new `mode` field on `MentalModelTrigger`.
- Add regression test for the empty-content guard: when reflect_async
  returns "" and the structured-delta call also fails, refresh must NOT
  overwrite existing content (was destroying working documents).
- Add regression test for the upsert observation cleanup: directly invoke
  `handle_document_tracking` with pre-populated source memories +
  observation, assert the observation is gone after the upsert and the
  surviving co-source memory is reset for re-consolidation.
- Lint hook reformatted long log strings in consolidator.py /
  memory_engine.py / fact_storage.py and ran prettier across the new
  control-plane TS code.

* fix(rust-cli): set mode=Full on MentalModelTriggerInput; refresh generated artefacts

- Generated Rust client now requires `mode: Mode` (not Option) on the
  MentalModelTriggerInput struct since the Python field has a default. Set
  to `Mode::Full` at the call sites in `commands/mental_model.rs`.
- Re-run `generate-openapi.sh` and `generate-docs-skill.sh` after rebasing
  on origin/main so the spec includes upstream additions
  (`failed_consolidation` from #1100). Without this, the new spec dropped
  the field and `check-openapi-compatibility` failed.
- `skills/hindsight-docs/references/openapi.json` is the doc-skill copy of
  the spec; was missing from the previous commit.

* chore: regenerate bank-template-schema.json

Auto-generated from BankTemplateConfig; updated by the structured-doc /
mental-model trigger changes earlier in this PR. ``verify-generated-files``
CI step caught it.

* docs(mental-models): document delta refresh mode

Add a "Refresh Mode" section to the mental-models API docs covering the
new ``mode: "full" | "delta"`` trigger field — strategy explanation,
fallback rules (no existing content / source_query change), empty-answer
preservation, and a quick "when to use which" table.
2026-04-16 18:39:45 +02:00
Nicolò Boschi e1e5f36cee feat(control-plane): surface failed-consolidation count and drilldown (#1100)
* feat(control-plane): surface failed-consolidation count and drilldown

Adds a "Failed" cell to the Consolidation card on the bank General page
that shows how many memories are stuck with consolidation_failed_at. When
non-zero, the cell opens a dialog listing the affected memories with a
"Recover all" action that resets the failed flag and queues a
consolidation run so the worker actually retries them.

Backend: additive only — `failed_consolidation` on BankStatsResponse and
an optional `consolidation_state` filter (failed|pending|done) on
/memories/list. Existing fields and callers are unchanged.

* fix(cli): pass consolidation_state arg through list_memories

* chore: regenerate docs-skill openapi reference
2026-04-16 16:23:35 +02:00
Nicolò Boschi 70d60e96cf feat(cli): add named connection profiles (-p/--profile) (#1080)
* feat(cli): add named connection profiles (-p/--profile)

Adds named profiles stored at ~/.hindsight/cli-profiles/<name>.toml
so a single hindsight binary can target multiple deployments without
stomping on the shared ~/.hindsight/config file. Profiles are plain
TOML (api_url, api_key) with 0600 permissions on Unix.

- New global flag `-p/--profile <NAME>` (also reads $HINDSIGHT_PROFILE)
- New `hindsight profile {create,list,show,delete}` subcommands
- Config precedence: env > profile > ~/.hindsight/config > default
- Missing profile produces an actionable error pointing to
  `hindsight profile create <name> --api-url <url>`
- Unit tests cover round-trip save/load, name validation, list order,
  missing-file error, and 0600 permission bit

* test(cli): end-to-end tests for profile CRUD + docs

- Add tests/cli_profile.rs covering create/list/show/delete against a
  temporary HOME (no API server required), plus `-p` precedence over
  ~/.hindsight/config and the HINDSIGHT_PROFILE env var.
- Fix silent error swallowing in main(): surface anyhow errors via
  ui::print_error before exiting so users see why a command failed
  (previously `profile show missing` just exited 1 with no message).
- Document named profiles in hindsight-docs/docs/sdks/cli.md with the
  new precedence rules.

* fix(cli): regen docs skill + gate profile integration tests to unix

- Run generate-docs-skill.sh so skills/hindsight-docs/references/sdks/cli.md
  picks up the new Named Profiles section (fixes verify-generated-files).
- Gate tests/cli_profile.rs with #![cfg(unix)]: these tests set \$HOME to
  redirect dirs::home_dir() at a tempdir, which only works on Unix.
  On Windows dirs::home_dir() resolves via the shell API (FOLDERID_Profile)
  and ignores env vars, so letting them run there would pollute the real
  user profile. The Windows runtime path is still exercised through the
  config::tests::* unit tests that drive save_profile_to_dir /
  load_profile_from_dir with explicit tempdirs.
2026-04-15 14:58:11 +02:00
Nicolò Boschi f2fc8f9f26 feat(api): add recall controls to mental model trigger (#1048)
* feat(api): add recall controls to mental model trigger

Internal recall during mental model refresh used to hardcode
include_chunks=True with fixed token budgets, wasting prompt budget on
chunks that some refreshes don't need.

Adds three knobs exposed both as hierarchical config (env -> tenant ->
bank) and as per-mental-model overrides on the trigger JSONB field:

- recall_include_chunks / trigger.include_chunks
- recall_max_tokens / trigger.recall_max_tokens
- recall_chunks_max_tokens / trigger.recall_chunks_max_tokens

Trigger value (when set) wins over bank/global config. Both refresh
paths (task handler and synchronous refresh_mental_model) forward the
overrides into reflect_async.

* feat(control-plane): expose recall trigger fields in mental model dialogs

Adds form fields under the Options tab for the three new trigger
overrides (include_chunks, recall_max_tokens, recall_chunks_max_tokens)
in both the create and update mental model dialogs. Empty/Default means
inherit the bank/global config.

* fix(control-plane): cap mental model dialog height and add scroll

* style(control-plane): theme scrollbars to match app surface

* refactor(control-plane): group mental model options into Refresh/Tags/Recall sections

* refactor(control-plane): move Fact Types into Recall, add Other Mental Models section

* fix(cli): pass new recall trigger fields in MentalModelTriggerInput

* chore: regenerate hindsight-docs skill openapi/configuration

* test(hierarchical-config): bump configurable field count for new recall fields
2026-04-14 13:27:16 +02:00