mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
perf/profiling-env
14 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f747d96c38 |
feat(recall): fuzzy tag matching on tag_groups leaves (#4026) (#4028)
* feat(recall): fuzzy tag matching on tag_groups leaves (#4026) Tags increasingly hold user-facing names, and tag filtering is exact array containment. A caller filtering by what a query mentioned passes `typsecript`, and the memory tagged `typescript` is dropped before ranking runs — so the recall returns empty even though ranking would have found it. Better ranking cannot fix that; the match itself has to tolerate the misspelling. A `tag_groups` leaf gains one optional field, `resolve`, defaulting to `exact` (today's behaviour). Set to `fuzzy`, its tags are matched against the bank's tags by similarity instead of literally. That is the whole API change: no new config, no new response field, no new TagsMatch values. Matching is trigram similarity at 0.45 via `entity_resolver._trigram_similarity`, already verified byte-identical to Postgres `similarity()` (#3107), so Postgres, Oracle and store-owned backends behave the same. Resolves: typescropt/typescript 0.57, kubernets/kubernetes 0.62, user:alcie/user:alice 0.47. Does not: mango/mongo 0.33, k9s/k8s 0.14. Known limit, pinned by a test: similarity is length-sensitive. A short tag has few trigrams and one edit destroys three of them, so kakfa/kafka scores 0.20 and does not resolve. Fuzzy matching is effective on descriptive tags and close to inert on very short ones — and that same property is what keeps different short words apart. Resolution runs above the SQL layer, rewriting the leaf into ordinary exact leaves so only those reach the query builders. The ~20 SQL call sites, the Python mirrors used on the graph path, the GIN(tags) index, the store protocol and the Oracle dialect are untouched. Per mode, for tokens t1..tn resolving to E1..En: any/any_strict becomes one leaf over the union; all/all_strict becomes an AND of one OR-leaf per token, so a memory must carry some spelling of each; exact becomes an OR over the cross product, one tag per token, bounded at 32 branches and checked before enumeration, with combinations carrying fewer distinct tags than tokens dropped. Failing closed: a tag that resolves to nothing stays in the filter as itself, leaving the leaf unsatisfiable. Returning an empty list would read as "no tag filtering" in the builders and hand back the whole bank. The vocabulary comes from the existing `list_tags` store method, so there is no schema change. A bank holding more than 5000 distinct tags is rejected with a 422 rather than resolved against a truncated vocabulary. Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk * fix(clients): make tag_groups reachable through the wrapper SDKs `Hindsight.recall(tag_groups=...)` and `.reflect(tag_groups=...)` raised ModuleNotFoundError for every caller. The wrapper imported `hindsight_client_api.models.recall_request_tag_groups_inner`, which the generator does not emit: it produces one union model per tag_groups shape and names it after the first schema that used it, so the class is `MentalModelTriggerInputTagGroupsInner`. Nothing caught it because the wrapper's tests never passed tag_groups and the import sits inside the `if tag_groups is not None` branch, so it only fires when the feature is used. Fixed at both call sites, with mirrored regression tests on the Python and TypeScript wrappers asserting a tag group reaches the request body with the leaf's `resolve` intact — the pair the review checklist asks for, since a capability that exists in one wrapper and not the other is invisible to client-coverage-check (it validates request-body fields, not wrapper surface). Also thread tag_groups through the control-plane recall and reflect proxy routes and their client types. Both accepted every other tag filter and silently dropped this one, so no control-plane caller could use compound tag filtering at all — fuzzy or exact. Two follow-ups from reviewing #4026: - Reject `resolve="fuzzy"` in a mental-model trigger's tag_groups. A trigger's scope is read by two paths that resolve differently: the refresh runs through reflect, which resolves fuzzy leaves, while the staleness check and the scope watermark build SQL straight from the stored groups and do not. A stored fuzzy leaf would build content from the resolved tags while never being marked stale by them, and would drift as the bank's tag vocabulary changes. - Promote `entity_resolver._trigram_similarity` to `trigram_similarity`. Two subsystems now share it — entity resolution and fuzzy tag matching — so the leading underscore misrepresented a real contract between them. Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk |
||
|
|
921ae824a4 |
fix(reflect): fail the run when a retrieval tool fails, and record refused refreshes (#4003)
A mental-model refresh could replace a document built over months with "I don't have information about that" (#2894). When a reflect tool raised, reflect handed the exception back to the model as a tool result and let the loop continue; the model answered from whatever it had — usually nothing — and that answer was indistinguishable from a run over a bank that genuinely holds nothing on the topic. Non-empty prose, so every emptiness guard on the write path let it through, and the operation was recorded as completed. Reflect now fails the run instead: - A tool that RAISES is an infrastructure failure (the database, the embedder, the reranker), not something the model can retry its way out of, so it raises ReflectToolExecutionError. One failure in a parallel batch fails the run. A tool that RETURNS {"error": ...} for a malformed or unavailable call is the model's own mistake and is still fed back to it, unchanged. - A non-context-overflow LLM error that survives one retry re-raises rather than falling through to a forced final synthesis built on evidence the failed turn never finished gathering. Context overflow keeps its deliberate degradation: that is a prompt-budgeting problem with the gathered evidence intact. - OperationCancelledError passes through untouched on both paths, so a client disconnect stays a 499 (#2122) instead of becoming a 500. The line this draws is between "we could not look" and "we looked and there is nothing there". A retrieval that succeeds and returns nothing is an answer, and still rewrites the document. The refresh re-raises both reflect failures as a typed MentalModelRefreshError (refresh_failed_reflect_error, with reasons retrieval_failed / no_answer), so they reach the operation's typed `details` through the same failure-metadata hook every other refusal already used. Content, structured document and watermark are all left untouched, so the retry re-reads the same window. Every refused refresh — the new reflect-side ones and the existing _preserve_and_fail paths — now also appends a failure record to mental_model_history carrying the reason and the exception. Before this a failed refresh left no trace on the model at all: the History tab kept rendering the last SUCCESSFUL trace as though it were current, and the only record was prose on an async-operation row no mental-model view reads. Retention is applied per kind so a run of failures cannot evict the version history. Control plane: a new Errors tab on the mental-model modal shows those records as a timeline of events — reason, attempt count, time, exception — with a red dot on the tab when there are any. History goes back to being a version browser. Consecutive identical failures collapse into one event (the worker retries each refresh), but a successful refresh between two of them breaks the chain, so separate outages stay separate. Also fixes a metadata bug found while verifying this live: a refresh is retried on the same operation row, and the success path wrote no failure_reason to overwrite the failed attempt's — producing outcome=content_written alongside failure_reason=no_answer. The success now writes it as null explicitly. Claude-Session: https://claude.ai/code/session_01F3i5UVdZRK16AZ9oFQqsg4 |
||
|
|
82859af02b |
docs(reflect): clarify tag-scoped directive behavior (#3038)
Explain how tags, tags_match, tag_groups, and directive isolation interact across REST, MCP, versioned docs, and agent skills. Clarify the different defaults used by reflect and directive listing, then regenerate OpenAPI and supported client artifacts. |
||
|
|
4278f0989d |
feat(reflect,mental-models): surface structured output in the control plane (#3113)
* feat(reflect,mental-models): surface structured output in control plane
Reflect's response_schema -> structured_output was already implemented and
tested in the engine but never exposed in the UI. Surface it in the reflect
(think) view, and extend the same structured-output extraction to mental
models via a per-model response_schema stored in the trigger config.
- engine: refresh_mental_model reads trigger.response_schema, forwards it to
the internal reflect call, and persists the parsed structured_output onto
the stored reflect_response payload; fix stale 'not yet supported' docstrings
- api: add response_schema to MentalModelTrigger
- control-plane: reflect route + api.ts forward response_schema; think-view
gets a JSON-schema input and renders structured_output; create/update mental
model dialogs get a schema editor; detail modal renders structured_output
- tests: mental model structured-output plumbing (schema forwarded + persisted)
- regenerate OpenAPI spec + client SDKs; add i18n keys for all locales
* feat(control-plane): show configured response_schema in mental model config tab
Adds a read-only JSON card for the mental model's trigger.response_schema in
the detail modal's Configuration tab (mirrors the tag_groups card), plus the
regenerated go openapi.yaml.
* style: ruff format test_mental_model_structured_output
* fix(control-plane): don't route the JSON schema example through next-intl
The response_schema placeholder was a t() message whose value is literal JSON.
next-intl parses messages as ICU, so the '{' in the example was read as an
argument placeholder, the parse failed, and the field rendered the raw message
key instead of the example. Inline the JSON example directly on the placeholder
prop (i18n:check skips JSON-shaped placeholders) and drop the now-unused
*Placeholder message keys. Caught by running the control plane.
* feat(structured-output): validate response_schema + add a no-code schema builder
Validation (both reflect and mental models): a schema that is valid JSON but
not a usable object-with-properties silently produced empty structured_output
or blew up inside the LLM extraction call later. Now:
- engine: validate_response_schema() enforces the usable-shape contract
(object schema, non-empty properties, well-formed required); wired as Pydantic
field_validators on ReflectRequest.response_schema and
MentalModelTrigger.response_schema (invalid -> HTTP 422).
- control-plane: the reflect and mental-model forms validate the schema shape on
submit (not just JSON.parse) and surface the specific error.
No-code schema builder: a 'Build schema' button on both the reflect view and the
mental-model dialogs opens a dialog with Visual and Code modes. Visual mode edits
a flat field list (name, type, array item-type, description, required); Code mode
edits raw JSON. The two stay in sync and Apply is gated on a usable schema. Shared
frontend lib (response-schema.ts) mirrors the backend contract.
tests: test_response_schema_validation.py (16 cases: validator + model integration).
* refactor(control-plane): schema only via the builder, show set/unset status
Removes the inline response_schema JSON textarea from the reflect view and the
mental-model dialogs. Editing now happens exclusively in the schema builder; the
page shows only whether a schema is set (field count + names, with Edit/Remove)
or a Build schema button when none. Extracts the shared ResponseSchemaField
component used identically by reflect and both mental-model dialogs.
* fix(mental-models): derive structured_output from final content, not reflect's answer
In delta mode reflect only sees facts created since the last refresh, so its
answer (and any structured_output it derived) reflects just the delta — while the
stored content is the delta-merged document. Persisting the reflect-derived value
made structured_output inconsistent with the markdown.
Now the mental-model refresh no longer passes response_schema to reflect; instead
it extracts structured_output from the FINAL stored content (correct for both full
and delta), and carries the previous value forward untouched when a delta refresh
preserves content (no new facts). Adds a delta test asserting extraction runs
against the merged document, not reflect's partial answer.
* fix(schema-builder): allow switching an empty schema from Code back to Visual
An empty schema serialises to properties:{}, which schemaToFields mapped to an
empty array — and the Code->Visual guard treated 'empty' the same as 'not
representable', blocking the switch. schemaToFields now returns [] (representable)
for a missing/empty properties map and null only for genuinely unrepresentable
schemas; the switch seeds a blank field when empty.
* docs(reflect): document structured output (response_schema) + schema builder
Adds a Structured Output section to the reflect docs: how response_schema returns
both text and a structured_output projection of the same answer, the schema rules,
mental-model structured output (extracted from the final/merged document), and the
no-code Build schema editor. Regenerates the docs-skill mirror.
* feat(schema-builder): recursive visual editor for nested objects & arrays
The visual editor was flat — object/array fields had no way to define their inner
shape. Reworks the field model into a recursive tree (each field has a node; an
object node nests fields, an array node nests an item node) so you can build
nested objects and arrays-of-objects entirely in the visual editor. Code<->Visual
round-trips losslessly; schemas using features the editor can't represent (enum,
oneOf, $ref, tuple items, …) stay in code mode rather than being silently
flattened.
* fix(structured-output): recursive model for nested schemas + fail refresh loudly
Two problems surfaced by nested schemas on Gemini:
1. _generate_structured_output mapped object/array properties to bare dict/list,
which serialize with additionalProperties — rejected by the Gemini API. So any
schema with a nested object/array silently failed extraction. Now it builds a
proper recursive Pydantic model (nested objects -> nested models, arrays ->
typed lists), matching how retain's structured output already works on Gemini.
2. On extraction failure the mental-model refresh silently persisted content with
no structured_output, clobbering the previously-stored value. Now, when a
response_schema is configured and extraction yields nothing, the refresh raises
MentalModelRefreshError — prior content and structured_output are preserved and
the refresh can be retried.
Verified live on Gemini: a nested {location:object, people:array} schema now
extracts (structured_output present) instead of failing on additionalProperties.
Adds a fail-loud regression test.
* fix(schema-builder): readable error text in dark mode
text-destructive resolves to a dark red (#C0183A) in dark mode, which is
low-contrast on the dark dialog background. Use the codebase's standard
readable pattern (text-red-600 dark:text-red-400) for the builder's validation
error and the invalid-schema notice.
* fix(cli): set response_schema on MentalModelTriggerInput literals
Adding response_schema to MentalModelTrigger regenerated the Rust
MentalModelTriggerInput struct with a new field; the hand-written CLI struct
literals must initialize it (E0063). Sets response_schema: None in the three
construction sites (create/update mental model, knowledge-base pin).
* docs(api): document mental-model response_schema; fix stale reflect text-empty claim; test schema lib
- api/mental-models: document the trigger.response_schema flag + a Structured
Output section (extraction from final content, fail-loud, validation).
- api/reflect: correct the stale claim that text is empty with response_schema —
reflect returns both text and structured_output.
- control-plane: vitest unit tests for the response-schema lib (validation +
recursive fields<->schema round-trip).
- regenerate docs-skill mirror.
* chore: regenerate bank-template-schema for MentalModelTrigger.response_schema
The bank template schema embeds MentalModelTrigger; adding response_schema to
the trigger changed the generated schema. Regenerated so verify-generated-files
passes.
|
||
|
|
2f075dedea |
feat(tags): officially surface tags_match=exact across UI, docs, clients (#2230)
The `exact` set-equality match mode landed in the API + generated clients in #2149 but was never exposed in the control plane, documented, or added to the hand-maintained SDK wrappers. This completes the feature. Control plane: add `exact` to the TagsMatch type/unions and to the tags_match dropdowns in think-view, search-debug-view, and both mental-model trigger forms; add translated labels to all 10 locales. Docs: document `exact` in the recall tags_match table + tag_groups, the reflect tags value list, and the observations scope-listing guide; regenerate the docs skill mirror. Clients: add `exact` to the hand-maintained Python and TypeScript wrapper Literals/unions and docstrings (generated clients already had it; Rust is generated from openapi.json at build time). Supersedes #2159. |
||
|
|
a56cd044e5 |
feat: 4-tab code parity across all documentation examples (#613)
* feat: independent versioning for integrations
- Add per-integration changelog pages at /changelog/integrations/<name>
- Move main changelog to changelog/index.md (URL unchanged)
- Add --integration flag to generate-changelog for LLM-based per-integration changelog generation
- Add scripts/release-integration.sh <name> <version> for cutting integration releases
- Add .github/workflows/release-integration.yml to publish on integrations/** tags
- Remove integrations from main release.sh and release.yml cycle
* fix: add agno and hermes integration docs to version-0.4 for production build
* chore: apply ruff formatting to generate_changelog.py
* feat: add 4-tab code parity across all documentation examples
Every code snippet Tabs block now has Python, Node.js, CLI, and Go variants.
Raw HTTP/curl tabs replaced with proper SDK calls.
New example files:
- Go: retain.go, recall.go, reflect.go, memory-banks.go, directives.go,
mental-models.go, documents.go, main-methods.go
- Shell: memory-banks.sh, directives.sh, mental-models.sh
- Node.js: mental-models.mjs
Extended example files with missing sections:
- recall.mjs/sh: world/experience/observation types, token-budget, all tag modes
- reflect.sh: reflect-with-params, reflect-disposition, reflect-sources, reflect-with-tags
- reflect.mjs: reflect-with-tags, fixed reflect-sources API usage
- retain.mjs/sh: retain-conversation, retain-batch, retain-files-batch
SDK/CLI additions:
- TypeScript: getMentalModelHistory method
- CLI recall: --tags, --tags-match flags
- CLI reflect: --tags, --tags-match, --include-facts flags
- CLI directive update: --is-active flag
- CLI bank set-config: --retain-mission, --retain-extraction-mode,
--observations-mission, --reflect-mission, --disposition-* flags
Build validation:
- scripts/check-code-parity.mjs validates 4-tab parity across all MDX files
- Integrated into npm run build — fails if any Tabs block is missing a variant
* fix: fix doc examples for Go, Node.js, CLI + add mental model with-id examples
- Fix Go Budget constants: BUDGET_HIGH/LOW/MID → HIGH/LOW/MID
- Fix Go documents.go: ListDocuments returns []map[string]interface{}, use map access
- Fix Go retain.go: use correct relative path for sample.pdf
- Fix Node.js createMentalModel: use positional args (name, sourceQuery) not object
- Add CLI 'history' subcommand for mental models (api.rs, main.rs, mental_model.rs)
- Rebuild TypeScript/Python clients to support id param in createMentalModel
- Add create-mental-model-with-id examples across all 4 languages and docs
* fix: move id param to end of create_mental_model signature for backwards compat
|
||
|
|
278344b3b3 |
doc: improve api explanation (#415)
* doc: improve api explanation * doc: improve api explanation * doc: improve api explanation * fix: add include_facts to reflect client, fix retain.sh temp files, fix main-methods based_on access * fix: create report.pdf in working directory for retain.sh file upload examples |
||
|
|
90e370ef35 |
fix: misc fixes for observations and mental models (#209)
* fix: misc fixes for observations and mental models * feat: improve graph retrieval for observations - Update LinkExpansionRetriever to traverse through source_memory_ids for observation entity connections (avoiding data duplication) - Remove entity link copy from world facts to observations in consolidator - Add tests for link expansion graph retrieval - Add directives_applied field to ReflectResult - Include user's other changes (CLI, docs, client updates) * fix: CI test failures - Add mental_model_id parameter to create_mental_model function - Fix ToolCallTrace not including reason field from ToolCall - Improve test_link_expansion_observation_graph_retrieval to wait for consolidation with retry * chore: reduce link expansion log verbosity * Revert "chore: reduce link expansion log verbosity" This reverts commit 3ce759391cead1012157785fa78fef16ef9bfe3b. * feat: add semantic/temporal/entity links as fallback in graph retrieval - Add fallback query for semantic, temporal, and entity links from memory_links - Check both directions (outgoing and incoming links) - Weight fallback results at 0.5x to prioritize entity links via unit_entities - Fixes graph retrieval returning 0 when data has cross-cluster temporal connections * fix: enable observations fixture for link expansion test - Add enable_observations fixture to ensure observations are created - Increase wait time from 10 to 30 seconds for CI reliability |
||
|
|
5b52a84fff |
chore: internal renames (#204)
This commit renames the terminology across the entire codebase: - "mental models" (fact_type='mental_model' in memory_units) → "observations" - "reflections" table (stored reflect responses) → "mental_models" Changes include: - Database migration to rename tables, indexes, and constraints - API endpoints: /reflections → /mental-models, /mental-models → /observations - Config: ENABLE_MENTAL_MODELS → ENABLE_OBSERVATIONS - Response models and Pydantic classes - Reflect agent tools and prompts - Control plane UI and routes - Documentation and examples - Regenerated OpenAPI spec and client SDKs (Python, TypeScript) - Rust CLI: reflection commands → mental-model commands - LiteLLM: updated fact_types documentation |
||
|
|
522b71aab8 |
doc: mental models (#199)
* doc: mental models * doc: mental models |
||
|
|
4476a10aa3 |
doc: refinement for 0.3.0 new features (#159)
* doc: refinement for 0.3.0 new features * fix * fix * fixes |
||
|
|
29a542dc23 |
feat: Add per-request LLM token usage metrics (#117)
* feat: Record LLM token metrics via Prometheus Wire up the existing token metrics infrastructure to actually record token usage from LLM calls. The MetricsCollector already had record_tokens() method and Prometheus counters (hindsight.tokens.input, hindsight.tokens.output), but they were never being populated. Changes: - Import get_metrics_collector in llm_wrapper.py - Call record_tokens() after successful LLM calls for: - OpenAI/Groq (using response.usage.prompt_tokens, completion_tokens) - Anthropic (using response.usage.input_tokens, output_tokens) - Gemini (using response.usage_metadata.prompt_token_count, candidates_token_count) - Add test file to verify token metrics are recorded Note: Ollama's native API doesn't return token usage, so metrics are not recorded for that provider. The token metrics will now be available via /metrics endpoint: - hindsight_tokens_input_total - hindsight_tokens_output_total * feat: add per-request token usage tracking to retain and reflect endpoints - Add TokenUsage model with input_tokens, output_tokens, total_tokens - Return usage metrics in retain response (sync operations only) - Return usage metrics in reflect response - Update Python, TypeScript, and Rust clients - Add API documentation for usage fields - Add changelog entry |
||
|
|
d49e8201b4 |
feat: add max_tokens and structured output to /reflect (#74)
* feat: add structured output to /reflect * feat: add structured output to /reflect * imrpove * add max_toksn * fix rust client * fix rust client * fix rust client * try fix * try fix * no stricts |
||
|
|
d405b4feed |
ci: finalize test for the documentation code (#57)
* Fix main-methods.py: entities is a dict, use .items() and .canonical_name * Migrate docs to use CodeSnippet components - Convert quickstart.md, retain.md, recall.md, reflect.md, memory-banks.md to .mdx - Use CodeSnippet to pull code from validated example scripts - Add missing 'name' parameter to create_bank calls - Fix main-methods.py entities iteration (dict not list) - Remove retain-new.mdx demo file * Migrate existing docs to match testing pattern with code snippet and add CLI tests to the CI * Fix doc-id issue + add main-method tests * CLI fixes * Update openAPI json * Fix rust build issues * increase sleep time for Hindsight to process the document * Added a polling sleep instead of fixed * Delete immediately fails, so create the doc a earlier in the test to get the doc ready * Add debug logs * Remove debug logs |