Files
Nicolò Boschi f747d96c38 feat(recall): fuzzy tag matching on tag_groups leaves (#4026) (#4028)
* feat(recall): fuzzy tag matching on tag_groups leaves (#4026)

Tags increasingly hold user-facing names, and tag filtering is exact array
containment. A caller filtering by what a query mentioned passes `typsecript`,
and the memory tagged `typescript` is dropped before ranking runs — so the
recall returns empty even though ranking would have found it. Better ranking
cannot fix that; the match itself has to tolerate the misspelling.

A `tag_groups` leaf gains one optional field, `resolve`, defaulting to `exact`
(today's behaviour). Set to `fuzzy`, its tags are matched against the bank's
tags by similarity instead of literally. That is the whole API change: no new
config, no new response field, no new TagsMatch values.

Matching is trigram similarity at 0.45 via `entity_resolver._trigram_similarity`,
already verified byte-identical to Postgres `similarity()` (#3107), so Postgres,
Oracle and store-owned backends behave the same. Resolves: typescropt/typescript
0.57, kubernets/kubernetes 0.62, user:alcie/user:alice 0.47. Does not: mango/mongo
0.33, k9s/k8s 0.14.

Known limit, pinned by a test: similarity is length-sensitive. A short tag has
few trigrams and one edit destroys three of them, so kakfa/kafka scores 0.20 and
does not resolve. Fuzzy matching is effective on descriptive tags and close to
inert on very short ones — and that same property is what keeps different short
words apart.

Resolution runs above the SQL layer, rewriting the leaf into ordinary exact
leaves so only those reach the query builders. The ~20 SQL call sites, the
Python mirrors used on the graph path, the GIN(tags) index, the store protocol
and the Oracle dialect are untouched. Per mode, for tokens t1..tn resolving to
E1..En: any/any_strict becomes one leaf over the union; all/all_strict becomes an
AND of one OR-leaf per token, so a memory must carry some spelling of each;
exact becomes an OR over the cross product, one tag per token, bounded at 32
branches and checked before enumeration, with combinations carrying fewer
distinct tags than tokens dropped.

Failing closed: a tag that resolves to nothing stays in the filter as itself,
leaving the leaf unsatisfiable. Returning an empty list would read as "no tag
filtering" in the builders and hand back the whole bank.

The vocabulary comes from the existing `list_tags` store method, so there is no
schema change. A bank holding more than 5000 distinct tags is rejected with a
422 rather than resolved against a truncated vocabulary.

Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk

* fix(clients): make tag_groups reachable through the wrapper SDKs

`Hindsight.recall(tag_groups=...)` and `.reflect(tag_groups=...)` raised
ModuleNotFoundError for every caller. The wrapper imported
`hindsight_client_api.models.recall_request_tag_groups_inner`, which the
generator does not emit: it produces one union model per tag_groups shape and
names it after the first schema that used it, so the class is
`MentalModelTriggerInputTagGroupsInner`. Nothing caught it because the wrapper's
tests never passed tag_groups and the import sits inside the `if tag_groups is
not None` branch, so it only fires when the feature is used.

Fixed at both call sites, with mirrored regression tests on the Python and
TypeScript wrappers asserting a tag group reaches the request body with the
leaf's `resolve` intact — the pair the review checklist asks for, since a
capability that exists in one wrapper and not the other is invisible to
client-coverage-check (it validates request-body fields, not wrapper surface).

Also thread tag_groups through the control-plane recall and reflect proxy routes
and their client types. Both accepted every other tag filter and silently
dropped this one, so no control-plane caller could use compound tag filtering at
all — fuzzy or exact.

Two follow-ups from reviewing #4026:

- Reject `resolve="fuzzy"` in a mental-model trigger's tag_groups. A trigger's
  scope is read by two paths that resolve differently: the refresh runs through
  reflect, which resolves fuzzy leaves, while the staleness check and the scope
  watermark build SQL straight from the stored groups and do not. A stored fuzzy
  leaf would build content from the resolved tags while never being marked stale
  by them, and would drift as the bank's tag vocabulary changes.

- Promote `entity_resolver._trigram_similarity` to `trigram_similarity`. Two
  subsystems now share it — entity resolution and fuzzy tag matching — so the
  leading underscore misrepresented a real contract between them.

Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk
2026-09-02 16:50:53 +02:00

271 lines
13 KiB
Plaintext

---
sidebar_position: 3
---
# Reflect
Generate a grounded, disposition-aware response using an agentic reasoning loop.
When you call **reflect**, Hindsight runs an agentic loop that autonomously searches the memory bank using multiple retrieval tools, applies the bank's disposition traits to shape the reasoning style, and produces a final answer grounded in what it found. Unlike recall — which returns raw facts — reflect returns a synthesized response written by the LLM.
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
import CodeSnippet from '@site/src/components/CodeSnippet';
{/* Import raw source files */}
import reflectPy from '!!raw-loader!@site/examples/api/reflect.py';
import reflectMjs from '!!raw-loader!@site/examples/api/reflect.mjs';
import reflectSh from '!!raw-loader!@site/examples/api/reflect.sh';
import reflectGo from '!!raw-loader!@site/examples/api/reflect.go';
:::info How Reflect Works
Learn about disposition-driven reasoning in the [Reflect Architecture](/developer/reflect) guide.
:::
:::tip Prerequisites
Make sure you've completed the [Quick Start](./quickstart) to install the client and start the server.
:::
## Basic Usage
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={reflectPy} section="reflect-basic" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={reflectMjs} section="reflect-basic" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
<CodeSnippet code={reflectSh} section="reflect-basic" language="bash" />
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={reflectGo} section="reflect-basic" language="go" />
</TabItem>
</Tabs>
---
## Parameters
### query
The question or prompt to reflect on. This is the only required field. If you have situational context that should influence the answer, include it directly in the query rather than as a separate field.
### budget
Controls how thoroughly the agent explores the memory bank before answering. Accepted values are `low` (default), `mid`, and `high`. At `low`, the agent does a shallow search optimized for speed. At `mid`, it checks multiple sources when the question warrants it. At `high`, it performs deep exploration across all knowledge levels and may use multiple query variations to find indirect connections. Use `high` for complex questions that require synthesizing information from many sources.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={reflectPy} section="reflect-with-params" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={reflectMjs} section="reflect-with-params" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
<CodeSnippet code={reflectSh} section="reflect-with-params" language="bash" />
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={reflectGo} section="reflect-with-params" language="go" />
</TabItem>
</Tabs>
### max_tokens
Limits the length of the final generated response. Defaults to `4096`. This does not affect how much the agent can retrieve during the agentic loop — only the final answer length.
### response_schema
An optional JSON Schema **object** with a non-empty `properties` map (nested objects and arrays are supported). When provided, the response includes a `structured_output` field **in addition to** the markdown `text`: the agent reasons to its answer, then a second pass extracts that answer into JSON matching your schema. `structured_output` is therefore a faithful projection of `text` — you get the readable answer *and* a typed object to program against, never one instead of the other. Invalid schemas (not an object, or no properties) are rejected before the call runs.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={reflectPy} section="reflect-structured-output" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={reflectMjs} section="reflect-structured-output" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
<CodeSnippet code={reflectSh} section="reflect-structured-output" language="bash" />
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={reflectGo} section="reflect-structured-output" language="go" />
</TabItem>
</Tabs>
### tags
Defines the visibility scope used throughout the reflect agent. It filters raw
facts, observations, and mental models that the agent can retrieve. The same
`tags` and `tags_match` values also select which tagged directives are injected
into the reflect prompt.
`tags` defaults to `null`, `tags_match` defaults to `any`, and `tag_groups`
defaults to `null`. For non-empty tags, raw facts, observations, and mental
models use the same matching modes as [recall tags](./recall#tags). Directives
have one additional rule: untagged directives are global and remain eligible
whenever a tag scope is supplied, including with a strict or exact match.
| Reflect configuration | Raw facts and observations | Mental models | Active directives |
|-----------------------|----------------------------|---------------|-------------------|
| Omit `tags`, `tags_match`, and `tag_groups` | All tagged and untagged data | All tagged and untagged models | Untagged/global directives only |
| `tags: []`, default `tags_match: "any"` | All tagged and untagged data | All tagged and untagged models | Untagged/global directives only |
| No tags, `tags_match: "exact"` | Untagged/global data only | Untagged/global models only | Untagged/global directives only |
| Non-empty `tags`, `any` or `all` | Matching tagged data plus untagged/global data | Matching models plus untagged/global models | Matching tagged directives plus untagged/global directives |
| Non-empty `tags`, `any_strict` or `all_strict` | Matching tagged data only | Matching tagged models only | Matching tagged directives plus untagged/global directives |
| Non-empty `tags`, `exact` | Data with exactly the requested tag set | Models with exactly the requested tag set | Exactly matching tagged directives plus untagged/global directives |
| Non-empty `tag_groups`, default top-level `tags_match` | Data matching the compound expression | Models matching the compound expression | Matching tagged directives plus untagged/global directives |
A `tag_groups` leaf may set `resolve: "fuzzy"`, as in
[recall](./recall#fuzzy-leaves); reflect resolves it once, before the agentic
loop starts, so every tool it runs filters on the same tags.
The first row is intentionally asymmetric: an unscoped reflect can search all
memories, but it does not load tagged directives. To create a directive that
applies to every reflect call, leave its `tags` empty. To create a scoped
directive, assign tags and pass a matching scope to `reflect`.
:::note `isolation_mode`
`isolation_mode` is an internal `list_directives` option, not a public reflect
request parameter. Reflect always enables it. When neither `tags` nor
`tag_groups` is supplied, it limits directive loading to untagged directives.
There is currently no per-request switch to disable it.
:::
:::note MCP omitted tags
The MCP `reflect` tool forwards `tags_match` only when `tags` is present. To
request the empty exact scope through MCP, pass `tags: []` together with
`tags_match: "exact"`.
:::
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={reflectPy} section="reflect-with-tags" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={reflectMjs} section="reflect-with-tags" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
<CodeSnippet code={reflectSh} section="reflect-with-tags" language="bash" />
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={reflectGo} section="reflect-with-tags" language="go" />
</TabItem>
</Tabs>
#### Common scope examples
Unscoped reflect searches all memories but applies only global directives:
```json
{
"query": "Summarize the current project status"
}
```
A project scope includes global data and directives alongside matching
`project:a` data and directives:
```json
{
"query": "Summarize the current project status",
"tags": ["project:a"]
}
```
A strict project scope excludes untagged memories, observations, and mental
models. Global directives still apply:
```json
{
"query": "Summarize the current project status",
"tags": ["project:a"],
"tags_match": "all_strict"
}
```
### tag_groups
Provides compound tag filtering with recursive `and`, `or`, and `not`
expressions. It affects the same reflect data sources and directive selection as
flat `tags`. `tag_groups` and `tags` are mutually exclusive in the public REST
request. Each leaf supplies its own matching mode and defaults to
`any_strict`; the top-level groups are AND-ed. Normally leave the top-level
`tags_match` at its default, `any`. Setting it to `exact` while using
`tag_groups` additionally constrains facts, observations, and mental models to
the global flat scope before applying the compound expression.
The MCP `reflect` tool currently exposes flat `tags` and `tags_match`, but not
`tag_groups`.
### include
Controls optional supplementary data returned alongside the main response.
#### include.facts
When enabled, the response includes a `based_on` object listing the memories, mental models, and directives the agent actually used to construct the answer. Only sources retrieved during the agent loop can appear here — citations are validated to prevent hallucinated references. Useful for transparency and verification.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={reflectPy} section="reflect-sources" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={reflectMjs} section="reflect-sources" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
<CodeSnippet code={reflectSh} section="reflect-sources" language="bash" />
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={reflectGo} section="reflect-sources" language="go" />
</TabItem>
</Tabs>
#### include.tool_calls
When enabled, the response includes a `trace` object with the full execution log of every tool call and LLM call made during the agentic loop, including inputs, outputs, and durations. Set `output: false` to include only tool inputs for a smaller payload. Useful for debugging why the agent reached a particular conclusion.
---
## Response
### text
The synthesized answer as a well-formatted markdown string. This is the primary output of reflect. Still returned when `response_schema` is provided — `structured_output` is derived from it, not a replacement for it.
### structured_output
The LLM's response parsed according to the `response_schema` provided in the request. Only present when `response_schema` was set. `null` otherwise.
### based_on
The sources the agent used to construct the answer. Only present when `include.facts` was enabled. Contains three fields:
- `memories` — a list of memory facts (world, experience, observation) that were retrieved and cited. Each item has `id`, `text`, `type`, `context`, `occurred_start`, and `occurred_end`.
- `mental_models` — a list of mental models that were used. Each item has `id`, `text`, and `context`.
- `directives` — a list of directives that were enforced during reasoning. Each item has `id`, `name`, and `content`.
### usage
Token usage for all LLM calls made during the agentic loop: `input_tokens`, `output_tokens`, and `total_tokens`. Useful for cost tracking.
### trace
The full execution log of the agentic loop. Only present when `include.tool_calls` was enabled. Contains:
- `tool_calls` — each tool invocation with `tool` name (`lookup`, `recall`, `learn`, `expand`), `input`, `output` (if `output: true`), `duration_ms`, and `iteration` number.
- `llm_calls` — each LLM call with `scope` (e.g., `"agent_1"`, `"final"`) and `duration_ms`.
## When Reflect Fails
Reflect answers from evidence it gathered, so a run that could not gather it does not answer at all — it fails with a **500**, rather than returning a confident reply built on nothing:
- **A retrieval tool raised.** The database, the embedder or the reranker was unavailable. One failure in a batch fails the run: an answer written around a tool that never returned is indistinguishable from one over a bank that genuinely holds nothing on the topic, and callers store it as a real answer.
- **The model produced no answer.** The `done` call arrived empty, or the final synthesis returned nothing.
- **The model or transport cannot drive tool calls.** Reflect is driven entirely by structured tool calls; a transport that silently drops the tool definitions fails loudly so you can switch to a tool-calling-capable model.
- **The provider kept failing.** A non-context-overflow LLM error is retried once inside the loop and then given up on.
A **successful retrieval that returns nothing is not a failure**: the bank genuinely has nothing on the topic, and reflect says so in its answer.
Two cases are deliberately not failures. A run whose *context window* overflows synthesizes from the evidence it has — the prompt was too big for the model, which is a budgeting problem, not a broken dependency. And a tool call the model got *wrong* — a missing argument, a tool that does not exist — is returned to it as an error to fix, not raised.