mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
f747d96c38
* feat(recall): fuzzy tag matching on tag_groups leaves (#4026) Tags increasingly hold user-facing names, and tag filtering is exact array containment. A caller filtering by what a query mentioned passes `typsecript`, and the memory tagged `typescript` is dropped before ranking runs — so the recall returns empty even though ranking would have found it. Better ranking cannot fix that; the match itself has to tolerate the misspelling. A `tag_groups` leaf gains one optional field, `resolve`, defaulting to `exact` (today's behaviour). Set to `fuzzy`, its tags are matched against the bank's tags by similarity instead of literally. That is the whole API change: no new config, no new response field, no new TagsMatch values. Matching is trigram similarity at 0.45 via `entity_resolver._trigram_similarity`, already verified byte-identical to Postgres `similarity()` (#3107), so Postgres, Oracle and store-owned backends behave the same. Resolves: typescropt/typescript 0.57, kubernets/kubernetes 0.62, user:alcie/user:alice 0.47. Does not: mango/mongo 0.33, k9s/k8s 0.14. Known limit, pinned by a test: similarity is length-sensitive. A short tag has few trigrams and one edit destroys three of them, so kakfa/kafka scores 0.20 and does not resolve. Fuzzy matching is effective on descriptive tags and close to inert on very short ones — and that same property is what keeps different short words apart. Resolution runs above the SQL layer, rewriting the leaf into ordinary exact leaves so only those reach the query builders. The ~20 SQL call sites, the Python mirrors used on the graph path, the GIN(tags) index, the store protocol and the Oracle dialect are untouched. Per mode, for tokens t1..tn resolving to E1..En: any/any_strict becomes one leaf over the union; all/all_strict becomes an AND of one OR-leaf per token, so a memory must carry some spelling of each; exact becomes an OR over the cross product, one tag per token, bounded at 32 branches and checked before enumeration, with combinations carrying fewer distinct tags than tokens dropped. Failing closed: a tag that resolves to nothing stays in the filter as itself, leaving the leaf unsatisfiable. Returning an empty list would read as "no tag filtering" in the builders and hand back the whole bank. The vocabulary comes from the existing `list_tags` store method, so there is no schema change. A bank holding more than 5000 distinct tags is rejected with a 422 rather than resolved against a truncated vocabulary. Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk * fix(clients): make tag_groups reachable through the wrapper SDKs `Hindsight.recall(tag_groups=...)` and `.reflect(tag_groups=...)` raised ModuleNotFoundError for every caller. The wrapper imported `hindsight_client_api.models.recall_request_tag_groups_inner`, which the generator does not emit: it produces one union model per tag_groups shape and names it after the first schema that used it, so the class is `MentalModelTriggerInputTagGroupsInner`. Nothing caught it because the wrapper's tests never passed tag_groups and the import sits inside the `if tag_groups is not None` branch, so it only fires when the feature is used. Fixed at both call sites, with mirrored regression tests on the Python and TypeScript wrappers asserting a tag group reaches the request body with the leaf's `resolve` intact — the pair the review checklist asks for, since a capability that exists in one wrapper and not the other is invisible to client-coverage-check (it validates request-body fields, not wrapper surface). Also thread tag_groups through the control-plane recall and reflect proxy routes and their client types. Both accepted every other tag filter and silently dropped this one, so no control-plane caller could use compound tag filtering at all — fuzzy or exact. Two follow-ups from reviewing #4026: - Reject `resolve="fuzzy"` in a mental-model trigger's tag_groups. A trigger's scope is read by two paths that resolve differently: the refresh runs through reflect, which resolves fuzzy leaves, while the staleness check and the scope watermark build SQL straight from the stored groups and do not. A stored fuzzy leaf would build content from the resolved tags while never being marked stale by them, and would drift as the bank's tag vocabulary changes. - Promote `entity_resolver._trigram_similarity` to `trigram_similarity`. Two subsystems now share it — entity resolution and fuzzy tag matching — so the leading underscore misrepresented a real contract between them. Claude-Session: https://claude.ai/code/session_011KDT484YujNcBxHzbfzNbk
271 lines
13 KiB
Plaintext
271 lines
13 KiB
Plaintext
---
|
|
sidebar_position: 3
|
|
---
|
|
|
|
# Reflect
|
|
|
|
Generate a grounded, disposition-aware response using an agentic reasoning loop.
|
|
|
|
When you call **reflect**, Hindsight runs an agentic loop that autonomously searches the memory bank using multiple retrieval tools, applies the bank's disposition traits to shape the reasoning style, and produces a final answer grounded in what it found. Unlike recall — which returns raw facts — reflect returns a synthesized response written by the LLM.
|
|
|
|
import Tabs from '@theme/Tabs';
|
|
import TabItem from '@theme/TabItem';
|
|
import CodeSnippet from '@site/src/components/CodeSnippet';
|
|
|
|
{/* Import raw source files */}
|
|
import reflectPy from '!!raw-loader!@site/examples/api/reflect.py';
|
|
import reflectMjs from '!!raw-loader!@site/examples/api/reflect.mjs';
|
|
import reflectSh from '!!raw-loader!@site/examples/api/reflect.sh';
|
|
import reflectGo from '!!raw-loader!@site/examples/api/reflect.go';
|
|
|
|
:::info How Reflect Works
|
|
Learn about disposition-driven reasoning in the [Reflect Architecture](/developer/reflect) guide.
|
|
:::
|
|
|
|
:::tip Prerequisites
|
|
Make sure you've completed the [Quick Start](./quickstart) to install the client and start the server.
|
|
:::
|
|
|
|
## Basic Usage
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-basic" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-basic" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
<CodeSnippet code={reflectSh} section="reflect-basic" language="bash" />
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={reflectGo} section="reflect-basic" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
---
|
|
|
|
## Parameters
|
|
|
|
### query
|
|
|
|
The question or prompt to reflect on. This is the only required field. If you have situational context that should influence the answer, include it directly in the query rather than as a separate field.
|
|
|
|
### budget
|
|
|
|
Controls how thoroughly the agent explores the memory bank before answering. Accepted values are `low` (default), `mid`, and `high`. At `low`, the agent does a shallow search optimized for speed. At `mid`, it checks multiple sources when the question warrants it. At `high`, it performs deep exploration across all knowledge levels and may use multiple query variations to find indirect connections. Use `high` for complex questions that require synthesizing information from many sources.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-with-params" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-with-params" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
<CodeSnippet code={reflectSh} section="reflect-with-params" language="bash" />
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={reflectGo} section="reflect-with-params" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
### max_tokens
|
|
|
|
Limits the length of the final generated response. Defaults to `4096`. This does not affect how much the agent can retrieve during the agentic loop — only the final answer length.
|
|
|
|
### response_schema
|
|
|
|
An optional JSON Schema **object** with a non-empty `properties` map (nested objects and arrays are supported). When provided, the response includes a `structured_output` field **in addition to** the markdown `text`: the agent reasons to its answer, then a second pass extracts that answer into JSON matching your schema. `structured_output` is therefore a faithful projection of `text` — you get the readable answer *and* a typed object to program against, never one instead of the other. Invalid schemas (not an object, or no properties) are rejected before the call runs.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-structured-output" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-structured-output" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
<CodeSnippet code={reflectSh} section="reflect-structured-output" language="bash" />
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={reflectGo} section="reflect-structured-output" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
### tags
|
|
|
|
Defines the visibility scope used throughout the reflect agent. It filters raw
|
|
facts, observations, and mental models that the agent can retrieve. The same
|
|
`tags` and `tags_match` values also select which tagged directives are injected
|
|
into the reflect prompt.
|
|
|
|
`tags` defaults to `null`, `tags_match` defaults to `any`, and `tag_groups`
|
|
defaults to `null`. For non-empty tags, raw facts, observations, and mental
|
|
models use the same matching modes as [recall tags](./recall#tags). Directives
|
|
have one additional rule: untagged directives are global and remain eligible
|
|
whenever a tag scope is supplied, including with a strict or exact match.
|
|
|
|
| Reflect configuration | Raw facts and observations | Mental models | Active directives |
|
|
|-----------------------|----------------------------|---------------|-------------------|
|
|
| Omit `tags`, `tags_match`, and `tag_groups` | All tagged and untagged data | All tagged and untagged models | Untagged/global directives only |
|
|
| `tags: []`, default `tags_match: "any"` | All tagged and untagged data | All tagged and untagged models | Untagged/global directives only |
|
|
| No tags, `tags_match: "exact"` | Untagged/global data only | Untagged/global models only | Untagged/global directives only |
|
|
| Non-empty `tags`, `any` or `all` | Matching tagged data plus untagged/global data | Matching models plus untagged/global models | Matching tagged directives plus untagged/global directives |
|
|
| Non-empty `tags`, `any_strict` or `all_strict` | Matching tagged data only | Matching tagged models only | Matching tagged directives plus untagged/global directives |
|
|
| Non-empty `tags`, `exact` | Data with exactly the requested tag set | Models with exactly the requested tag set | Exactly matching tagged directives plus untagged/global directives |
|
|
| Non-empty `tag_groups`, default top-level `tags_match` | Data matching the compound expression | Models matching the compound expression | Matching tagged directives plus untagged/global directives |
|
|
|
|
A `tag_groups` leaf may set `resolve: "fuzzy"`, as in
|
|
[recall](./recall#fuzzy-leaves); reflect resolves it once, before the agentic
|
|
loop starts, so every tool it runs filters on the same tags.
|
|
|
|
The first row is intentionally asymmetric: an unscoped reflect can search all
|
|
memories, but it does not load tagged directives. To create a directive that
|
|
applies to every reflect call, leave its `tags` empty. To create a scoped
|
|
directive, assign tags and pass a matching scope to `reflect`.
|
|
|
|
:::note `isolation_mode`
|
|
`isolation_mode` is an internal `list_directives` option, not a public reflect
|
|
request parameter. Reflect always enables it. When neither `tags` nor
|
|
`tag_groups` is supplied, it limits directive loading to untagged directives.
|
|
There is currently no per-request switch to disable it.
|
|
:::
|
|
|
|
:::note MCP omitted tags
|
|
The MCP `reflect` tool forwards `tags_match` only when `tags` is present. To
|
|
request the empty exact scope through MCP, pass `tags: []` together with
|
|
`tags_match: "exact"`.
|
|
:::
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-with-tags" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-with-tags" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
<CodeSnippet code={reflectSh} section="reflect-with-tags" language="bash" />
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={reflectGo} section="reflect-with-tags" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
#### Common scope examples
|
|
|
|
Unscoped reflect searches all memories but applies only global directives:
|
|
|
|
```json
|
|
{
|
|
"query": "Summarize the current project status"
|
|
}
|
|
```
|
|
|
|
A project scope includes global data and directives alongside matching
|
|
`project:a` data and directives:
|
|
|
|
```json
|
|
{
|
|
"query": "Summarize the current project status",
|
|
"tags": ["project:a"]
|
|
}
|
|
```
|
|
|
|
A strict project scope excludes untagged memories, observations, and mental
|
|
models. Global directives still apply:
|
|
|
|
```json
|
|
{
|
|
"query": "Summarize the current project status",
|
|
"tags": ["project:a"],
|
|
"tags_match": "all_strict"
|
|
}
|
|
```
|
|
|
|
### tag_groups
|
|
|
|
Provides compound tag filtering with recursive `and`, `or`, and `not`
|
|
expressions. It affects the same reflect data sources and directive selection as
|
|
flat `tags`. `tag_groups` and `tags` are mutually exclusive in the public REST
|
|
request. Each leaf supplies its own matching mode and defaults to
|
|
`any_strict`; the top-level groups are AND-ed. Normally leave the top-level
|
|
`tags_match` at its default, `any`. Setting it to `exact` while using
|
|
`tag_groups` additionally constrains facts, observations, and mental models to
|
|
the global flat scope before applying the compound expression.
|
|
|
|
The MCP `reflect` tool currently exposes flat `tags` and `tags_match`, but not
|
|
`tag_groups`.
|
|
|
|
### include
|
|
|
|
Controls optional supplementary data returned alongside the main response.
|
|
|
|
#### include.facts
|
|
|
|
When enabled, the response includes a `based_on` object listing the memories, mental models, and directives the agent actually used to construct the answer. Only sources retrieved during the agent loop can appear here — citations are validated to prevent hallucinated references. Useful for transparency and verification.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-sources" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-sources" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
<CodeSnippet code={reflectSh} section="reflect-sources" language="bash" />
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={reflectGo} section="reflect-sources" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
#### include.tool_calls
|
|
|
|
When enabled, the response includes a `trace` object with the full execution log of every tool call and LLM call made during the agentic loop, including inputs, outputs, and durations. Set `output: false` to include only tool inputs for a smaller payload. Useful for debugging why the agent reached a particular conclusion.
|
|
|
|
---
|
|
|
|
## Response
|
|
|
|
### text
|
|
|
|
The synthesized answer as a well-formatted markdown string. This is the primary output of reflect. Still returned when `response_schema` is provided — `structured_output` is derived from it, not a replacement for it.
|
|
|
|
### structured_output
|
|
|
|
The LLM's response parsed according to the `response_schema` provided in the request. Only present when `response_schema` was set. `null` otherwise.
|
|
|
|
### based_on
|
|
|
|
The sources the agent used to construct the answer. Only present when `include.facts` was enabled. Contains three fields:
|
|
|
|
- `memories` — a list of memory facts (world, experience, observation) that were retrieved and cited. Each item has `id`, `text`, `type`, `context`, `occurred_start`, and `occurred_end`.
|
|
- `mental_models` — a list of mental models that were used. Each item has `id`, `text`, and `context`.
|
|
- `directives` — a list of directives that were enforced during reasoning. Each item has `id`, `name`, and `content`.
|
|
|
|
### usage
|
|
|
|
Token usage for all LLM calls made during the agentic loop: `input_tokens`, `output_tokens`, and `total_tokens`. Useful for cost tracking.
|
|
|
|
### trace
|
|
|
|
The full execution log of the agentic loop. Only present when `include.tool_calls` was enabled. Contains:
|
|
|
|
- `tool_calls` — each tool invocation with `tool` name (`lookup`, `recall`, `learn`, `expand`), `input`, `output` (if `output: true`), `duration_ms`, and `iteration` number.
|
|
- `llm_calls` — each LLM call with `scope` (e.g., `"agent_1"`, `"final"`) and `duration_ms`.
|
|
|
|
## When Reflect Fails
|
|
|
|
Reflect answers from evidence it gathered, so a run that could not gather it does not answer at all — it fails with a **500**, rather than returning a confident reply built on nothing:
|
|
|
|
- **A retrieval tool raised.** The database, the embedder or the reranker was unavailable. One failure in a batch fails the run: an answer written around a tool that never returned is indistinguishable from one over a bank that genuinely holds nothing on the topic, and callers store it as a real answer.
|
|
- **The model produced no answer.** The `done` call arrived empty, or the final synthesis returned nothing.
|
|
- **The model or transport cannot drive tool calls.** Reflect is driven entirely by structured tool calls; a transport that silently drops the tool definitions fails loudly so you can switch to a tool-calling-capable model.
|
|
- **The provider kept failing.** A non-context-overflow LLM error is retried once inside the loop and then given up on.
|
|
|
|
A **successful retrieval that returns nothing is not a failure**: the bank genuinely has nothing on the topic, and reflect says so in its answer.
|
|
|
|
Two cases are deliberately not failures. A run whose *context window* overflows synthesizes from the evidence it has — the prompt was too big for the model, which is a budgeting problem, not a broken dependency. And a tool call the model got *wrong* — a missing argument, a tool that does not exist — is returned to it as an error to fix, not raised.
|