Files
Nicolò Boschi d20893a83f fix(api): paginate GET /observations/scopes and GET /webhooks (#4075)
* fix(api): paginate GET /observations/scopes and GET /webhooks

Both endpoints returned the whole collection: the scopes endpoint grouped
every observation in a bank and shipped one row per distinct tag set, and
the webhooks endpoint selected every webhook row. Neither is bounded by
construction — a bank has as many scopes as it has distinct tag sets — so
the payload and the per-row work grew with the data.

Both now take `limit` (default 100, max 1000) and `offset` and return
`total` / `limit` / `offset` alongside the page, the same shape
`list_tags` and the other paged list endpoints use. The bound reaches the
SQL: the scope histogram is grouped, ordered and paged in one query with
a separate `COUNT(DISTINCT scope)` for `total`, and the webhook listing
gets `LIMIT/OFFSET` plus a count query. Webhook rows now order by
`created_at, id` so a page boundary cannot fall inside a group of rows
sharing a timestamp.

Consumers page with them: the control-plane webhooks view walks every
page (it renders the whole list and its count), the two proxy routes
forward `limit`/`offset`, and the CLI prints "Showing N of M webhooks"
rather than presenting the first page as the whole list.

Clients, OpenAPI spec and the docs skill regenerated.

* fix(tests): page the in-memory store's observation_scope_counts stub

The MemoriesExtension conformance test asserts every stub method accepts
the interface's params, so widening the interface with limit/offset left
the InMemoryMemories stub behind (test-api shards 2/3 and free-threaded).
It now groups, orders and pages its own observation rows and returns the
same {scopes, total, limit, offset} shape as its list_tags neighbour,
rather than the bare list it used to hand back.

* fix(control-plane): page the webhooks view instead of loading every page

The first cut walked every page in loadWebhooks and rendered the lot, so
the bound existed but nothing in the UI showed it — a bank with hundreds
of webhooks still built one unbounded table, and a reader had no way to
tell there was more than a screenful.

It now fetches one 50-row page and exposes the pager entities-view uses:
first/prev/next/last, "1 / 3", and a "51-100 of 120" range line, with the
header count coming from `total` rather than the loaded page. The controls
hide when everything fits on one page.

Two paging hazards handled: a bank switch resets to page 1 (the page
number belongs to the bank being left, and carrying it over lands on an
offset the new bank may not reach), and creating a webhook jumps to the
page it lands on — rows are oldest-first, so reloading the current page
would close the dialog and appear to do nothing. Deleting the last row of
the final page steps back a page rather than stranding an empty one.
2026-09-04 10:34:54 +02:00

276 lines
15 KiB
Plaintext

---
sidebar_position: 5
---
import CodeSnippet from '@site/src/components/CodeSnippet';
import recallPy from '!!raw-loader!@site/examples/api/recall.py';
import memoryBanksPy from '!!raw-loader!@site/examples/api/memory-banks.py';
# Observations: Knowledge Consolidation
After memories are retained, Hindsight automatically consolidates related facts into **observations** — deduplicated, evidence-grounded beliefs the bank has built up from multiple memories. Each observation tracks its supporting evidence (with exact quotes) and a proof count, and is refined rather than overwritten when new evidence arrives.
```mermaid
graph LR
A[New Facts] --> B[Consolidation Engine]
B --> C{Existing Observation?}
C -->|Yes| D[Refine Observation]
C -->|No| E[Create Observation]
D --> F[Observations]
E --> F
```
---
## What Are Observations?
Observations are **consolidated knowledge** built from multiple facts. Unlike raw facts — which are individual pieces of information — observations represent deduplicated beliefs, preferences, and learnings grounded in accumulated evidence. They are not summaries the LLM invents on the fly: each observation is backed by specific source memories, carries a proof count, and evolves as new evidence supports, contradicts, or extends it.
| Raw Facts | Observation |
|-----------|--------------|
| "Alice prefers Python" | "Alice is a Python-focused developer who values readability and simplicity" |
| "Alice dislikes verbose code" | |
| "Alice recommends type hints" | |
Observations provide:
- **Deduplication**: One durable belief instead of many overlapping facts
- **Grounding**: Every observation references the specific memories (with quotes) that support it
- **Evolution**: Refined as evidence strengthens, weakens, or contradicts it — history is preserved
- **Freshness awareness**: when newer memories haven't been consolidated yet, `reflect` treats the affected observations as stale and verifies them against raw facts
- **Efficiency**: Condensed knowledge for faster retrieval
---
## How Consolidation Works
### Automatic Background Processing
After `retain()` completes, the consolidation engine runs automatically:
1. **New facts analyzed** — Each new fact is compared against existing observations
2. **Pattern detection** — Related facts are grouped and synthesized
3. **Observation creation/update** — New observations are created or existing ones refined
4. **Evidence tracking** — Each observation maintains references to supporting facts
### Near-Duplicate Reconciliation
Consolidation can still produce two observations that say the same thing in slightly different words — for example when a weaker model writes a near-identical observation instead of refining the existing one, or when refining an observation reshapes its wording so it overlaps another one. Left alone, these near-duplicates clutter recall with redundant beliefs.
When enabled, Hindsight reconciles them automatically. Whenever an observation is created **or** updated, it is compared against the existing observations it most closely resembles. If one is highly similar, a focused check decides whether to **merge** them into a single belief (folding both sets of supporting evidence together) or **keep** them separate. Because the check reads the full text of both, observations that differ in a meaningful detail — a number, a negation, a named entity or language — are correctly kept apart rather than collapsed.
This is controlled by the [`HINDSIGHT_API_CONSOLIDATION_DEDUP_THRESHOLD`](/developer/configuration#observations) setting: the cosine similarity at or above which two observations are reconciled. It is **enabled by default** (`0.97`); a lower value reconciles more aggressively, and `1.0` disables it. Reconciliation runs on PostgreSQL deployments only — it is skipped on Oracle regardless of the threshold.
Reconciliation only compares observations **within the same tag scope**. If you tag retains with a unique per-call value (e.g. a `session-id`), each session lands in its own scope and never dedups against the others — producing one near-identical observation per session. To consolidate across those volatile tags, retain with [`observation_scopes: "shared"`](/developer/api/retain#shared), which scopes observations to one global, untagged belief while leaving the session tag on the source facts for recall filtering.
### Disabling Auto-Consolidation
Set `HINDSIGHT_API_ENABLE_AUTO_CONSOLIDATION=false` (or configure per-bank via the [bank config API](/developer/api/memory-banks#observations-configuration)) to prevent consolidation from running automatically after retain, delete, and update operations. When disabled, consolidation only runs when you explicitly call the [consolidate endpoint](#trigger-consolidation).
This is useful when you want full control over consolidation timing — for example, batching many retains before consolidating, or running consolidation only for specific scopes.
### Targeted Consolidation
By default, consolidation processes **all** unconsolidated memories in a bank. You can scope it to specific tag sets using the `observation_scopes` parameter on the consolidate endpoint:
```python
# Consolidate only memories tagged with user:alice
client.consolidate(
bank_id="my-bank",
observation_scopes=[["user:alice"]]
)
# Consolidate memories for alice OR the engineering team
client.consolidate(
bank_id="my-bank",
observation_scopes=[["user:alice"], ["team:engineering"]]
)
```
Each scope is a list of tags. A memory matches a scope if its tags **contain all** tags in that scope. For example, scope `["user:alice"]` matches memories tagged `["user:alice", "team:eng"]`.
When `observation_scopes` is omitted, all unconsolidated memories are processed (backward compatible).
### Evidence-Based Evolution
Observations evolve as new evidence arrives:
| Event | What the bank learns | Observation state |
|-------|---------------------|----------------|
| **Day 1** | "Redis is open source under BSD license" | "Redis is excellent for caching — fast, reliable, and OSS-friendly" (2 supporting facts) |
| **Day 2** | "Redis has great community support" | Observation reinforced (3 supporting facts) |
| **Day 30** | "Redis changed license to SSPL" | Observation refined: "Redis is technically strong, but has license concerns for cloud" |
| **Day 45** | "Valkey forked Redis under BSD" | New observation: "Consider Valkey for new projects requiring true OSS" |
### Handling Contradictory Evidence
What happens when a new fact contradicts an existing observation?
The consolidation engine doesn't blindly overwrite — it **reconciles** the contradiction by capturing the evolution:
**Example: User preference changes**
| Time | Fact | Observation |
|------|------|--------------|
| Week 1 | "User says they love React" | "User prefers React for frontend development" |
| Week 2 | "User praises React's component model" | "User is enthusiastic about React, particularly its component model" |
| Week 3 | "User says they've switched to Vue and won't use React anymore" | "User was previously a React enthusiast who appreciated its component model, but has now switched to Vue and no longer uses React" |
Notice how the final observation captures the **full journey** — not just "User prefers Vue" but the complete evolution of their preference. This nuanced understanding means:
- Your agent won't recommend React tutorials to someone who explicitly moved away from it
- Your agent understands *why* this matters (they were enthusiastic before, so this is a deliberate choice)
- Your agent can reference this history when relevant ("I know you used to work with React...")
The system:
1. **Detects the conflict** — New fact contradicts existing observation
2. **Preserves history** — Incorporates the previous understanding into the new observation
3. **Creates nuanced observation** — Synthesizes a richer understanding that captures the change
4. **Updates freshness** — Marks the observation as recently updated
**Example: Correcting misinformation**
| Time | Fact | Observation |
|------|------|--------------|
| Day 1 | "Alice works at Google" | "Alice is a Google employee" |
| Day 10 | "Alice actually works at Meta, not Google" | "Alice works at Meta (previously thought to work at Google)" |
When a fact explicitly corrects previous information, the observation is updated to reflect the correction while noting the previous understanding. The raw facts are always preserved, so you can trace back to see what was originally stated and when it was corrected.
---
## Observations in Retrieval
Observations are automatically included in both `recall()` and `reflect()` operations:
### In Recall
Observations are returned alongside raw facts, filtered by the `types` parameter:
<CodeSnippet code={recallPy} section="recall-with-observations" language="python" />
### In Reflect
The reflect agent uses **hierarchical retrieval**:
1. **[Mental Models](/developer/api/mental-models)** — User-curated summaries (highest priority)
2. **Observations** — Consolidated knowledge with freshness awareness
3. **Raw Facts** — Ground truth for verification
The agent automatically queries observations and uses them to inform its reasoning.
---
## Freshness Awareness
Observations track when they were last updated. During reflect, the agent considers freshness:
- **Fresh observations**: Used directly for reasoning
- **Stale observations**: Agent verifies against current facts before relying on them
This ensures responses stay accurate even as the underlying data changes.
---
## Observation Scopes
By default, observations are scoped to all of a memory's tags combined. The `observation_scopes` retain parameter lets you control this — building separate observations per tag, per combination, or with a custom list of scopes. This is key when a single memory carries multiple tags and you want each tag to accumulate its own observations independently.
See [`observation_scopes` in the Retain API](./api/retain#observation_scopes) for the full explanation and options.
To inspect the scopes that already exist in a bank, call `GET /v1/default/banks/{bank_id}/observations/scopes`. The response lists each exact tag set with its observation count; the empty tag list is the global scope. The listing is paged (`limit`, default 100, and `offset`), with `total` reporting how many distinct scopes the bank holds. Use a returned scope as `tags` with `tags_match: "exact"` when you need to filter to that precise observation scope without also matching observations that carry extra tags. To recall **only** the global scope — the untagged observations written by `observation_scopes: "shared"` — pass an empty list with exact matching: `tags: []`, `tags_match: "exact"`.
---
## Observations Mission
You can define exactly what this bank should synthesise by setting an **observations mission** (`observations_mission`). This replaces the built-in durable-knowledge rules with your own instructions, letting you control what shape observations take.
```
e.g. Observations are stable facts about people and projects.
Always include preferences, skills, and recurring patterns.
Ignore one-off events and ephemeral state.
```
Leave it blank to use the server default — durable, specific facts that stay true over time (preferences, skills, relationships, recurring patterns), with ephemeral state filtered out.
**Examples:**
| `observations_mission` | What gets synthesised |
|------------------------|----------------------|
| *(unset — default)* | Durable facts: preferences, skills, relationships, recurring patterns |
| *"Observations are weekly summaries of sprint outcomes and blockers"* | Broad event summaries grouped by time period |
| *"Observations are stable facts about named individuals only"* | Person-centric knowledge, tied to specific people |
| *"Observations are recurring patterns in customer support interactions"* | Failure modes, common requests, pain points |
Set `observations_mission` via the [bank config API](/developer/api/memory-banks#observations-configuration) or the [`HINDSIGHT_API_OBSERVATIONS_MISSION`](/developer/configuration#observations) environment variable.
---
## Observation Lifecycle & Invalidation
### When Memories Are Deleted
Observations are derived from source memories. When source memories are removed, Hindsight automatically keeps observations consistent:
| Action | Effect on observations |
|--------|----------------------|
| Delete a document | All observations derived from the document's memories are deleted |
| Delete individual memories (by type) | Observations sourced from those memories are deleted |
| Delete an entire bank | All observations are deleted along with everything else |
After deletion, the **remaining source memories** that fed the affected observations have their consolidation state reset, so they will be re-consolidated on the next consolidation run and produce fresh observations.
### Clearing Observations for a Specific Memory
You can clear all observations derived from a single memory without deleting the memory itself. This is useful when you want to force re-synthesis of a memory's contribution to consolidated knowledge.
Use the `DELETE /v1/default/banks/{bank_id}/memories/{memory_id}/observations` endpoint. This will:
1. Delete all observations that list the memory as a source
2. Reset `consolidated_at` on the memory itself and any other source memories that contributed to those observations
3. Trigger a consolidation job so fresh observations are produced automatically
### Resetting All Observations
To wipe all consolidated knowledge and start over:
```python
# Clear all observations for a bank
client.clear_observations(bank_id="my-bank")
```
This resets the consolidation state for all source memories in the bank, so the next consolidation run will re-derive all observations from scratch.
---
## Trigger Consolidation {#trigger-consolidation}
Use the consolidate endpoint to manually trigger consolidation:
```http
POST /v1/default/banks/{bank_id}/consolidate
Content-Type: application/json
{
"observation_scopes": [["user:alice"], ["team:engineering"]]
}
```
The request body is optional. When omitted (or sent as an empty body), all unconsolidated memories in the bank are processed.
| Parameter | Type | Description |
|-----------|------|-------------|
| `observation_scopes` | `list[list[str]]` \| `null` | Optional list of tag scopes. Only memories whose tags contain all tags in at least one scope are processed. Omit for a full-bank sweep. |
## Configuration
Observation consolidation runs automatically by default. You can disable auto-consolidation with [`HINDSIGHT_API_ENABLE_AUTO_CONSOLIDATION`](/developer/configuration#observations) and trigger it on-demand via the [consolidate endpoint](#trigger-consolidation). Monitor consolidation progress via the [Operations API](./api/operations).
---
## Next Steps
- [**Retain**](./retain) — How facts are stored and trigger consolidation
- [**Recall**](./retrieval) — How observations are retrieved
- [**Reflect**](./reflect) — How the agentic loop uses observations
- [**Mental Models**](./api/mental-models) — User-curated summaries for common queries