Files
Nicolò Boschi 43f545e9b8 fix(documents): skip the retag cascade when a tags PATCH changes nothing (#3912) (#3931)
* fix(documents): skip the retag cascade when a tags PATCH changes nothing (#3912)

`update_document` never compared the incoming tags to the ones the document
already carries — the path was `if tags is not None:`. So a PATCH re-sending an
identical tags array, which is exactly what an idempotent tag-normalisation
sweep does on every run after its first, paid the full retag cascade for a write
that changes nothing.

That cascade is not cheap and is not meant to be: it deletes every observation
built over the document's memories and then resets `consolidated_at` on each of
those observations' OTHER sources. Both halves are required for correctness —
consolidation scopes a memory by its tag set, so an observation formed under the
old tags is no longer valid, and deleting it strands the co-sourced memories it
carried unless they are requeued. The consequence is that the blast radius of
one PATCH is the co-source degree of the affected observations, not the
document's own memory count, and a sweep multiplies that by its document count.

Read the current tags before overwriting them and compare as SETS —
consolidation scopes by tag set, so a reordered array changes nothing it can
observe. When the set is unchanged, skip the memory-unit retag, the observation
deletion, the source requeue and the consolidation submit. A tag set that
differs at all still runs the cascade unchanged; tags that could not be read
(document absent, or a store record not carrying them) are never treated as
unchanged, so a real retag is never silently skipped. Both the SQL and
store-owned branches are covered.

Also drops a redundant `get_document_record` round-trip on the store-owned path,
which the new pre-read already fetched.

Tests: a repeat PATCH leaves the observation and every co-source consolidated; a
reordered array is not a change; a superset still invalidates; and a real retag
stamps `memory_units.updated_at` while a no-op leaves it alone.

* chore(docs-skill): regenerate for the documents.mdx re-consolidation note

The bundled hindsight-docs skill is generated from hindsight-docs/docs/**;
verify-generated-files caught references/developer/api/documents.md drifting
from the callout edited in the previous commit.

Claude-Session: https://claude.ai/code/session_018HDqrzHgqZqsGc7EDqoTEu
2026-08-31 16:16:24 +02:00

253 lines
8.2 KiB
Plaintext

---
sidebar_position: 8
---
# Documents
Track and manage document sources in your memory bank. Documents provide traceability — knowing where memories came from.
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
import CodeSnippet from '@site/src/components/CodeSnippet';
{/* Import raw source files */}
import documentsPy from '!!raw-loader!@site/examples/api/documents.py';
import documentsMjs from '!!raw-loader!@site/examples/api/documents.mjs';
import documentsGo from '!!raw-loader!@site/examples/api/documents.go';
:::tip Prerequisites
Make sure you've completed the [Quick Start](./quickstart) and understand [how retain works](./retain).
:::
## What Are Documents?
Documents are containers for retained content. They help you:
- **Track sources** — Know which PDF, conversation, or file a memory came from
- **Update content** — Re-retain a document to update its facts
- **Delete in bulk** — Remove all memories from a document at once
- **Organize memories** — Group related facts by source
## Chunks
When you retain content, Hindsight splits it into chunks before extracting facts. These chunks are stored alongside the extracted memories, preserving the original text segments.
**Why chunks matter:**
- **Context preservation** — Chunks contain the raw text that generated facts, useful when you need the exact wording
- **Richer recall** — Including chunks in recall provides surrounding context for matched facts
:::tip Include Chunks in Recall
Use `include_chunks=True` in your recall calls to get the original text chunks alongside fact results. See [Recall](./recall) for details.
:::
## Retain with Document ID
Associate retained content with a document:
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-retain" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-retain" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# Retain content with document ID
hindsight memory retain my-bank "Meeting notes content..." --doc-id notes-2024-03-15
# Batch retain from files
hindsight memory retain-files my-bank docs/
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-retain" language="go" />
</TabItem>
</Tabs>
## Update Documents
Re-retaining with the same document_id **replaces** the old content:
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-update" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# Original
hindsight memory retain my-bank "Project deadline: March 31" --doc-id project-plan
# Update
hindsight memory retain my-bank "Project deadline: April 15 (extended)" --doc-id project-plan
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-update" language="go" />
</TabItem>
</Tabs>
## Get Document
Retrieve a document's original text and metadata. This is useful for expanding document context after a recall operation returns memories with document references.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-get" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-get" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
hindsight document get my-bank meeting-2024-03-15
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-get" language="go" />
</TabItem>
</Tabs>
## Update Document
Update mutable fields on an existing document without re-processing the content. Currently supports updating `tags`.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-update" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# Replace tags with new values
hindsight document update-tags my-bank meeting-2024-03-15 --tags team-a --tags team-b
# Remove all tags
hindsight document update-tags my-bank meeting-2024-03-15
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-update" language="go" />
</TabItem>
</Tabs>
:::info Observations are re-consolidated
When tags change, any consolidated observations derived from the document's memories are invalidated and queued for re-consolidation under the new tags. Co-source memories from other documents that shared those observations are also reset.
This is required for correctness rather than incidental: consolidation scopes a memory by its tag set, so an observation built under the old tags is no longer valid, and deleting it would strand every other memory that observation was consolidated from unless those are requeued too. The size of that requeue is the number of memories co-sourced with this document's — on a densely co-sourced bank it can be many times the document's own memory count.
Tags are compared as a **set** against the document's current tags, and an update that leaves the set unchanged — including one that only reorders the array — performs no retag and queues no re-consolidation. A repeatable tag-normalisation sweep therefore only pays the re-consolidation cost on the run that actually changes something.
:::
## Delete Document
Remove a document and all its associated memories:
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-delete" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-delete" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
hindsight document delete my-bank meeting-2024-03-15
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-delete" language="go" />
</TabItem>
</Tabs>
:::warning
Deleting a document permanently removes all memories extracted from it. This action cannot be undone.
:::
## List Documents
List documents in a bank with optional filtering by ID and tags.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-list" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-list" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# List all documents
hindsight document list my-bank
# Filter by ID substring
hindsight document list my-bank --q report
# Filter by tags
hindsight document list my-bank --tags team-a --tags team-b
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-list" language="go" />
</TabItem>
</Tabs>
### Filtering Options
| Parameter | Description |
|---|---|
| `q` | Case-insensitive substring match on document ID. `report` matches `report-2024`, `annual-report`, etc. |
| `tags` | Filter by document tags. Accepts multiple values. |
| `tags_match` | How to match tags (default: `any_strict`). See below. |
| `limit` / `offset` | Pagination. Default limit is 100. |
**`tags_match` modes:**
| Mode | Behaviour |
|---|---|
| `any_strict` *(default)* | Document must have **at least one** of the specified tags. Untagged docs excluded. |
| `any` | Same as `any_strict` but also includes untagged documents. |
| `all_strict` | Document must have **all** specified tags. Untagged docs excluded. |
| `all` | Same as `all_strict` but also includes untagged documents. |
## Document Response Format
```json
{
"id": "meeting-2024-03-15",
"bank_id": "my-bank",
"original_text": "Alice presented the Q4 roadmap...",
"content_hash": "abc123def456",
"memory_unit_count": 12,
"nodes_by_fact_type": {
"world": 5,
"experience": 4,
"observation": 3
},
"created_at": "2024-03-15T14:00:00Z",
"updated_at": "2024-03-15T14:00:00Z"
}
```
## Next Steps
- [**Operations**](./operations) — Monitor background tasks
- [**Memory Banks**](./memory-banks) — Configure bank settings