Files
Nicolò Boschi 565303d913 docs(documents): say the tags PATCH replaces the array, and test clearing it (#4272)
* test(documents): cover clearing a document's tags with an empty array

The tags PATCH replaces the array rather than merging it, so `tags: []` is how
a caller drops every tag. Every guard on that path is written `is not None`
rather than a truthiness check so the empty list survives it, but nothing
exercised it: a regression to `if tags:` would have turned a clear into a
silent no-op and a 200.

Adds an engine test (clears the document's tags and its units', runs the same
observation-invalidation cascade, and is a no-op when repeated) and an HTTP test
(PATCH `{"tags": []}` is 200, an omitted `tags` is still 422).

* docs(documents): say that the tags PATCH replaces the array

The endpoint description and the docs page both said only that tags are
"propagated to all associated memory units", which leaves the question a caller
actually has — does sending a tag ADD it, and how do I drop one — unanswered.
The replace semantics were documented in exactly one place: two comments in the
CLI tab of the docs page, which an API or SDK user never reads.

States it where they will see it: the array replaces rather than merges, an
omitted tag is dropped, `[]` clears them all, and only an omitted FIELD is the
422. Regenerates the spec, the clients and the docs skill.
2026-09-09 17:13:50 +02:00

255 lines
8.6 KiB
Plaintext

---
sidebar_position: 8
---
# Documents
Track and manage document sources in your memory bank. Documents provide traceability — knowing where memories came from.
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
import CodeSnippet from '@site/src/components/CodeSnippet';
{/* Import raw source files */}
import documentsPy from '!!raw-loader!@site/examples/api/documents.py';
import documentsMjs from '!!raw-loader!@site/examples/api/documents.mjs';
import documentsGo from '!!raw-loader!@site/examples/api/documents.go';
:::tip Prerequisites
Make sure you've completed the [Quick Start](./quickstart) and understand [how retain works](./retain).
:::
## What Are Documents?
Documents are containers for retained content. They help you:
- **Track sources** — Know which PDF, conversation, or file a memory came from
- **Update content** — Re-retain a document to update its facts
- **Delete in bulk** — Remove all memories from a document at once
- **Organize memories** — Group related facts by source
## Chunks
When you retain content, Hindsight splits it into chunks before extracting facts. These chunks are stored alongside the extracted memories, preserving the original text segments.
**Why chunks matter:**
- **Context preservation** — Chunks contain the raw text that generated facts, useful when you need the exact wording
- **Richer recall** — Including chunks in recall provides surrounding context for matched facts
:::tip Include Chunks in Recall
Use `include_chunks=True` in your recall calls to get the original text chunks alongside fact results. See [Recall](./recall) for details.
:::
## Retain with Document ID
Associate retained content with a document:
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-retain" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-retain" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# Retain content with document ID
hindsight memory retain my-bank "Meeting notes content..." --doc-id notes-2024-03-15
# Batch retain from files
hindsight memory retain-files my-bank docs/
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-retain" language="go" />
</TabItem>
</Tabs>
## Update Documents
Re-retaining with the same document_id **replaces** the old content:
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-update" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# Original
hindsight memory retain my-bank "Project deadline: March 31" --doc-id project-plan
# Update
hindsight memory retain my-bank "Project deadline: April 15 (extended)" --doc-id project-plan
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-update" language="go" />
</TabItem>
</Tabs>
## Get Document
Retrieve a document's original text and metadata. This is useful for expanding document context after a recall operation returns memories with document references.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-get" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-get" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
hindsight document get my-bank meeting-2024-03-15
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-get" language="go" />
</TabItem>
</Tabs>
## Update Document
Update mutable fields on an existing document without re-processing the content. Currently supports updating `tags`.
The `tags` array **replaces** the document's tags — it is not merged into them. Send the complete set you want the document to end up with: any tag you leave out is dropped, and an empty array clears them all. To remove a single tag, read the document's current tags, drop the one you want gone, and send the rest. Omitting the field entirely is not an update at all and is rejected with a `422`.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-update" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# Replace tags with new values
hindsight document update-tags my-bank meeting-2024-03-15 --tags team-a --tags team-b
# Remove all tags
hindsight document update-tags my-bank meeting-2024-03-15
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-update" language="go" />
</TabItem>
</Tabs>
:::info Observations are re-consolidated
When tags change, any consolidated observations derived from the document's memories are invalidated and queued for re-consolidation under the new tags. Co-source memories from other documents that shared those observations are also reset.
This is required for correctness rather than incidental: consolidation scopes a memory by its tag set, so an observation built under the old tags is no longer valid, and deleting it would strand every other memory that observation was consolidated from unless those are requeued too. The size of that requeue is the number of memories co-sourced with this document's — on a densely co-sourced bank it can be many times the document's own memory count.
Tags are compared as a **set** against the document's current tags, and an update that leaves the set unchanged — including one that only reorders the array — performs no retag and queues no re-consolidation. A repeatable tag-normalisation sweep therefore only pays the re-consolidation cost on the run that actually changes something.
:::
## Delete Document
Remove a document and all its associated memories:
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-delete" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-delete" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
hindsight document delete my-bank meeting-2024-03-15
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-delete" language="go" />
</TabItem>
</Tabs>
:::warning
Deleting a document permanently removes all memories extracted from it. This action cannot be undone.
:::
## List Documents
List documents in a bank with optional filtering by ID and tags.
<Tabs>
<TabItem value="python" label="Python">
<CodeSnippet code={documentsPy} section="document-list" language="python" />
</TabItem>
<TabItem value="node" label="Node.js">
<CodeSnippet code={documentsMjs} section="document-list" language="javascript" />
</TabItem>
<TabItem value="cli" label="CLI">
```bash
# List all documents
hindsight document list my-bank
# Filter by ID substring
hindsight document list my-bank --q report
# Filter by tags
hindsight document list my-bank --tags team-a --tags team-b
```
</TabItem>
<TabItem value="go" label="Go">
<CodeSnippet code={documentsGo} section="document-list" language="go" />
</TabItem>
</Tabs>
### Filtering Options
| Parameter | Description |
|---|---|
| `q` | Case-insensitive substring match on document ID. `report` matches `report-2024`, `annual-report`, etc. |
| `tags` | Filter by document tags. Accepts multiple values. |
| `tags_match` | How to match tags (default: `any_strict`). See below. |
| `limit` / `offset` | Pagination. Default limit is 100. |
**`tags_match` modes:**
| Mode | Behaviour |
|---|---|
| `any_strict` *(default)* | Document must have **at least one** of the specified tags. Untagged docs excluded. |
| `any` | Same as `any_strict` but also includes untagged documents. |
| `all_strict` | Document must have **all** specified tags. Untagged docs excluded. |
| `all` | Same as `all_strict` but also includes untagged documents. |
## Document Response Format
```json
{
"id": "meeting-2024-03-15",
"bank_id": "my-bank",
"original_text": "Alice presented the Q4 roadmap...",
"content_hash": "abc123def456",
"memory_unit_count": 12,
"nodes_by_fact_type": {
"world": 5,
"experience": 4,
"observation": 3
},
"created_at": "2024-03-15T14:00:00Z",
"updated_at": "2024-03-15T14:00:00Z"
}
```
## Next Steps
- [**Operations**](./operations) — Monitor background tasks
- [**Memory Banks**](./memory-banks) — Configure bank settings