mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
43f545e9b8
* fix(documents): skip the retag cascade when a tags PATCH changes nothing (#3912) `update_document` never compared the incoming tags to the ones the document already carries — the path was `if tags is not None:`. So a PATCH re-sending an identical tags array, which is exactly what an idempotent tag-normalisation sweep does on every run after its first, paid the full retag cascade for a write that changes nothing. That cascade is not cheap and is not meant to be: it deletes every observation built over the document's memories and then resets `consolidated_at` on each of those observations' OTHER sources. Both halves are required for correctness — consolidation scopes a memory by its tag set, so an observation formed under the old tags is no longer valid, and deleting it strands the co-sourced memories it carried unless they are requeued. The consequence is that the blast radius of one PATCH is the co-source degree of the affected observations, not the document's own memory count, and a sweep multiplies that by its document count. Read the current tags before overwriting them and compare as SETS — consolidation scopes by tag set, so a reordered array changes nothing it can observe. When the set is unchanged, skip the memory-unit retag, the observation deletion, the source requeue and the consolidation submit. A tag set that differs at all still runs the cascade unchanged; tags that could not be read (document absent, or a store record not carrying them) are never treated as unchanged, so a real retag is never silently skipped. Both the SQL and store-owned branches are covered. Also drops a redundant `get_document_record` round-trip on the store-owned path, which the new pre-read already fetched. Tests: a repeat PATCH leaves the observation and every co-source consolidated; a reordered array is not a change; a superset still invalidates; and a real retag stamps `memory_units.updated_at` while a no-op leaves it alone. * chore(docs-skill): regenerate for the documents.mdx re-consolidation note The bundled hindsight-docs skill is generated from hindsight-docs/docs/**; verify-generated-files caught references/developer/api/documents.md drifting from the callout edited in the previous commit. Claude-Session: https://claude.ai/code/session_018HDqrzHgqZqsGc7EDqoTEu
253 lines
8.2 KiB
Plaintext
253 lines
8.2 KiB
Plaintext
---
|
|
sidebar_position: 8
|
|
---
|
|
|
|
# Documents
|
|
|
|
Track and manage document sources in your memory bank. Documents provide traceability — knowing where memories came from.
|
|
|
|
import Tabs from '@theme/Tabs';
|
|
import TabItem from '@theme/TabItem';
|
|
import CodeSnippet from '@site/src/components/CodeSnippet';
|
|
|
|
{/* Import raw source files */}
|
|
import documentsPy from '!!raw-loader!@site/examples/api/documents.py';
|
|
import documentsMjs from '!!raw-loader!@site/examples/api/documents.mjs';
|
|
import documentsGo from '!!raw-loader!@site/examples/api/documents.go';
|
|
|
|
:::tip Prerequisites
|
|
Make sure you've completed the [Quick Start](./quickstart) and understand [how retain works](./retain).
|
|
:::
|
|
|
|
## What Are Documents?
|
|
|
|
Documents are containers for retained content. They help you:
|
|
|
|
- **Track sources** — Know which PDF, conversation, or file a memory came from
|
|
- **Update content** — Re-retain a document to update its facts
|
|
- **Delete in bulk** — Remove all memories from a document at once
|
|
- **Organize memories** — Group related facts by source
|
|
|
|
## Chunks
|
|
|
|
When you retain content, Hindsight splits it into chunks before extracting facts. These chunks are stored alongside the extracted memories, preserving the original text segments.
|
|
|
|
**Why chunks matter:**
|
|
- **Context preservation** — Chunks contain the raw text that generated facts, useful when you need the exact wording
|
|
- **Richer recall** — Including chunks in recall provides surrounding context for matched facts
|
|
|
|
:::tip Include Chunks in Recall
|
|
Use `include_chunks=True` in your recall calls to get the original text chunks alongside fact results. See [Recall](./recall) for details.
|
|
:::
|
|
|
|
## Retain with Document ID
|
|
|
|
Associate retained content with a document:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-retain" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-retain" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# Retain content with document ID
|
|
hindsight memory retain my-bank "Meeting notes content..." --doc-id notes-2024-03-15
|
|
|
|
# Batch retain from files
|
|
hindsight memory retain-files my-bank docs/
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={documentsGo} section="document-retain" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Update Documents
|
|
|
|
Re-retaining with the same document_id **replaces** the old content:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-update" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# Original
|
|
hindsight memory retain my-bank "Project deadline: March 31" --doc-id project-plan
|
|
|
|
# Update
|
|
hindsight memory retain my-bank "Project deadline: April 15 (extended)" --doc-id project-plan
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={documentsGo} section="document-update" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Get Document
|
|
|
|
Retrieve a document's original text and metadata. This is useful for expanding document context after a recall operation returns memories with document references.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-get" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-get" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
hindsight document get my-bank meeting-2024-03-15
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={documentsGo} section="document-get" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Update Document
|
|
|
|
Update mutable fields on an existing document without re-processing the content. Currently supports updating `tags`.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-update" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# Replace tags with new values
|
|
hindsight document update-tags my-bank meeting-2024-03-15 --tags team-a --tags team-b
|
|
|
|
# Remove all tags
|
|
hindsight document update-tags my-bank meeting-2024-03-15
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={documentsGo} section="document-update" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
:::info Observations are re-consolidated
|
|
When tags change, any consolidated observations derived from the document's memories are invalidated and queued for re-consolidation under the new tags. Co-source memories from other documents that shared those observations are also reset.
|
|
|
|
This is required for correctness rather than incidental: consolidation scopes a memory by its tag set, so an observation built under the old tags is no longer valid, and deleting it would strand every other memory that observation was consolidated from unless those are requeued too. The size of that requeue is the number of memories co-sourced with this document's — on a densely co-sourced bank it can be many times the document's own memory count.
|
|
|
|
Tags are compared as a **set** against the document's current tags, and an update that leaves the set unchanged — including one that only reorders the array — performs no retag and queues no re-consolidation. A repeatable tag-normalisation sweep therefore only pays the re-consolidation cost on the run that actually changes something.
|
|
:::
|
|
|
|
## Delete Document
|
|
|
|
Remove a document and all its associated memories:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-delete" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-delete" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
hindsight document delete my-bank meeting-2024-03-15
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={documentsGo} section="document-delete" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
:::warning
|
|
Deleting a document permanently removes all memories extracted from it. This action cannot be undone.
|
|
:::
|
|
|
|
## List Documents
|
|
|
|
List documents in a bank with optional filtering by ID and tags.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-list" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-list" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# List all documents
|
|
hindsight document list my-bank
|
|
|
|
# Filter by ID substring
|
|
hindsight document list my-bank --q report
|
|
|
|
# Filter by tags
|
|
hindsight document list my-bank --tags team-a --tags team-b
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="go" label="Go">
|
|
<CodeSnippet code={documentsGo} section="document-list" language="go" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
### Filtering Options
|
|
|
|
| Parameter | Description |
|
|
|---|---|
|
|
| `q` | Case-insensitive substring match on document ID. `report` matches `report-2024`, `annual-report`, etc. |
|
|
| `tags` | Filter by document tags. Accepts multiple values. |
|
|
| `tags_match` | How to match tags (default: `any_strict`). See below. |
|
|
| `limit` / `offset` | Pagination. Default limit is 100. |
|
|
|
|
**`tags_match` modes:**
|
|
|
|
| Mode | Behaviour |
|
|
|---|---|
|
|
| `any_strict` *(default)* | Document must have **at least one** of the specified tags. Untagged docs excluded. |
|
|
| `any` | Same as `any_strict` but also includes untagged documents. |
|
|
| `all_strict` | Document must have **all** specified tags. Untagged docs excluded. |
|
|
| `all` | Same as `all_strict` but also includes untagged documents. |
|
|
|
|
## Document Response Format
|
|
|
|
```json
|
|
{
|
|
"id": "meeting-2024-03-15",
|
|
"bank_id": "my-bank",
|
|
"original_text": "Alice presented the Q4 roadmap...",
|
|
"content_hash": "abc123def456",
|
|
"memory_unit_count": 12,
|
|
"nodes_by_fact_type": {
|
|
"world": 5,
|
|
"experience": 4,
|
|
"observation": 3
|
|
},
|
|
"created_at": "2024-03-15T14:00:00Z",
|
|
"updated_at": "2024-03-15T14:00:00Z"
|
|
}
|
|
```
|
|
|
|
## Next Steps
|
|
|
|
- [**Operations**](./operations) — Monitor background tasks
|
|
- [**Memory Banks**](./memory-banks) — Configure bank settings
|