Round-robin interleave fusion for consolidation dedup recall (guarantees the semantic-#1 'twin' a slot so the LLM updates instead of duplicating), unified 'reranking' strategy param (cross_encoder/rrf/interleave), case-sensitive exact-dup guard, obs-dedup tool + benchmark wired into the perf dashboard (English dataset). Near-dup observation rate 4% -> 0% on the English hermes transcript (1/10 and 1/4), coverage 89% -> 94%, no false merges.
Observation duplication benchmark
Measures how many duplicate observations consolidation produces — a quality
signal for the consolidation pipeline (complements the perf-focused
consolidation benchmark).
For each document under datasets/, it ingests the content into a fresh bank,
runs consolidation, then reuses the observation-dedup tool
(hindsight_dev.obs_dedup) to score:
- exact duplicates — observations with identical (normalised) text in a scope
- near duplicates — cosine-similarity clusters at thresholds (0.97, 0.92)
Headline metric: duplication rate = redundant observations / total. Lower is better.
Run
./scripts/benchmarks/run-obs.sh
# or
cd hindsight-dev && uv run python -m benchmarks.obs.obs_benchmark
Requires a real LLM (set HINDSIGHT_API_LLM_PROVIDER / _MODEL / _API_KEY).
Results are written to benchmarks/results/obs_benchmark_<ts>.json.
Extending
Drop more *.txt transcripts into datasets/ — one file per scenario. Keep them
synthetic and PII-free. Documents that restate the same durable facts across many
turns are the ones that stress consolidation's dedup behaviour.