Files
Nicolò Boschi 613a699e9f fix(consolidation): eliminate duplicate observations via interleave dedup recall (#1907)
Round-robin interleave fusion for consolidation dedup recall (guarantees the semantic-#1 'twin' a slot so the LLM updates instead of duplicating), unified 'reranking' strategy param (cross_encoder/rrf/interleave), case-sensitive exact-dup guard, obs-dedup tool + benchmark wired into the perf dashboard (English dataset). Near-dup observation rate 4% -> 0% on the English hermes transcript (1/10 and 1/4), coverage 89% -> 94%, no false merges.
2026-06-03 17:37:57 +02:00
..

Observation duplication benchmark

Measures how many duplicate observations consolidation produces — a quality signal for the consolidation pipeline (complements the perf-focused consolidation benchmark).

For each document under datasets/, it ingests the content into a fresh bank, runs consolidation, then reuses the observation-dedup tool (hindsight_dev.obs_dedup) to score:

  • exact duplicates — observations with identical (normalised) text in a scope
  • near duplicates — cosine-similarity clusters at thresholds (0.97, 0.92)

Headline metric: duplication rate = redundant observations / total. Lower is better.

Run

./scripts/benchmarks/run-obs.sh
# or
cd hindsight-dev && uv run python -m benchmarks.obs.obs_benchmark

Requires a real LLM (set HINDSIGHT_API_LLM_PROVIDER / _MODEL / _API_KEY). Results are written to benchmarks/results/obs_benchmark_<ts>.json.

Extending

Drop more *.txt transcripts into datasets/ — one file per scenario. Keep them synthetic and PII-free. Documents that restate the same durable facts across many turns are the ones that stress consolidation's dedup behaviour.