Files
rohitg00__agentmemory/docs/benchmarks/TEMPLATE.md
Rohit Ghumare 7fb72f4010 feat(eval): pluggable benchmark harness with in-house coding-agent corpus (#562)
* feat(eval): pluggable benchmark harness with in-house coding-agent corpus

Adds eval/ tree (outside files field so npm tarball stays thin) with Adapter
interface, three reference adapters (grep / vector / agentmemory-hybrid),
two benchmarks (LongMemEval _s public, coding-agent-life-v1 in-house 15
sessions), scoring (P@K, R@K, hit, top-gold-rank), NDJSON output,
sandbox script.

coding-agent-life-v1 published scorecard at
docs/benchmarks/2026-05-20-coding-agent-life-v1.md:
agentmemory-hybrid R@5=0.967 P@5=0.578 (100% hit) vs grep R@5=0.967 P@5=0.267.
2.2x better precision on identical input, sandbox-reproducible.

Adapter contract: init(sessions, config) -> State; query(q, state, k) -> RankedDoc[]

npm scripts:
  npm run eval:coding-life   (no download, no API key for grep)
  npm run eval:longmemeval   (needs OPENAI key + 278MB download)

eval/scripts/sandbox.sh boots clean agentmemory + iii-engine on ports
3411/3412 with isolated data dir; tears down on exit.

README headline updated. 1072/1072 tests pass + 5 new eval tests.

* fix(eval): address review findings on benchmark harness

- agentmemory adapter: prefer row.sessionId before observationToSession lookup
- vector adapter: validate embedBatch response (length, indexes, non-empty rows)
- coding-life: positive-int guard on --k; wrap query loop in try/finally so teardown runs
- longmemeval: positive-int guards on --k/--limit/--stratify; per-question try/finally
- load: throw on haystack_session_ids vs haystack_sessions length mismatch
- score: P@K denominator is k (requested cutoff) not topK.length
- sandbox.sh: guard rm -rf with non-empty + /tmp/ prefix check
- README: drop unsafe rm "$(which iii)"; instruct ~/.local/bin + PATH instead; add language tag to repo-layout fenced block
- sessions.json: fix "two-phase" -> "three-phase" wording mismatch
2026-05-20 14:11:52 +01:00

1.4 KiB

—

Commit: <sha> Bench: LongMemEval _s / coding-agent-life-v1 / ... N: 500 / 15 / ... K: 5 Hardware: macos-15 / ubuntu-22.04 / ... OpenAI model: text-embedding-3-small Anthropic model: N/A (no LLM in retrieval loop)

Headline

agentmemory-hybrid: R@5 = XX.XX%, P@5 = XX.XX%, p50 latency = XXms

Beats grep baseline by +X.Xpt R@5, vector by +X.Xpt R@5.

Per-adapter

Adapter P@5 R@5 Hit rate p50 latency
grep
vector
agentmemory-hybrid

Per-question-type

Type grep R@5 vector R@5 agentmemory R@5
single-session-bug
single-session-refactor
preference
multi-session-causal
temporal

Methodology

  • Sessions ingested via POST /agentmemory/remember with type=eval-session
  • Queries hit POST /agentmemory/smart-search with limit=k*4
  • No LLM in retrieval loop. Direct rank from hybrid scoring.
  • Ranks dedup by sessionId before truncating to K
  • Latency measured as init+query for LongMemEval (per-question fresh state), query-only for coding-life (shared state)

Reproduce

git checkout <sha>
npm install --legacy-peer-deps
OPENAI_API_KEY=sk-... AGENTMEMORY_BASE_URL=http://localhost:3111 \
  npm run eval:longmemeval -- --stratify 10

Notes

<what surprised, what regressed, what's load-bearing>