mirror of
https://github.com/rohitg00/agentmemory.git
synced 2026-09-14 20:16:33 +08:00
7fb72f4010
* feat(eval): pluggable benchmark harness with in-house coding-agent corpus Adds eval/ tree (outside files field so npm tarball stays thin) with Adapter interface, three reference adapters (grep / vector / agentmemory-hybrid), two benchmarks (LongMemEval _s public, coding-agent-life-v1 in-house 15 sessions), scoring (P@K, R@K, hit, top-gold-rank), NDJSON output, sandbox script. coding-agent-life-v1 published scorecard at docs/benchmarks/2026-05-20-coding-agent-life-v1.md: agentmemory-hybrid R@5=0.967 P@5=0.578 (100% hit) vs grep R@5=0.967 P@5=0.267. 2.2x better precision on identical input, sandbox-reproducible. Adapter contract: init(sessions, config) -> State; query(q, state, k) -> RankedDoc[] npm scripts: npm run eval:coding-life (no download, no API key for grep) npm run eval:longmemeval (needs OPENAI key + 278MB download) eval/scripts/sandbox.sh boots clean agentmemory + iii-engine on ports 3411/3412 with isolated data dir; tears down on exit. README headline updated. 1072/1072 tests pass + 5 new eval tests. * fix(eval): address review findings on benchmark harness - agentmemory adapter: prefer row.sessionId before observationToSession lookup - vector adapter: validate embedBatch response (length, indexes, non-empty rows) - coding-life: positive-int guard on --k; wrap query loop in try/finally so teardown runs - longmemeval: positive-int guards on --k/--limit/--stratify; per-question try/finally - load: throw on haystack_session_ids vs haystack_sessions length mismatch - score: P@K denominator is k (requested cutoff) not topK.length - sandbox.sh: guard rm -rf with non-empty + /tmp/ prefix check - README: drop unsafe rm "$(which iii)"; instruct ~/.local/bin + PATH instead; add language tag to repo-layout fenced block - sessions.json: fix "two-phase" -> "three-phase" wording mismatch
1.4 KiB
1.4 KiB
—
Commit: <sha>
Bench: LongMemEval _s / coding-agent-life-v1 / ...
N: 500 / 15 / ...
K: 5
Hardware: macos-15 / ubuntu-22.04 / ...
OpenAI model: text-embedding-3-small
Anthropic model: N/A (no LLM in retrieval loop)
Headline
agentmemory-hybrid: R@5 = XX.XX%, P@5 = XX.XX%, p50 latency = XXms
Beats grep baseline by +X.Xpt R@5, vector by +X.Xpt R@5.
Per-adapter
| Adapter | P@5 | R@5 | Hit rate | p50 latency |
|---|---|---|---|---|
| grep | ||||
| vector | ||||
| agentmemory-hybrid |
Per-question-type
| Type | grep R@5 | vector R@5 | agentmemory R@5 |
|---|---|---|---|
| single-session-bug | |||
| single-session-refactor | |||
| preference | |||
| multi-session-causal | |||
| temporal |
Methodology
- Sessions ingested via
POST /agentmemory/rememberwithtype=eval-session - Queries hit
POST /agentmemory/smart-searchwithlimit=k*4 - No LLM in retrieval loop. Direct rank from hybrid scoring.
- Ranks dedup by sessionId before truncating to K
- Latency measured as init+query for LongMemEval (per-question fresh state), query-only for coding-life (shared state)
Reproduce
git checkout <sha>
npm install --legacy-peer-deps
OPENAI_API_KEY=sk-... AGENTMEMORY_BASE_URL=http://localhost:3111 \
npm run eval:longmemeval -- --stratify 10
Notes
<what surprised, what regressed, what's load-bearing>