← engrava.ai

Reproducible benchmark results.

This page renders results from github.com/sovantica/engrava-benchmark, a public, MIT-licensed reproduction of standard agent-memory benchmarks. Numbers here are read from that repository at a pinned commit at build time — there is no live API and no client-side fetch on this page.

Scores are only comparable to each other within the same segment below — the same benchmark, dataset revision, split, harness, reader, judge, and scorer. Different segments are never merged into one ranked table, because changing any of those axes changes what the number means.

No language model runs inside the memory pipeline itself — retrieval is embeddings only, and the reader and judge that turn retrieved context into a graded answer sit outside it. Every verified row was produced by the public engrava package under the official scorer, and carries its own audit trail: hypotheses, judge labels and a checksummed artifact, committed in the repository alongside the row. That is what lets a third party check a score without re-running it.

Engrava 0.5.0 and 0.6.0 sit in separate segments because they ran on different harness commits, the only axis separating them. 0.6.0 is the lower of the two, by four graded questions out of five hundred. Judge verdicts carry non-zero error, and the benchmark’s methodology treats small gaps as noise even inside a single segment; these two rows are not in one. Neither an improvement nor a regression is demonstrated.

longmemeval v1 · s_full_500

Dataset revision:
longmemeval_s_cleaned@sha256:d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442
Harness:
longmemeval-official @ engrava-benchmark@0d1fde5
Reader:
gpt-4o-2024-08-06 @ api.openai.com
Judge:
gpt-4o-2024-08-06 @ api.openai.com
Scorer:
longmemeval@9e0b455f4ef0e2ab8f2e582289761153549043fc
Verified results for longmemeval v1 — s_full_500, harness longmemeval-official @ engrava-benchmark@0d1fde5. Reader gpt-4o-2024-08-06 at api.openai.com, judge gpt-4o-2024-08-06 at api.openai.com. Micro and macro are a display axis, not a ranking.
SystemTierProvenanceMicroMacronReproduce
Engrava 0.5.0EngravaSovantica-run82.40%micro82.58%macro500ArtifactRow JSON

Secondary breakdown (per result, not ranked)

Engrava 0.5.0 (Engrava, Sovantica-run) — abstention accuracy 73.33% (n=30)

Single-session (user)Single-session (assistant)Single-session (preference)Knowledge updateTemporal reasoningMulti-session
100.00%n=7089.29%n=5666.67%n=3084.62%n=7883.46%n=13371.43%n=133

longmemeval v1 · s_full_500

Dataset revision:
longmemeval_s_cleaned@sha256:d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442
Harness:
longmemeval-official @ engrava-benchmark@a45dde9
Reader:
gpt-4o-2024-08-06 @ api.openai.com
Judge:
gpt-4o-2024-08-06 @ api.openai.com
Scorer:
longmemeval@9e0b455f4ef0e2ab8f2e582289761153549043fc
Verified results for longmemeval v1 — s_full_500, harness longmemeval-official @ engrava-benchmark@a45dde9. Reader gpt-4o-2024-08-06 at api.openai.com, judge gpt-4o-2024-08-06 at api.openai.com. Micro and macro are a display axis, not a ranking.
SystemTierProvenanceMicroMacronReproduce
Engrava 0.6.0EngravaSovantica-run81.60%micro81.76%macro500ArtifactRow JSON

Secondary breakdown (per result, not ranked)

Engrava 0.6.0 (Engrava, Sovantica-run) — abstention accuracy 76.67% (n=30)

Single-session (user)Single-session (assistant)Single-session (preference)Knowledge updateTemporal reasoningMulti-session
97.14%n=7085.71%n=5670.00%n=3082.05%n=7881.95%n=13373.68%n=133