Reproducible benchmark results.
This page renders results from github.com/sovantica/engrava-benchmark, a public, MIT-licensed reproduction of standard agent-memory benchmarks. Numbers here are read from that repository at a pinned commit at build time — there is no live API and no client-side fetch on this page.
Scores are only comparable to each other within the same segment below — the same benchmark, dataset revision, split, harness, reader, judge, and scorer. Different segments are never merged into one ranked table, because changing any of those axes changes what the number means.
No language model runs inside the memory pipeline itself — retrieval is embeddings only, and the reader and judge that turn retrieved context into a graded answer sit outside it. Every verified row was produced by the public engrava package under the official scorer, and carries its own audit trail: hypotheses, judge labels and a checksummed artifact, committed in the repository alongside the row. That is what lets a third party check a score without re-running it.
Engrava 0.5.0 and 0.6.0 sit in separate segments because they ran on different harness commits, the only axis separating them. 0.6.0 is the lower of the two, by four graded questions out of five hundred. Judge verdicts carry non-zero error, and the benchmark’s methodology treats small gaps as noise even inside a single segment; these two rows are not in one. Neither an improvement nor a regression is demonstrated.
longmemeval v1 · s_full_500
- Dataset revision:
- longmemeval_s_cleaned@sha256:d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442
- Harness:
- longmemeval-official @ engrava-benchmark@0d1fde5
- Reader:
- gpt-4o-2024-08-06 @ api.openai.com
- Judge:
- gpt-4o-2024-08-06 @ api.openai.com
- Scorer:
- longmemeval@9e0b455f4ef0e2ab8f2e582289761153549043fc
| System | Tier | Provenance | Micro | Macro | n | Reproduce |
|---|---|---|---|---|---|---|
| Engrava 0.5.0 | Engrava | Sovantica-run | 82.40%micro | 82.58%macro | 500 | ArtifactRow JSON |
Secondary breakdown (per result, not ranked)
Engrava 0.5.0 (Engrava, Sovantica-run) — abstention accuracy 73.33% (n=30)
| Single-session (user) | Single-session (assistant) | Single-session (preference) | Knowledge update | Temporal reasoning | Multi-session |
|---|---|---|---|---|---|
| 100.00%n=70 | 89.29%n=56 | 66.67%n=30 | 84.62%n=78 | 83.46%n=133 | 71.43%n=133 |
longmemeval v1 · s_full_500
- Dataset revision:
- longmemeval_s_cleaned@sha256:d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442
- Harness:
- longmemeval-official @ engrava-benchmark@a45dde9
- Reader:
- gpt-4o-2024-08-06 @ api.openai.com
- Judge:
- gpt-4o-2024-08-06 @ api.openai.com
- Scorer:
- longmemeval@9e0b455f4ef0e2ab8f2e582289761153549043fc
| System | Tier | Provenance | Micro | Macro | n | Reproduce |
|---|---|---|---|---|---|---|
| Engrava 0.6.0 | Engrava | Sovantica-run | 81.60%micro | 81.76%macro | 500 | ArtifactRow JSON |
Secondary breakdown (per result, not ranked)
Engrava 0.6.0 (Engrava, Sovantica-run) — abstention accuracy 76.67% (n=30)
| Single-session (user) | Single-session (assistant) | Single-session (preference) | Knowledge update | Temporal reasoning | Multi-session |
|---|---|---|---|---|---|
| 97.14%n=70 | 85.71%n=56 | 70.00%n=30 | 82.05%n=78 | 81.95%n=133 | 73.68%n=133 |