Featured exam · long_term_memory · revision 0.1.0
Cross-session fact update
Updating a cross-session fact: answer from the newer authoritative roster, not stale carried memory
vvdex.knowledge.memory-fact-update-1 · 3 model lanes · 10 rollouts per lane · campaign fc-d89e429d2781
A state-and-memory task across sessions: a fact was recorded earlier, an authoritative source has since changed it, and the model must answer from the newer authoritative roster rather than the stale carried memory, and acknowledge the change.
The failure mode this exposes is stale memory winning over fresh authority: repeating what was remembered instead of what the current source says, or updating the answer without recording that it changed.
A written task statement plus the workspace it names. Recorded limits: 20 tool steps; tools list_files, read_file, write_file, edit_file, run_tests, submit, run_command. CLI temperature and stochastic seed controls are not established by this record; attempt indices are bookkeeping. The reference solution and the hidden tests are never mounted in the model's workspace.
Task statement withheld — proprietary evaluation material.
The model received a written task statement and the workspace it names. Both are VVDex evaluation material: the statement is the exam, and on the retrieval exam it carries the question set, so publishing it would publish a reusable benchmark. What the exam asks for, and how it is graded, are described in plain English above; every count on this page is unaffected by the withholding.
Hidden tests check that the answer reflects the authoritative source and that the change is acknowledged in the model's written output; both are required.
Grading is one step of five. A certified pass additionally requires: verified containment, complete provenance (exam identity and revision fingerprint on the record), no integrity invariant fired (graded tree equals the submitted workspace), and no verdict-bearing harness suspicion. A grader pass alone is never presented as a result.
Every lane in this run was verified contained: wire lanes by construction (the provider sees only what the harness sends), CLI lanes by auditing the CLI's own session store for tool use, answer-key access, cross-session memory and network use after the run.
This run's self-check (API-failure share, identical-failure-across-providers, starvation, containment) fired no rule.
The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
Every rollout carries three separate outcomes (grader, integrity, certified), and the record is sealed with a digest over its canonical JSON. Changing any verdict-bearing field breaks verification.
| Lane | Grader | Integrity | Certified outcome | Wall clock |
|---|---|---|---|---|
| cli/claude-sonnet | certified 10 of 10 graded rollouts, 100% (95% CI 72–100%) | 10 rollouts | ||
| cli/codex | certified 10 of 10 graded rollouts, 100% (95% CI 72–100%) | 10 rollouts | ||
| codestral-latest | certified 2 of 10 graded rollouts, 20% (95% CI 6–51%) | 10 rollouts | ||
- codestral-latest: 8 rollout(s) submitted an answer the hidden tests rejected; the names of the failing hidden assertions are withheld: they are the grader's own, and the grader stays private whoever owns the source under test
The action sequence the harness recorded for each lane: tool and path only. Diff contents are not shown here on any exam. Public-source ownership may permit upstream source or task material and approved public-source diffs. It never declassifies VVDex grader, hidden-test, reference, answer-key or other evaluation internals.
cli:cli/claude-sonnet · 10 rollout(s) · 5 steps · most common shape: read_file ×2 → write_file → run_tests → submit
- attempt index 0 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 1 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 2 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 3 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 4 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 5 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 6 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 7 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 8 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 9 · 5 steps · submitted_passed · read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
cli:cli/codex · 10 rollout(s) · 6–7 steps · most common shape: read_file ×3 → write_file → run_tests → submit
- attempt index 0 · 6 steps · submitted_passed · read_file ISSUE.md → read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 1 · 6 steps · submitted_passed · read_file ISSUE.md → read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 2 · 7 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 3 · 7 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 4 · 6 steps · submitted_passed · read_file ISSUE.md → read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 5 · 6 steps · submitted_passed · read_file ISSUE.md → read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 6 · 6 steps · submitted_passed · read_file ISSUE.md → read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 7 · 6 steps · submitted_passed · read_file ISSUE.md → read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 8 · 6 steps · submitted_passed · read_file ISSUE.md → read_file memory/notes.md → read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 9 · 6 steps · submitted_passed · read_file ISSUE.md → read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
mistral:codestral-latest · 10 rollout(s) · 4–5 steps · most common shape: read_file → write_file → run_tests → submit
- attempt index 0 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 1 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 2 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 3 · 5 steps · submitted_passed · read_file facts/current.json → read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 4 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 5 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 6 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
- attempt index 7 · 4 steps · submitted_failed · read_file facts/current.json → write_file answers/answer.json → run_tests → submit
- attempt index 8 · 5 steps · submitted_passed · read_file facts/current.json → read_file memory/notes.md → run_tests → write_file answers/answer.json → submit
- attempt index 9 · 4 steps · submitted_failed · read_file memory/notes.md → write_file answers/answer.json → run_tests → submit
VVDex proprietary evaluation material. Public results are published for viewing, evaluation and verification; publication grants no licence to the exam, its task package, its corpus or world, its hidden tests or its grader. Terms & ownership.
Submission content withheld — proprietary evaluation material. On this exam the model's submitted content IS the graded answer, so nothing derived from a submission is published anywhere — not on this page, not in the campaign report, and not in the published JSON: no diff, no answer value or answer JSON body, no reference or gold value, no expected hidden output, none of the model's own words, and not the names of the hidden assertions a rollout failed, because a grader's name can carry the value it asserts. What is published is the tool-and-path walk above and the grader's verdicts; every count on this page is the grader's and is unaffected.
Rule applied to this exam: answer-key disclosure (ownership vvdex_proprietary — engine registry; family long_term_memory): the exam is VVDex evaluation IP and the graded output is an answer or a keyed value, so no submission content, diff, model claim or hidden-assertion name is published.
Retired for future evaluation on 2026-09-02: post-run public disclosure of answer-bearing submission content. Historical sealed results on this revision are unaffected: the disclosure happened after those runs, so it could not have influenced them. This exam revision will not be used to measure a model again; it stays published so the results already recorded on it stay checkable.
campaign fc-d89e429d2781
evaluation record fr-20260901-46eb9548
record digest 1127e8356d42d31767a3ddc4d6c6cd0a515f5bbf7239ef0a82129e00fc2e6900
exam fingerprint 2052eec0ff2d3044dc269e3599ae1bfaa45d0d8cd5df9554796105538d226c21
The exam's own promotion chain — reference solution passes, baseline fails, grader controls distinguished, attack probes blocked, the runtime and isolation policy it ran under — is published as a machine-readable certification receipt. It states the same exam fingerprint as this page (2052eec0ff2d3044dc269e3599ae1bfaa45d0d8cd5df9554796105538d226c21); that fingerprint is the join key, and a receipt stating a different one describes a different exam revision. Fields the recorded evidence does not substantiate read unavailable rather than a number.
This exam's chapter, and the campaign's per-exam matrix that places this result beside the other exams in the same sitting, are in the campaign report.
The record's canonical bytes are not published beside this page: the sealed record carries content the disclosure policy withholds (rollouts[].failedTests, transcripts[].diff.excerpt, rollouts[].lastWords), and the canonical bytes are the sealed record verbatim; the record digest is published instead. The digest above is the digest of the sealed record as it exists in the Forge; a holder of that record — the customer who commissioned the run, or an auditor under agreement — recomputes it with shasum -a 256 on the canonical bytes, or with vvdex-env records verify, and reaches the value printed here. Exams whose graded artefact is public source do publish their bytes; the browser verifier works on those.
The engine's own check on a full record: vvdex-env records verify --report <evalId>.md. The campaign document is regenerated from records only: vvdex-env records campaign --workspace <ws>; a document that disagrees with its records fails generation.
- Each lane sat this exam 10 times, on independent seeds. Rates carry 95% Wilson intervals. Where two lanes' intervals overlap, this page does not rank them; where they do not, it says so in the results above.
- Lane errors (provider quota, rate limits) are recorded as lane facts and excluded from capability in both directions.
- Withheld cells are neither passes nor model failures.
- The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
- CERTIFIED PASS: the grader accepted the submission, containment was verified, provenance is complete and no integrity invariant fired.
- model fail: the submission failed the hidden tests under verified containment; a real, attributable model result.
- lane error: the provider or CLI lane failed (rate limit, quota, transport) before an attributable measurement existed; not a model result.
- withheld: the run cannot prove its own conditions (containment unproven, provenance incomplete, or a harness suspicion the record cannot clear); neither a pass nor a model failure.
- INVALID: an integrity invariant fired (isolation breach, evidence mismatch); the row is not a model result.
- harness suspect: the run's own self-check fired; each rule is classified as verdict-bearing or diagnostic, and verdict-bearing suspicion withholds the affected cells.