Featured exam · rag_grounding · revision 0.1.0
Grounded question answering
Answering questions from a document collection with outdated and distractor sources, and refusing when the evidence is insufficient
vvdex.knowledge.grounded-rag-1 · 3 model lanes · 10 rollouts per lane · campaign fc-d89e429d2781
Grounded question answering over a small document collection. Some documents are current, some are superseded versions, some are unrelated. The model must answer only from the corpus, cite exactly the supporting sources, and say when the corpus does not contain an answer.
The failure mode this exposes is confident fabrication and mis-citation: answering from a superseded document, citing a document that does not support the answer, or inventing an answer for a question the corpus cannot answer. Rule 4 of the exam calls the last one its worst failure.
A written task statement plus the workspace it names. Recorded limits: 30 tool steps; tools list_files, read_file, write_file, edit_file, run_tests, submit, run_command. CLI temperature and stochastic seed controls are not established by this record; attempt indices are bookkeeping. The reference solution and the hidden tests are never mounted in the model's workspace.
Task statement withheld — proprietary evaluation material.
The model received a written task statement and the workspace it names. Both are VVDex evaluation material: the statement is the exam, and on the retrieval exam it carries the question set, so publishing it would publish a reusable benchmark. What the exam asks for, and how it is graded, are described in plain English above; every count on this page is unaffected by the withholding.
Hidden tests compare each answer to the key (case-insensitive, whitespace and trailing period ignored), require the citation set to match exactly, and require `insufficient` to be set exactly when the corpus has no answer.
Grading is one step of five. A certified pass additionally requires: verified containment, complete provenance (exam identity and revision fingerprint on the record), no integrity invariant fired (graded tree equals the submitted workspace), and no verdict-bearing harness suspicion. A grader pass alone is never presented as a result.
Every lane in this run was verified contained: wire lanes by construction (the provider sees only what the harness sends), CLI lanes by auditing the CLI's own session store for tool use, answer-key access, cross-session memory and network use after the run.
This run's self-check (API-failure share, identical-failure-across-providers, starvation, containment) fired no rule.
The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
Every rollout carries three separate outcomes (grader, integrity, certified), and the record is sealed with a digest over its canonical JSON. Changing any verdict-bearing field breaks verification.
| Lane | Grader | Integrity | Certified outcome | Wall clock |
|---|---|---|---|---|
| cli/claude-sonnet | certified 1 of 10 graded rollouts, 10% (95% CI 2–40%) | 10 rollouts | ||
| cli/codex | certified 10 of 10 graded rollouts, 100% (95% CI 72–100%) | 10 rollouts | ||
| codestral-latest | certified 0 of 9 graded rollouts; outside the denominator: 1 lane error | 10 rollouts | ||
- cli/claude-sonnet: 9 rollout(s) submitted an answer the hidden tests rejected; the names of the failing hidden assertions are withheld: they are the grader's own, and the grader stays private whoever owns the source under test
- codestral-latest: 9 rollout(s) submitted an answer the hidden tests rejected; the names of the failing hidden assertions are withheld: they are the grader's own, and the grader stays private whoever owns the source under test
The action sequence the harness recorded for each lane: tool and path only. Diff contents are not shown here on any exam. Public-source ownership may permit upstream source or task material and approved public-source diffs. It never declassifies VVDex grader, hidden-test, reference, answer-key or other evaluation internals.
cli:cli/claude-sonnet · 10 rollout(s) · 6–12 steps · most common shape: list_files → read_file ×8 → write_file → run_tests → submit
- attempt index 0 · 6 steps · submitted_failed · list_files . → read_file questions.json → run_command → write_file answers/answers.json → run_tests → submit
- attempt index 1 · 12 steps · submitted_passed · list_files . → read_file questions.json → read_file corpus/data-protection-policy-v3.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/incident-report-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/employee-handbook.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 2 · 12 steps · submitted_failed · list_files . → read_file questions.json → read_file corpus/data-protection-policy-v3.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/incident-report-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/employee-handbook.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 3 · 8 steps · submitted_failed · list_files . → read_file questions.json → run_command → run_command → run_command → invalid_tool_call (none) → run_tests → submit
- attempt index 4 · 12 steps · submitted_failed · list_files . → read_file questions.json → read_file corpus/org-chart-2026.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/incident-report-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/employee-handbook.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 5 · 11 steps · submitted_failed · list_files . → read_file questions.json → read_file corpus/org-chart-2026.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/incident-report-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/employee-handbook.txt → write_file answers/answers.json → run_tests → submit
- attempt index 6 · 12 steps · submitted_failed · list_files . → read_file questions.json → read_file corpus/data-protection-policy-v3.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/incident-report-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/employee-handbook.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 7 · 6 steps · submitted_failed · list_files . → read_file questions.json → run_command → write_file answers/answers.json → run_tests → submit
- attempt index 8 · 8 steps · submitted_failed · list_files . → read_file questions.json → run_command → run_command → run_command → invalid_tool_call (none) → run_tests → submit
- attempt index 9 · 12 steps · submitted_failed · list_files . → read_file questions.json → read_file corpus/data-protection-policy-v3.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/incident-report-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/employee-handbook.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
cli:cli/codex · 10 rollout(s) · 14 steps · most common shape: read_file → list_files → read_file ×9 → write_file → run_tests → submit
- attempt index 0 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v3.txt → read_file corpus/data-protection-policy-v1.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 1 · 14 steps · submitted_passed · list_files . → read_file ISSUE.md → read_file questions.json → read_file corpus/data-protection-policy-v3.txt → read_file corpus/data-protection-policy-v1.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 2 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 3 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 4 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 5 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 6 · 14 steps · submitted_passed · list_files . → read_file ISSUE.md → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 7 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 8 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
- attempt index 9 · 14 steps · submitted_passed · read_file ISSUE.md → list_files . → read_file questions.json → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file corpus/incident-report-2026.txt → read_file corpus/org-chart-2026.txt → read_file corpus/retention-addendum-2026.txt → read_file corpus/security-whitepaper.txt → read_file corpus/vendor-list.txt → write_file answers/answers.json → run_tests → submit
mistral:codestral-latest · 10 rollout(s) · 4–13 steps · most common shape: read_file ×2 → list_files ×2 → read_file → write_file → run_tests → submit
- attempt index 0 · 10 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → list_files corpus → read_file corpus/data-protection-policy-v3.txt → read_file corpus/org-chart-2026.txt → run_tests → invalid_tool_call (none) → write_file answers/answers.json → run_tests → submit
- attempt index 1 · 8 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → list_files answers → list_files corpus → read_file corpus/data-protection-policy-v3.txt → write_file answers/answers.json → run_tests → submit
- attempt index 2 · 9 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → read_file → list_files corpus → read_file corpus/retention-addendum-2026.txt → read_file corpus/data-protection-policy-v3.txt → write_file answers/answers.json → run_tests → submit
- attempt index 3 · 13 steps · submitted_failed · read_file ISSUE.md → list_files answers → read_file answers/.keep → list_files corpus → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → read_file corpus/employee-handbook.txt → read_file questions.json → run_tests → invalid_tool_call (none) → write_file answers/answers.json → run_tests → submit
- attempt index 4 · 4 steps · model_api_failure · read_file ISSUE.md → read_file questions.json → list_files corpus → invalid_tool_call (none)
- attempt index 5 · 9 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → list_files answers → read_file answers/.keep → list_files corpus → read_file corpus/retention-addendum-2026.txt → write_file answers/answers.json → run_tests → submit
- attempt index 6 · 7 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → list_files corpus → read_file corpus/retention-addendum-2026.txt → write_file answers/answers.json → run_tests → submit
- attempt index 7 · 8 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → list_files answers → list_files corpus → read_file corpus/data-protection-policy-v3.txt → write_file answers/answers.json → run_tests → submit
- attempt index 8 · 9 steps · submitted_failed · read_file ISSUE.md → list_files answers → read_file answers/.keep → list_files corpus → read_file corpus/data-protection-policy-v1.txt → read_file corpus/data-protection-policy-v3.txt → write_file answers/answers.json → run_tests → submit
- attempt index 9 · 7 steps · submitted_failed · read_file ISSUE.md → read_file questions.json → list_files corpus → read_file corpus/data-protection-policy-v3.txt → write_file answers/answers.json → run_tests → submit
VVDex proprietary evaluation material. Public results are published for viewing, evaluation and verification; publication grants no licence to the exam, its task package, its corpus or world, its hidden tests or its grader. Terms & ownership.
Submission content withheld — proprietary evaluation material. On this exam the model's submitted content IS the graded answer, so nothing derived from a submission is published anywhere — not on this page, not in the campaign report, and not in the published JSON: no diff, no answer value or answer JSON body, no reference or gold value, no expected hidden output, none of the model's own words, and not the names of the hidden assertions a rollout failed, because a grader's name can carry the value it asserts. What is published is the tool-and-path walk above and the grader's verdicts; every count on this page is the grader's and is unaffected.
Rule applied to this exam: answer-key disclosure (ownership vvdex_proprietary — engine registry; family rag_grounding): the exam is VVDex evaluation IP and the graded output is an answer or a keyed value, so no submission content, diff, model claim or hidden-assertion name is published.
Retired for future evaluation on 2026-09-02: post-run public disclosure of answer-bearing submission content. Historical sealed results on this revision are unaffected: the disclosure happened after those runs, so it could not have influenced them. This exam revision will not be used to measure a model again; it stays published so the results already recorded on it stay checkable.
campaign fc-d89e429d2781
evaluation record fr-20260902-2e7f4171
record digest 157ec324a925c0b10fc8769d01f628b6143c40a00fa2dc70ea309a70e55a37a8
exam fingerprint 4880a527ef97ca2f3fe6b57be026c3dc1be51e9e104d79fc13a3e558331dda92
The exam's own promotion chain — reference solution passes, baseline fails, grader controls distinguished, attack probes blocked, the runtime and isolation policy it ran under — is published as a machine-readable certification receipt. It states the same exam fingerprint as this page (4880a527ef97ca2f3fe6b57be026c3dc1be51e9e104d79fc13a3e558331dda92); that fingerprint is the join key, and a receipt stating a different one describes a different exam revision. Fields the recorded evidence does not substantiate read unavailable rather than a number.
This exam's chapter, and the campaign's per-exam matrix that places this result beside the other exams in the same sitting, are in the campaign report.
The record's canonical bytes are not published beside this page: the sealed record carries content the disclosure policy withholds (rollouts[].failedTests, transcripts[].diff.excerpt, rollouts[].lastWords), and the canonical bytes are the sealed record verbatim; the record digest is published instead. The digest above is the digest of the sealed record as it exists in the Forge; a holder of that record — the customer who commissioned the run, or an auditor under agreement — recomputes it with shasum -a 256 on the canonical bytes, or with vvdex-env records verify, and reaches the value printed here. Exams whose graded artefact is public source do publish their bytes; the browser verifier works on those.
The engine's own check on a full record: vvdex-env records verify --report <evalId>.md. The campaign document is regenerated from records only: vvdex-env records campaign --workspace <ws>; a document that disagrees with its records fails generation.
- Each lane sat this exam 10 times, on independent seeds. Rates carry 95% Wilson intervals. Where two lanes' intervals overlap, this page does not rank them; where they do not, it says so in the results above.
- Lane errors (provider quota, rate limits) are recorded as lane facts and excluded from capability in both directions.
- Withheld cells are neither passes nor model failures.
- The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
- CERTIFIED PASS: the grader accepted the submission, containment was verified, provenance is complete and no integrity invariant fired.
- model fail: the submission failed the hidden tests under verified containment; a real, attributable model result.
- lane error: the provider or CLI lane failed (rate limit, quota, transport) before an attributable measurement existed; not a model result.
- withheld: the run cannot prove its own conditions (containment unproven, provenance incomplete, or a harness suspicion the record cannot clear); neither a pass nor a model failure.
- INVALID: an integrity invariant fired (isolation breach, evidence mismatch); the row is not a model result.
- harness suspect: the run's own self-check fired; each rule is classified as verdict-bearing or diagnostic, and verdict-bearing suspicion withholds the affected cells.