VVDex Forge AI exams. Evidence. Independent grading.

Featured exam · annotation · revision 0.1.0

Record normalization and duplicate QA

vvdex.annotation.structured-data-1 · 1 model lanes · 10 rollouts per lane · campaign fc-ec21239c00aa

All 10 sealed evaluation records contribute to the outcome counts. The trace, runtime details and single record digest below describe one sample record, selected in record order, without selecting for outcome. The campaign summary binds every record digest and the complete task results.

What is being tested

An exam from the annotation family.

Why this is difficult

See the task statement below for the behaviour under test.

What the model receives

A written task statement plus the workspace it names. Recorded limits: 24 tool steps; tools list_files, read_file, write_file, edit_file, run_tests, submit. CLI temperature and stochastic seed controls are not established by this record; attempt indices are bookkeeping. The reference solution and the hidden tests are never mounted in the model's workspace.

Task statement withheld — proprietary evaluation material.

The model received a written task statement and the workspace it names. Both are VVDex evaluation material: the statement is the exam, and on the retrieval exam it carries the question set, so publishing it would publish a reusable benchmark. What the exam asks for, and how it is graded, are described in plain English above; every count on this page is unaffected by the withholding.

How success is measured

Hidden tests are executed after submission in an isolated container.

Grading is one step of five. A certified pass additionally requires: verified containment, complete provenance (exam identity and revision fingerprint on the record), no integrity invariant fired (graded tree equals the submitted workspace), and no verdict-bearing harness suspicion. A grader pass alone is never presented as a result.

How Forge protects evaluation integrity

Every lane in this run was verified contained: wire lanes by construction (the provider sees only what the harness sends), CLI lanes by auditing the CLI's own session store for tool use, answer-key access, cross-session memory and network use after the run.

This run's self-check (API-failure share, identical-failure-across-providers, starvation, containment) fired no rule.

The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.

Every rollout carries three separate outcomes (grader, integrity, certified), and the record is sealed with a digest over its canonical JSON. Changing any verdict-bearing field breaks verification.

What happened
LaneGraderIntegrityCertified outcomeWall clock
cli/codex-gpt-5.5certified 10 of 10 graded rollouts, 100% (95% CI 72–100%)10 rollouts
Why a model failed

No additional failure details are published. The outcome table above is authoritative.

What each lane did

All 10 sealed evaluation records contribute to the outcome counts. The trace, runtime details and single record digest below describe one sample record, selected in record order, without selecting for outcome. The campaign summary binds every record digest and the complete task results.

The action sequence the harness recorded for each lane: tool and path only. Diff contents are not shown here on any exam. Public-source ownership may permit upstream source or task material and approved public-source diffs. It never declassifies VVDex grader, hidden-test, reference, answer-key or other evaluation internals.

cli:cli/codex-gpt-5.5 · 1 rollout(s) · 8 steps · most common shape: list_files → read_file ×4 → write_file → run_tests → submit
  • attempt index 0 · 8 steps · submitted_passed · list_files . → read_file TASK.md → read_file annotation-task.json → read_file assets/records.json → read_file submission.json → write_file submission.json → run_tests → submit
Ownership and what this page withholds

VVDex proprietary evaluation material. Public results are published for viewing, evaluation and verification; publication grants no licence to the exam, its task package, its corpus or world, its hidden tests or its grader. Terms & ownership.

Submission content withheld — proprietary evaluation material. On this exam the model's submitted content IS the graded answer, so nothing derived from a submission is published anywhere — not on this page, not in the campaign report, and not in the published JSON: no diff, no answer value or answer JSON body, no reference or gold value, no expected hidden output, none of the model's own words, and not the names of the hidden assertions a rollout failed, because a grader's name can carry the value it asserts. What is published is the tool-and-path walk above and the grader's verdicts; every count on this page is the grader's and is unaffected.

Rule applied to this exam: answer-key disclosure (ownership vvdex_proprietary — no declaration — proprietary by default; family annotation): the exam is VVDex evaluation IP and the graded output is an answer or a keyed value, so no submission content, diff, model claim or hidden-assertion name is published.

Can this result be verified

campaign fc-ec21239c00aa
evaluation record fr-20260906-0524b448
record digest 6c7ef1ad19e86c265db5ca3ae26545f32a3aa7610be592847156999eaad0a613
exam fingerprint 5d6b3708dba2cdf08218a44997058413d0a3aae738299bce8c346bb2b309fc28

The exam's own promotion chain — reference solution passes, baseline fails, grader controls distinguished, attack probes blocked, the runtime and isolation policy it ran under — is published as a machine-readable certification receipt. It states the same exam fingerprint as this page (5d6b3708dba2cdf08218a44997058413d0a3aae738299bce8c346bb2b309fc28); that fingerprint is the join key, and a receipt stating a different one describes a different exam revision. Fields the recorded evidence does not substantiate read unavailable rather than a number.

This exam's chapter, and the campaign's per-exam matrix that places this result beside the other exams in the same sitting, are in the campaign report.

The record's canonical bytes are not published beside this page: the sealed record carries content the disclosure policy withholds (transcripts[].diff.excerpt, rollouts[].lastWords), and the canonical bytes are the sealed record verbatim; the record digest is published instead. The digest above is the digest of the sealed record as it exists in the Forge; a holder of that record — the customer who commissioned the run, or an auditor under agreement — recomputes it with shasum -a 256 on the canonical bytes, or with vvdex-env records verify, and reaches the value printed here. Exams whose graded artefact is public source do publish their bytes; the browser verifier works on those.

The engine's own check on a full record: vvdex-env records verify --report <evalId>.md. The campaign document is regenerated from records only: vvdex-env records campaign --workspace <ws>; a document that disagrees with its records fails generation.

Limitations
  • Each lane sat this exam 10 times, on independent seeds. Rates carry 95% Wilson intervals. Where two lanes' intervals overlap, this page does not rank them; where they do not, it says so in the results above.
  • Lane errors (provider quota, rate limits) are recorded as lane facts and excluded from capability in both directions.
  • Withheld cells are neither passes nor model failures.
  • The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
What the outcomes mean
  • CERTIFIED PASS: the grader accepted the submission, containment was verified, provenance is complete and no integrity invariant fired.
  • model fail: the submission failed the hidden tests under verified containment; a real, attributable model result.
  • lane error: the provider or CLI lane failed (rate limit, quota, transport) before an attributable measurement existed; not a model result.
  • withheld: the run cannot prove its own conditions (containment unproven, provenance incomplete, or a harness suspicion the record cannot clear); neither a pass nor a model failure.
  • INVALID: an integrity invariant fired (isolation breach, evidence mismatch); the row is not a model result.
  • harness suspect: the run's own self-check fired; each rule is classified as verdict-bearing or diagnostic, and verdict-bearing suspicion withholds the affected cells.