Featured exam · annotation · revision 0.2.0
Robot catch temporal and tracking annotation
vvdex.annotation.robot-video-1 · 1 model lanes · one rollout per lane · campaign fc-403cfc23f9e5
An exam from the annotation family.
See the task statement below for the behaviour under test.
A written task statement plus the workspace it names. Recorded limits: 24 tool steps; tools list_files, read_file, write_file, edit_file, run_tests, submit. Recorded temperature: 0.2. The reference solution and the hidden tests are never mounted in the model's workspace.
Task statement withheld — proprietary evaluation material.
The model received a written task statement and the workspace it names. Both are VVDex evaluation material: the statement is the exam, and on the retrieval exam it carries the question set, so publishing it would publish a reusable benchmark. What the exam asks for, and how it is graded, are described in plain English above; every count on this page is unaffected by the withholding.
Hidden tests are executed after submission in an isolated container.
Grading is one step of five. A certified pass additionally requires: verified containment, complete provenance (exam identity and revision fingerprint on the record), no integrity invariant fired (graded tree equals the submitted workspace), and no verdict-bearing harness suspicion. A grader pass alone is never presented as a result.
Every lane in this run was verified contained: wire lanes by construction (the provider sees only what the harness sends), CLI lanes by auditing the CLI's own session store for tool use, answer-key access, cross-session memory and network use after the run.
This run's self-check (API-failure share, identical-failure-across-providers, starvation, containment) fired no rule.
The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
Every rollout carries three separate outcomes (grader, integrity, certified), and the record is sealed with a digest over its canonical JSON. Changing any verdict-bearing field breaks verification.
| Lane | Grader | Integrity | Certified outcome | Wall clock |
|---|---|---|---|---|
| gemini-3.6-flash | fail | verified | model fail | 125s |
- gemini-3.6-flash: The submission failed a hidden grader check.
The action sequence the harness recorded for each lane: tool and path only. Diff contents are not shown here on any exam. Public-source ownership may permit upstream source or task material and approved public-source diffs. It never declassifies VVDex grader, hidden-test, reference, answer-key or other evaluation internals.
gemini:gemini-3.6-flash · 1 rollout(s) · 7 steps · most common shape: list_files → read_file ×3 → write_file → run_tests → submit
- seed 0 · 7 steps · submitted_failed · list_files . → read_file TASK.md → read_file annotation-task.json → read_file submission.json → write_file submission.json → run_tests → submit
VVDex proprietary evaluation material. Public results are published for viewing, evaluation and verification; publication grants no licence to the exam, its task package, its corpus or world, its hidden tests or its grader. Terms & ownership.
Submission content withheld — proprietary evaluation material. On this exam the model's submitted content IS the graded answer, so nothing derived from a submission is published anywhere — not on this page, not in the campaign report, and not in the published JSON: no diff, no answer value or answer JSON body, no reference or gold value, no expected hidden output, none of the model's own words, and not the names of the hidden assertions a rollout failed, because a grader's name can carry the value it asserts. What is published is the tool-and-path walk above and the grader's verdicts; every count on this page is the grader's and is unaffected.
Rule applied to this exam: answer-key disclosure (ownership vvdex_proprietary — no declaration — proprietary by default; family annotation): the exam is VVDex evaluation IP and the graded output is an answer or a keyed value, so no submission content, diff, model claim or hidden-assertion name is published.
campaign fc-403cfc23f9e5
evaluation record fr-20260906-92669fc1
record digest f9e6baef1c1acb4e45e2bd7e027b368b5e594f33037c643dc8ee9edb75daa2e9
exam fingerprint 85e026007af17db692614084f12ff40c5bbf2348bcfda0a01deb95aad8928f4c
The exam's own promotion chain — reference solution passes, baseline fails, grader controls distinguished, attack probes blocked, the runtime and isolation policy it ran under — is published as a machine-readable certification receipt. It states the same exam fingerprint as this page (85e026007af17db692614084f12ff40c5bbf2348bcfda0a01deb95aad8928f4c); that fingerprint is the join key, and a receipt stating a different one describes a different exam revision. Fields the recorded evidence does not substantiate read unavailable rather than a number.
This exam's chapter, and the campaign's per-exam matrix that places this result beside the other exams in the same sitting, are in the campaign report.
The record's canonical bytes are not published beside this page: the sealed record carries content the disclosure policy withholds (transcripts[].diff.excerpt, rollouts[].lastWords), and the canonical bytes are the sealed record verbatim; the record digest is published instead. The digest above is the digest of the sealed record as it exists in the Forge; a holder of that record — the customer who commissioned the run, or an auditor under agreement — recomputes it with shasum -a 256 on the canonical bytes, or with vvdex-env records verify, and reaches the value printed here. Exams whose graded artefact is public source do publish their bytes; the browser verifier works on those.
The engine's own check on a full record: vvdex-env records verify --report <evalId>.md. The campaign document is regenerated from records only: vvdex-env records campaign --workspace <ws>; a document that disagrees with its records fails generation.
- Each lane sat this exam once. These are individual certified trials; no rate, ranking or "better than" claim is supported by one rollout per lane.
- Lane errors (provider quota, rate limits) are recorded as lane facts and excluded from capability in both directions.
- Withheld cells are neither passes nor model failures.
- The certification chain for this exam revision is recorded in this package: reference solution passes at rate 1.0, baseline fails at rate 0.0.
- CERTIFIED PASS: the grader accepted the submission, containment was verified, provenance is complete and no integrity invariant fired.
- model fail: the submission failed the hidden tests under verified containment; a real, attributable model result.
- lane error: the provider or CLI lane failed (rate limit, quota, transport) before an attributable measurement existed; not a model result.
- withheld: the run cannot prove its own conditions (containment unproven, provenance incomplete, or a harness suspicion the record cannot clear); neither a pass nor a model failure.
- INVALID: an integrity invariant fired (isolation breach, evidence mismatch); the row is not a model result.
- harness suspect: the run's own self-check fired; each rule is classified as verdict-bearing or diagnostic, and verdict-bearing suspicion withholds the affected cells.