VVDex ForgeAI exams. Evidence. Independent grading.

Catalog · as of 2026-09-02

Explore the AI exam inventory.

Explore task families and their certification status. Published reports show results; runnable exam packages have a separate release decision.

This is the historical inventory snapshot dated 2026-09-02. The five new Annotation fixtures and their public evidence are reported separately; this release does not change the catalog's runnable-exam listing.

Counts on this page are read from /data/catalog.json, the same file the listing rule is applied to.

Certified exams
25
Each has passed the full chain.
Families with an exam
14
Of 14 declared.
Listed publicly
0
Runnable packages are not publicly listed.
Listed exams

Reports are public. Runnable packages are separate.

This snapshot contains 25 certified AI exams and no publicly listed runnable packages. You can still explore published model results and the newer annotation certification reports.

Explore evaluated AI exams →   Explore annotation →

Families
Task familyCertifiedListedStatus
Public open-source SWEA bug fixed upstream by a merged pull request, rebuilt as an exam: the parent tree is the starting point, the maintainers' fix is the reference solution, and a hidden test suite grades the rule the fix establishes rather than the diff.110active
Terminal operationsThe model gets a shell inside a network-none box and is graded on the end state of the workspace, never on the commands it chose to get there.20active
Algorithms, graded by propertyGraded by derandomised property tests rather than examples, so a solution that hard-codes the worked examples in the task text fails structurally.10active
Code qualityHidden tests are the floor. The reward also requires deterministic quality axes measured inside the grade box: lint, dependency surface, complexity delta, and execution of the changed lines.10active
Security and permission boundariesOut-of-scope secrets carrying unique honeytokens, an injected-instruction document, and an approval-gated action whose approval token is absent. Boundary respect is scored from the submitted workspace.10active
Tool use and MCPDeclared tools with a permission table (allowed / approval / forbidden), dispatched through a checked loop that never touches the real world, recorded to a ledger outside the model's writable paths.10active
Retrieval and groundingA corpus with authoritative sources, distractors and outdated sources. Citation correctness is graded as strictly as the answer, and saying the evidence is insufficient is itself a graded axis.10active
Cross-session memoryThe exam runs across sessions with state carried between them. A fact is deliberately changed in the authoritative source, and the grader checks whether the model updates from the newer evidence or lets stale carried memory win.10active
Structured document workStructured inputs and a rules statement in, a reconciled output file out. The grader checks the end state exactly, including the credit note, the overpayment, the rounding and the excluded currency.10active
Customer supportA scripted scenario, never a free-form model user, so the exam stays deterministic. The grader checks the action against policy and whether the model claimed an action it never performed.10active
Data and SQLThe model writes a query; the grader executes it against the real dataset in the grade box and compares the result set to ground truth computed at authoring time. Write-safety is part of the grade.10active
Internal VVDex SWEBuilt from our own codebase. Never listed publicly, by an engine rule and not by a decision taken per exam.10active
Harness evaluationThe model-plus-harness's behaviour under adversarial tool conditions: a declared, deterministic fault schedule applied inside the tool loop, graded from the recorded trace and the end state.10active
Browser and computer useA real web application driven in a real headless Chromium inside the network-none box; graded on the application's end state, its audit log and the engine's action ledger, all outside the writable paths.10active
How publication works

Certification and release are separate decisions.

An AI exam may pass its grading and integrity checks while its working package remains private. Public release requires a separate review and explicit approval for that exam.

Published evidence describes outcomes and provenance. Private reference answers, graders and customer materials stay out of the public catalog.

How grading works →   Terms and ownership →