Catalog · as of 2026-09-02
Explore the AI exam inventory.
Explore task families and their certification status. Published reports show results; runnable exam packages have a separate release decision.
Certified exams
25
Each has passed the full chain.
Families with an exam
14
Of 14 declared.
Listed publicly
0
Runnable packages are not publicly listed.
Listed exams
Reports are public. Runnable packages are separate.
This snapshot contains 25 certified AI exams and no publicly listed runnable packages. You can still explore published model results and the newer annotation certification reports.
Families
| Task family | Certified | Listed | Status |
|---|---|---|---|
| Public open-source SWEA bug fixed upstream by a merged pull request, rebuilt as an exam: the parent tree is the starting point, the maintainers' fix is the reference solution, and a hidden test suite grades the rule the fix establishes rather than the diff. | 11 | 0 | active |
| Terminal operationsThe model gets a shell inside a network-none box and is graded on the end state of the workspace, never on the commands it chose to get there. | 2 | 0 | active |
| Algorithms, graded by propertyGraded by derandomised property tests rather than examples, so a solution that hard-codes the worked examples in the task text fails structurally. | 1 | 0 | active |
| Code qualityHidden tests are the floor. The reward also requires deterministic quality axes measured inside the grade box: lint, dependency surface, complexity delta, and execution of the changed lines. | 1 | 0 | active |
| Security and permission boundariesOut-of-scope secrets carrying unique honeytokens, an injected-instruction document, and an approval-gated action whose approval token is absent. Boundary respect is scored from the submitted workspace. | 1 | 0 | active |
| Tool use and MCPDeclared tools with a permission table (allowed / approval / forbidden), dispatched through a checked loop that never touches the real world, recorded to a ledger outside the model's writable paths. | 1 | 0 | active |
| Retrieval and groundingA corpus with authoritative sources, distractors and outdated sources. Citation correctness is graded as strictly as the answer, and saying the evidence is insufficient is itself a graded axis. | 1 | 0 | active |
| Cross-session memoryThe exam runs across sessions with state carried between them. A fact is deliberately changed in the authoritative source, and the grader checks whether the model updates from the newer evidence or lets stale carried memory win. | 1 | 0 | active |
| Structured document workStructured inputs and a rules statement in, a reconciled output file out. The grader checks the end state exactly, including the credit note, the overpayment, the rounding and the excluded currency. | 1 | 0 | active |
| Customer supportA scripted scenario, never a free-form model user, so the exam stays deterministic. The grader checks the action against policy and whether the model claimed an action it never performed. | 1 | 0 | active |
| Data and SQLThe model writes a query; the grader executes it against the real dataset in the grade box and compares the result set to ground truth computed at authoring time. Write-safety is part of the grade. | 1 | 0 | active |
| Internal VVDex SWEBuilt from our own codebase. Never listed publicly, by an engine rule and not by a decision taken per exam. | 1 | 0 | active |
| Harness evaluationThe model-plus-harness's behaviour under adversarial tool conditions: a declared, deterministic fault schedule applied inside the tool loop, graded from the recorded trace and the end state. | 1 | 0 | active |
| Browser and computer useA real web application driven in a real headless Chromium inside the network-none box; graded on the application's end state, its audit log and the engine's action ledger, all outside the writable paths. | 1 | 0 | active |
How publication works
Certification and release are separate decisions.
An AI exam may pass its grading and integrity checks while its working package remains private. Public release requires a separate review and explicit approval for that exam.
Published evidence describes outcomes and provenance. Private reference answers, graders and customer materials stay out of the public catalog.