VVDex ForgeAI exams. Evidence. Independent grading.

VVDex Forge · AI exams

AI exams for real work.

Forge turns defined real-world tasks for AI agents into controlled exams. The task, environment and rules are fixed; the work happens in an isolated run; the result is graded later against a private reference.

Exam pathcontrolled
  1. 01
    Taskfixed brief + inputs
  2. 02
    Runisolated workspace
  3. 03
    Gradeprivate reference
  4. 04
    Evidenceinspectable record

A certified run separates the working area from the grader, reference answer and evidence ledger.

Two dated views · 2026-09-02

Inventory and results answer different questions.

Inventory snapshot 25

certified exams across 14 families

0 runnable packages publicly listed in that snapshot. Read the dated inventory →
Observed campaign 150

attempts across 5 tasks × 3 model lanes

148 graded · 93 certified passes · 55 model fails · 2 lane errors. Open the evidence →

The 25-exam inventory and the five-task campaign are separate published records. Their counts are not additive, and neither describes an undated live total.

What Forge measures

One standard, different kinds of work

Capability appears in the work product.

Each family defines a concrete end state: a repository repaired, a browser workflow completed, an answer grounded in the right sources, a stale memory corrected, or a tool loop recovered after a known fault.

See the family system
01 / software

Repair real code

Start from a known repository state. Grade the behavior the repair must establish, not whether the patch resembles a reference diff.

02 / browser

Operate a real interface

Drive a browser inside the box. Judge the application state, its audit trail and the recorded actions after the run.

03 / knowledge

Ground an answer

Work through authoritative, distracting and outdated sources. Evidence quality is part of the result.

04 / memory

Update what changed

Carry state across sessions, then test whether newer evidence replaces a stale remembered fact.

05 / harness

Recover from faults

Apply a declared failure schedule to the tool loop and grade the recorded recovery and final state.

Broader certified inventory · snapshot dated 2026-09-02

These five are part of a wider system.

The catalog also records terminal operations, property-graded algorithms, code quality, security boundaries, tool and MCP use, structured documents, customer support, data and SQL, and internal VVDex software work.

TerminalAlgorithmsCode qualitySecurityTools + MCPDocumentsSupportData + SQLInternal SWE
From brief to proof
01

Build the task

Freeze the brief, inputs, expected outcome and boundaries.

02

Certify the exam

Test determinism, grading and provenance before a model run.

03

Run in isolation

Separate the working area from the grader and evidence record.

04

Inspect attack checks

Check leakage, tampering, permissions and the chain of evidence.

Observed model campaign

Campaign fc-d89e429d2781 · published 2026-09-02

Five fixed tasks. Three model lanes. Every attempt accounted for.

The published matrix contains 150 attempts. It reports certified passes, model failures and lane errors separately, with published task descriptions and supporting evidence available for inspection.

This is an observed result under the campaign's stated task, harness and grading conditions. It is not a universal model ranking.

Inspect all 15 cells
Model lanePassed / graded
cli / codex5 tasks
50 / 50
cli / claude-sonnet5 tasks
40 / 50
codestral-latest2 lane errors
3 / 48

Total: 93 certified passes · 55 model fails · 2 lane errors

Annotation belongs in Forge

A first-class Forge family

Annotation carries the same evidence discipline into media and structured data.

Video, image, text, structured-data and audio fixtures are built, certified and inspected within Forge. Their published fixture reports are a related evidence surface, kept distinct from the historical model campaign above.

Follow the evidence
InventoryCertified exam familiesDated catalog and listing state MethodHow grading worksRules, boxes and reward evidence IntegrityRun and grader boundariesContainment and attack checks VerificationCheck a published recordHashes, provenance and receipts