VVDex ForgeAI exams. Evidence. Independent grading.

AI exams · Multimodal annotation

Know what your AI
can actually do.

Build reliable exams for AI agents. Run real tasks, grade the work independently, and inspect the evidence behind each result.

Software, browser actions, knowledge, memory and multimodal annotation—all part of Forge, with task-specific grading and a shared standard of evidence.

This site publishes evidence. The working tools run locally in VVDex Studio.

Inside ForgeTASK → EVIDENCE

Real work.
Independent judgment.

A software fix. A browser task. A grounded answer. Every exam needs a reliable way to tell whether the work succeeded.

  1. 01
    Define the taskGive the agent a workspace, tools and a clear objective.
  2. 02
    Test the testCheck reference work, near misses and attempts to manipulate grading.
  3. 03
    Run the AI agentKeep the task separate from its private grader and reference.
  4. 04
    Inspect the resultReview actions, outcomes, failures and the evidence behind them.
Follow a real software exam

What you can do

Two kinds of work.
The same standard of proof.

Evaluate an agent’s ability to complete a task, or check the quality of an annotation. Keep the result connected to the work that produced it.

01 / AI EXAMS

Give AI agents
work worth testing.

Run defined tasks across software, browser, knowledge and other task families. Follow the actions, inspect the submission and understand the final verdict.

Explore AI exams

02 / ANNOTATION QUALITY

Turn observations
into checkable data.

Inspect source media, label events and objects, validate submissions and compare annotation sets. Grade the result against a private reference.

Explore annotation quality

Explore the work

A broad range of tasks.
A clear route into each.

Start with a published AI exam or an annotation environment. Each explains the task, what is checked, and which evidence is available.

Observed model results

Two campaigns.
The work behind the numbers.

These campaigns cover different fixed tasks and run conditions. Read them separately; neither is a universal model ranking.

AI exam campaign · September 1–23 MODEL LANES

Five kinds of work.
150 recorded attempts.

Software, grounded knowledge, memory, fault recovery and browser actions. Ten attempts per task for each model lane.

Recorded lanePassed / graded
Codex CLI50 / 50
Claude Sonnet CLI40 / 50
Codestral3 / 48

Two Codestral attempts ended in lane errors and are outside the graded denominator. These are the recorded September 1–2 results.

Inspect the AI exam campaign

Annotation campaign · September 6

Four fixed tasks.
Forty retained attempts.

GPT-5.5 was evaluated on image, sampled-video, text and structured-data annotation. Audio was not evaluated.

Read the published report for task-level passes, failures, quality measurements and input limitations.

Read the annotation campaign →

Read the evidence

A passed check means
something specific.

A certified environment is not a successful model run. Forge keeps those two claims separate, so you can see exactly what was verified.

Is the test trustworthy?

Environment certification checks that reference work passes, incorrect work fails and common attempts to manipulate grading are blocked.

Read certification evidence

How did the model perform?

Campaign reports show observed attempts, passes, failures and limitations. Results belong to the specific model, task and run—not every possible use.

Read model results

Go beyond the headline. Published reports link outcomes to recorded evidence. Where record bytes are public, you can verify their digest in your browser.

Verify evidence →

VVDex Studio

The place to do the work.

Forge’s working tools run in the local Studio app: inspect an exam, run an AI agent, review source media and annotations, and examine independently graded results.

  1. Inspect the task
  2. Run the work
  3. Review the submission
  4. Grade independently
  5. Inspect the evidence
Built for work that can be checked.Understand the methodology
02

The certification chain

public.swe.pytoolz-toolz-issue-602 · docker · 2026-08-31

One real exam, walked through the chain that certified it. Select a stage; every string in its receipt was written by the machine and is reproduced character for character — A the Docker promotion receipt, B the factory log’s verbatim subprocess walk, C that wave’s freeze table.

Sources in full: /data/specimen-chain.json

Seed

ok

The exam is built into a fresh workspace inside a network-none container.

  • "seed": { "ok": true }A
  • family=public_oss_swe (active); entry=pytest; classification=public-ossB

Gold proof

isolated

Upstream’s own merged fix must pass the hidden grader; the untouched tree must fail it. Both, or the exam is worthless in one direction or the other.

  • "gold": { "ok": true, "promotionEligible": true, "isolated": true }A
  • isolated=True, gradeWithoutGold.ok=False, gradeWithGold.ok=True, leaks=[], graderVisibleToAgent=FalseB

Grader controls

distinguished

Seven deliberately wrong submissions, each of which has to land on a known verdict. The lookalike is the one that matters: a near-miss a competent engineer would actually write.

  • "grader": { "ok": true, "promotionEligible": true, "isolated": true, "controlsDistinguished": true }A
  • untouched F / gold P / incomplete F / lookalike F / plantedHook:conftest.py F / plantedReward F / goldAfterTamperAttempts PB
  • hostGraderUnchanged=True, plantedRewardIgnored=True, plantedHooksIgnored=TrueB

Attack suite

no trivial exploit

Ten scripted probes go after the grade instead of the task — rewriting tests, planting hooks, reaching for the answer key. The reference solution has to still pass afterwards.

  • "attack": { "ok": true, "promotionEligible": true, "trivialExploits": [], "passedForPromotion": true, "goldStillPassesAfterSuite": true }A
  • 10 probes, trivialExploits=[], goldStillPassesAfterSuite=TrueB

Evaluation

eligible

The exam is run end to end in the same container it was certified in, and the model-facing tree is scanned against the publication rule list.

  • "eval": { "ok": true, "promotionEligible": true }A
  • scan_publication.py on the model-facing tree plus the published graderB

Promoted

allOk

Every stage passed, so the exam is sealed against the image it was certified with. This is the point at which a score off it becomes quotable — against this freeze and no other.

  • "allOk": trueA
  • frozen agent files17C
  • treeSha256 (16)9999fb81b48354b3C
  • freeze hash (16)e036a66feefc4c98C
  • hidden-test lines190C
The VVDex Forge hallmark, struck: an anvil punched inside a certification seal. VVDEX FORGE CERTIFICATION CHAIN

The mark is not on the exam. The mark is on the fact that the chain above ran and every station answered. The freeze figures here are from the subprocess wave — it records imageDigest: null honestly, and the Docker pass re-freezes.

03

The sealed box

Diagram: the agent works in a network-none container; the grader and the reference solution live in a separate grade box the agent can never reach. AGENT CONTAINER network = none GRADE BOX hidden tests · reference solution gold submission gold never crosses

gold never mounts in the agent box · network=none · graderVisibleToAgent=False

Agent CLIs (Claude Code, Codex) run on the host, not in the agent container: their containment is audited after each run from the CLI's own session store, and a lane that cannot be verified is withheld. Full rule: methodology · from a grader verdict to a certified result.

Each stage, and what it does not prove, is written out in the methodology.

04

Families

Historical snapshot · 2026-09-02

Annotation · five certified fixtures

One family for reviewing and labeling multimodal data, with explicit task guidelines, ambiguity handling and deterministic quality checks. These fixtures extend Forge's coverage; their certification results are separate from the historical model campaigns.

Robotic video · Image objects · Text labeling · Structured data · Audio events

Public open-source SWE

11 certifiedactive

Terminal operations

2 certifiedactive

Algorithms, graded by property

1 certifiedactive

Code quality

1 certifiedactive

Security & permission boundaries

1 certifiedactive

Tool use & MCP

1 certifiedactive

Retrieval & grounding

1 certifiedactive

Cross-session memory

1 certifiedactive

Structured document work

1 certifiedactive

Customer support

1 certifiedactive

Data & SQL

1 certifiedactive

Internal VVDex SWE

1 certified · never listedactive

Harness evaluation

1 certifiedactive

Browser & computer use

1 certifiedactive

Counts and status from /data/catalog.json, as of 2026-09-02. 0 exams are listed: the catalog is fail-closed. The two families that stood blocked before 2026-09-02 — browser & computer use, harness evaluation — each carry a certified exam now; both were built with a deterministic grader first and certified through the same six-stage chain. Full register: the catalog.

05

Ten probes, zero exploits

Diagram: ten attack probes strike the exam from all sides and every one is deflected at the boundary; the receipt reads trivialExploits, empty list.

10 probes, trivialExploits=[], goldStillPassesAfterSuite=True B

The probes go after the grade, not the task: rewriting tests, planting hooks, reaching for the answer key. An exam ships only when no probe finds a trivial exploit and the reference solution still passes after the whole suite has run.

VVDex Studio · local workspace

Inspect the source.
Stand behind the result.

The working environment for Forge: review source media, prepare annotations, run AI evaluations and inspect the grade.

Studio runs locally. This website publishes selected evidence; it does not offer a hosted annotation editor or public Studio download.

AI exams

From task to recorded outcome.

  1. Inspect the exam

    Review the task, available tools, constraints and certification evidence before interpreting a model result.

  2. Set the run conditions

    Choose the model lane and record the intended attempts. Keep those conditions visible in the resulting evidence.

  3. Run the task

    Let the AI agent work in the exam environment while the private grader and reference remain separate.

  4. Read the outcomes

    Inspect passes, model failures and execution errors separately. Follow the recorded actions behind a result.

  5. Compare and report

    Compare results under their recorded conditions and preserve the evidence that supports the report.

Explore AI exams →

Annotation

From source to a checked submission.

  1. Inspect

    Open the environment and review its source. Play or seek supported video and audio, inspect images, or read text and structured records.

  2. Annotate

    Prepare a structured submission with the labels, spans, geometry or time intervals required by the task.

  3. Validate

    Check the submission’s shape and consistency. Review validation feedback before submitting it for a grade.

  4. Grade

    Submit for independent grading against a private reference. Read the pass verdict and the applicable quality measurements separately.

  5. Compare

    Compare annotation sets and inspect disagreements. Agreement between submissions is a different question from correctness against the reference.

Explore annotation →

Discuss an evaluation or annotation workflow.

Get in touch
07

Run your model against it

You bring the endpoint and the key; the exams and the grading are ours. The key is read from a file into the job, never from the environment and never from the command line, and it is scrubbed from every artifact a run produces — including the error text a rejecting gateway sends back. You get a written report and a durable record that attests to its own bytes.

What this is not

Task-specific results — published model campaign scores belong to their named exams and revisions; fixture certification alone says nothing about a model’s performance. Not a public exam list — zero listed, fail-closed, cleared one at a time. Not contamination-free — where a public fix pre-dates a model’s cutoff, that is recorded on the exam, not averaged into a headline. Not self-service — no sign-up, no dashboard, no billing.

Public by design · kept by design

Public: the methodology in full, the family list and counts, chosen campaign reports, the ledger verification recipe, and — once cleared — a listed exam’s title, family and badge. Kept: the factory, the scouts, the engine, the evidence chain behind a certification, the run history, and every customer’s results, which belong to the customer.