01 / AI EXAMS
Give AI agents
work worth testing.
Run defined tasks across software, browser, knowledge and other task families. Follow the actions, inspect the submission and understand the final verdict.
Explore AI examsAI exams · Multimodal annotation
Build reliable exams for AI agents. Run real tasks, grade the work independently, and inspect the evidence behind each result.
Software, browser actions, knowledge, memory and multimodal annotation—all part of Forge, with task-specific grading and a shared standard of evidence.
This site publishes evidence. The working tools run locally in VVDex Studio.
A software fix. A browser task. A grounded answer. Every exam needs a reliable way to tell whether the work succeeded.
What you can do
Evaluate an agent’s ability to complete a task, or check the quality of an annotation. Keep the result connected to the work that produced it.
01 / AI EXAMS
Run defined tasks across software, browser, knowledge and other task families. Follow the actions, inspect the submission and understand the final verdict.
Explore AI exams02 / ANNOTATION QUALITY
Inspect source media, label events and objects, validate submissions and compare annotation sets. Grade the result against a private reference.
Explore annotation qualityExplore the work
Start with a published AI exam or an annotation environment. Each explains the task, what is checked, and which evidence is available.
Observed model results
These campaigns cover different fixed tasks and run conditions. Read them separately; neither is a universal model ranking.
Software, grounded knowledge, memory, fault recovery and browser actions. Ten attempts per task for each model lane.
Two Codestral attempts ended in lane errors and are outside the graded denominator. These are the recorded September 1–2 results.
Inspect the AI exam campaignAnnotation campaign · September 6
GPT-5.5 was evaluated on image, sampled-video, text and structured-data annotation. Audio was not evaluated.
Read the published report for task-level passes, failures, quality measurements and input limitations.
Read the annotation campaign →Read the evidence
A certified environment is not a successful model run. Forge keeps those two claims separate, so you can see exactly what was verified.
Environment certification checks that reference work passes, incorrect work fails and common attempts to manipulate grading are blocked.
Read certification evidenceCampaign reports show observed attempts, passes, failures and limitations. Results belong to the specific model, task and run—not every possible use.
Read model resultsGo beyond the headline. Published reports link outcomes to recorded evidence. Where record bytes are public, you can verify their digest in your browser.
Verify evidence →VVDex Studio
Forge’s working tools run in the local Studio app: inspect an exam, run an AI agent, review source media and annotations, and examine independently graded results.
One real exam, walked through the chain that certified it. Select a stage; every string in its receipt was written by the machine and is reproduced character for character — A the Docker promotion receipt, B the factory log’s verbatim subprocess walk, C that wave’s freeze table.
The exam is built into a fresh workspace inside a network-none container.
"seed": { "ok": true }Afamily=public_oss_swe (active); entry=pytest; classification=public-ossBUpstream’s own merged fix must pass the hidden grader; the untouched tree must fail it. Both, or the exam is worthless in one direction or the other.
"gold": { "ok": true, "promotionEligible": true, "isolated": true }Aisolated=True, gradeWithoutGold.ok=False, gradeWithGold.ok=True, leaks=[], graderVisibleToAgent=FalseBSeven deliberately wrong submissions, each of which has to land on a known verdict. The lookalike is the one that matters: a near-miss a competent engineer would actually write.
"grader": { "ok": true, "promotionEligible": true, "isolated": true, "controlsDistinguished": true }Auntouched F / gold P / incomplete F / lookalike F / plantedHook:conftest.py F / plantedReward F / goldAfterTamperAttempts PBhostGraderUnchanged=True, plantedRewardIgnored=True, plantedHooksIgnored=TrueBTen scripted probes go after the grade instead of the task — rewriting tests, planting hooks, reaching for the answer key. The reference solution has to still pass afterwards.
"attack": { "ok": true, "promotionEligible": true, "trivialExploits": [], "passedForPromotion": true, "goldStillPassesAfterSuite": true }A10 probes, trivialExploits=[], goldStillPassesAfterSuite=TrueBThe exam is run end to end in the same container it was certified in, and the model-facing tree is scanned against the publication rule list.
"eval": { "ok": true, "promotionEligible": true }Ascan_publication.py on the model-facing tree plus the published graderBEvery stage passed, so the exam is sealed against the image it was certified with. This is the point at which a score off it becomes quotable — against this freeze and no other.
"allOk": trueAfrozen agent files17CtreeSha256 (16)9999fb81b48354b3Cfreeze hash (16)e036a66feefc4c98Chidden-test lines190CThe mark is not on the exam. The mark is on the fact that the chain above ran and
every station answered. The freeze figures here are from the subprocess wave — it
records imageDigest: null honestly, and the Docker pass re-freezes.
gold never mounts in the agent box · network=none ·
graderVisibleToAgent=False
Agent CLIs (Claude Code, Codex) run on the host, not in the agent container: their containment is audited after each run from the CLI's own session store, and a lane that cannot be verified is withheld. Full rule: methodology · from a grader verdict to a certified result.
One family for reviewing and labeling multimodal data, with explicit task guidelines, ambiguity handling and deterministic quality checks. These fixtures extend Forge's coverage; their certification results are separate from the historical model campaigns.
Robotic video · Image objects · Text labeling · Structured data · Audio events
10 probes, trivialExploits=[], goldStillPassesAfterSuite=True
B
VVDex Studio · local workspace
The working environment for Forge: review source media, prepare annotations, run AI evaluations and inspect the grade.
Studio runs locally. This website publishes selected evidence; it does not offer a hosted annotation editor or public Studio download.
AI exams
Review the task, available tools, constraints and certification evidence before interpreting a model result.
Choose the model lane and record the intended attempts. Keep those conditions visible in the resulting evidence.
Let the AI agent work in the exam environment while the private grader and reference remain separate.
Inspect passes, model failures and execution errors separately. Follow the recorded actions behind a result.
Compare results under their recorded conditions and preserve the evidence that supports the report.
Annotation
Open the environment and review its source. Play or seek supported video and audio, inspect images, or read text and structured records.
Prepare a structured submission with the labels, spans, geometry or time intervals required by the task.
Check the submission’s shape and consistency. Review validation feedback before submitting it for a grade.
Submit for independent grading against a private reference. Read the pass verdict and the applicable quality measurements separately.
Compare annotation sets and inspect disagreements. Agreement between submissions is a different question from correctness against the reference.
Discuss an evaluation or annotation workflow.
Get in touchYou bring the endpoint and the key; the exams and the grading are ours. The key is read from a file into the job, never from the environment and never from the command line, and it is scrubbed from every artifact a run produces — including the error text a rejecting gateway sends back. You get a written report and a durable record that attests to its own bytes.
Task-specific results — published model campaign scores belong to their named exams and revisions; fixture certification alone says nothing about a model’s performance. Not a public exam list — zero listed, fail-closed, cleared one at a time. Not contamination-free — where a public fix pre-dates a model’s cutoff, that is recorded on the exam, not averaged into a headline. Not self-service — no sign-up, no dashboard, no billing.
Public: the methodology in full, the family list and counts, chosen campaign reports, the ledger verification recipe, and — once cleared — a listed exam’s title, family and badge. Kept: the factory, the scouts, the engine, the evidence chain behind a certification, the run history, and every customer’s results, which belong to the customer.