VVDex Forge AI exams. Evidence. Independent grading.

Model results

See what happened.
Then inspect the evidence.

Real model attempts, recorded outcomes and the limits of each measurement. A failed attempt stays in the report. A technical error is identified separately.

Earlier published campaign · CURRENT

Lane pilots — fourteen API lanes and three CLIs on text and structured, one attempt each

2 tasks · 17 model lanes · 40 attempts · 13 certified passes

Models: cli/codex, cli/claude-sonnet, cli/cursor-grok, mistral-small-latest, mistral-medium-latest, magistral-small-latest, openai/gpt-oss-20b, openai/gpt-oss-120b, codestral-latest, mistral-code-latest, ministral-14b-latest, ministral-8b-latest, @cf/openai/gpt-oss-20b, nvidia/nemotron-3-super-120b-a12b:free, cohere/north-mini-code:free, minimax/minimax-m3:free, nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free. Read the report for individual task outcomes, exclusions and measurement limits.

Looking for annotation certification?

Certification checks whether an environment grades work reliably. It is separate from how a model performs on that environment.

Explore the five annotation environments →

Counts describe these campaigns only. They do not establish a general model ranking. Detailed reports explain the task scope and which attempts enter each calculation.