VVDex Forge AI exams. Evidence. Independent grading.

AI exam result · CURRENT

GPT-5.5 annotation evaluation — ten attempts on each of four tasks

Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.

Download result · 2 pages · Inspect technical evidence

Outcomes by AI exam

AI examAI agentPassed / attemptsTechnical errorsWithheldInvalid
Object geometry annotationGPT-5.5 · CLI5/10000
Robot catch temporal and tracking annotationGPT-5.5 · CLI0/10000
Record normalization and duplicate QAGPT-5.5 · CLI10/10000
Intent and entity annotationGPT-5.5 · CLI10/10000

Campaign coverage

40 of 40 planned attempts recorded. All planned attempts are recorded.

Image · GPT-5.5 · CLI

99.51/100 · Grader quality score

10 attempts; 5 passed; 5 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.

Input: image attachments.

Measured quality · scale 0–1
MeasureScore
Classification accuracy1.00
Box matching (F1)1.00

Video · GPT-5.5 · CLI

28.01/100 · Grader quality score

10 attempts; 0 passed; 10 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.

Input: sampled image frames; no native video.

Measured quality · scale 0–1
MeasureScore
Classification accuracy1.00
Event matching (F1)0.20
Segment matching (F1)0.87
Box matching (F1)0.80
Track consistency on matched boxes1.00

Structured · GPT-5.5 · CLI

100.00/100 · Grader quality score

10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.

Input: canonical text tools.

Measured quality · scale 0–1
MeasureScore

Text · GPT-5.5 · CLI

100.00/100 · Grader quality score

10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.

Input: canonical text tools.

Measured quality · scale 0–1
MeasureScore
Classification accuracy1.00

Interpretation and limits

  • Repeated attempts on these fixed tasks only; not a broad capability score.
  • All outcomes are retained; lane errors, withheld and invalid attempts are not model failures.
  • Quality is the existing grader measurement, separate from the pass gate.
  • Video uses sampled still images. Audio delivery is unsupported and was not attempted.
  • CLI attempt indices do not claim stochastic seed control. Subscription token and monetary usage are unavailable.
  • Private references, submissions, exact geometry, frame timestamps and raw sessions are withheld.
Technical evidence and detailed report

Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)

These optional files provide the full audit detail behind this result.

fc-ec21239c00aa