AI exam result · CURRENT
GPT-5.5 annotation evaluation — ten attempts on each of four tasks
Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.
Outcomes by AI exam
| AI exam | AI agent | Passed / attempts | Technical errors | Withheld | Invalid |
|---|---|---|---|---|---|
| Object geometry annotation | GPT-5.5 · CLI | 5/10 | 0 | 0 | 0 |
| Robot catch temporal and tracking annotation | GPT-5.5 · CLI | 0/10 | 0 | 0 | 0 |
| Record normalization and duplicate QA | GPT-5.5 · CLI | 10/10 | 0 | 0 | 0 |
| Intent and entity annotation | GPT-5.5 · CLI | 10/10 | 0 | 0 | 0 |
Campaign coverage
40 of 40 planned attempts recorded. All planned attempts are recorded.
Image · GPT-5.5 · CLI
99.51/100 · Grader quality score
10 attempts; 5 passed; 5 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: image attachments.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
| Box matching (F1) | 1.00 |
Video · GPT-5.5 · CLI
28.01/100 · Grader quality score
10 attempts; 0 passed; 10 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: sampled image frames; no native video.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
| Event matching (F1) | 0.20 |
| Segment matching (F1) | 0.87 |
| Box matching (F1) | 0.80 |
| Track consistency on matched boxes | 1.00 |
Structured · GPT-5.5 · CLI
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Text · GPT-5.5 · CLI
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Interpretation and limits
- Repeated attempts on these fixed tasks only; not a broad capability score.
- All outcomes are retained; lane errors, withheld and invalid attempts are not model failures.
- Quality is the existing grader measurement, separate from the pass gate.
- Video uses sampled still images. Audio delivery is unsupported and was not attempted.
- CLI attempt indices do not claim stochastic seed control. Subscription token and monetary usage are unavailable.
- Private references, submissions, exact geometry, frame timestamps and raw sessions are withheld.
Technical evidence and detailed report
Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)
These optional files provide the full audit detail behind this result.
fc-ec21239c00aa