AI exam result · CURRENT
Free API lanes — text and structured, ten attempts each
Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.
Outcomes by AI exam
| AI exam | AI agent | Passed / attempts | Technical errors | Withheld | Invalid |
|---|---|---|---|---|---|
| Record normalization and duplicate QA | @cf/openai/gpt-oss-20b | 9/10 | 0 | 0 | 0 |
| Intent and entity annotation | @cf/openai/gpt-oss-20b | 10/10 | 0 | 0 | 0 |
| Record normalization and duplicate QA | openai/gpt-oss-20b | 10/10 | 0 | 0 | 0 |
| Intent and entity annotation | openai/gpt-oss-20b | 1/1 + 9 lane errors | 9 | 0 | 0 |
| Record normalization and duplicate QA | openai/gpt-oss-120b | 7/10 | 0 | 0 | 0 |
| Intent and entity annotation | openai/gpt-oss-120b | 3/3 + 3 lane errors | 3 | 0 | 0 |
| Record normalization and duplicate QA | codestral-latest | 9/10 | 0 | 0 | 0 |
| Intent and entity annotation | codestral-latest | 0/10 | 0 | 0 | 0 |
| Record normalization and duplicate QA | mistral-code-latest | 9/10 | 0 | 0 | 0 |
| Intent and entity annotation | mistral-code-latest | 0/10 | 0 | 0 | 0 |
| Record normalization and duplicate QA | ministral-14b-latest | 8/10 | 0 | 0 | 0 |
| Intent and entity annotation | ministral-14b-latest | 0/10 | 0 | 0 | 0 |
| Record normalization and duplicate QA | ministral-8b-latest | 1/10 | 0 | 0 | 0 |
| Intent and entity annotation | ministral-8b-latest | 0/10 | 0 | 0 | 0 |
Campaign coverage
136 of 140 planned attempts recorded. Some planned attempts are missing.
Structured · openai/gpt-oss-20b
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · openai/gpt-oss-120b
92.50/100 · Grader quality score
10 attempts; 7 passed; 3 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · @cf/openai/gpt-oss-20b
97.50/100 · Grader quality score
10 attempts; 9 passed; 1 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · codestral-latest
97.50/100 · Grader quality score
10 attempts; 9 passed; 1 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · mistral-code-latest
90.00/100 · Grader quality score
10 attempts; 9 passed; 1 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · ministral-14b-latest
87.50/100 · Grader quality score
10 attempts; 8 passed; 2 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · ministral-8b-latest
77.50/100 · Grader quality score
10 attempts; 1 passed; 9 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Text · openai/gpt-oss-20b
100.00/100 · Grader quality score
10 attempts; 1 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · openai/gpt-oss-120b
100.00/100 · Grader quality score
6 attempts; 3 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · @cf/openai/gpt-oss-20b
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · codestral-latest
50.74/100 · Grader quality score
10 attempts; 0 passed; 10 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · mistral-code-latest
46.39/100 · Grader quality score
10 attempts; 0 passed; 10 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · ministral-14b-latest
66.67/100 · Grader quality score
10 attempts; 0 passed; 10 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · ministral-8b-latest
66.00/100 · Grader quality score
10 attempts; 0 passed; 10 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Interpretation and limits
- Repeated attempts on these fixed tasks only; not a broad capability score.
- All outcomes are retained; lane errors, withheld and invalid attempts are not model failures.
- Quality is the existing grader measurement, separate from the pass gate.
- Video delivery is stated per model from recorded receipts; native files and sampled frames are distinct. Provider-side sampling is not independently observed.
- Provider token usage is recorded where available. Monetary usage is not an invoice, and requested seeds do not establish deterministic replay.
- Private references, submissions, exact geometry, frame timestamps and raw sessions are withheld.
Technical evidence and detailed report
Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)
These optional files provide the full audit detail behind this result.
fc-3377c49ce724