AI exam result · CURRENT
Codex and Claude Sonnet CLIs — text and structured, ten attempts each
Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.
Outcomes by AI exam
| AI exam | AI agent | Passed / attempts | Technical errors | Withheld | Invalid |
|---|---|---|---|---|---|
| Record normalization and duplicate QA | cli/claude-sonnet | 10/10 | 0 | 0 | 0 |
| Intent and entity annotation | cli/claude-sonnet | 10/10 | 0 | 0 | 0 |
| Record normalization and duplicate QA | cli/codex | 10/10 | 0 | 0 | 0 |
| Intent and entity annotation | cli/codex | 10/10 | 0 | 0 | 0 |
Campaign coverage
40 of 40 planned attempts recorded. All planned attempts are recorded.
Structured · cli/codex
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Structured · cli/claude-sonnet
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Text · cli/codex
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Text · cli/claude-sonnet
100.00/100 · Grader quality score
10 attempts; 10 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Interpretation and limits
- Repeated attempts on these fixed tasks only; not a broad capability score.
- All outcomes are retained; lane errors, withheld and invalid attempts are not model failures.
- Quality is the existing grader measurement, separate from the pass gate.
- Video uses sampled still images. Audio delivery is unsupported and was not attempted.
- CLI attempt indices do not claim stochastic seed control. Subscription token and monetary usage are unavailable.
- Private references, submissions, exact geometry, frame timestamps and raw sessions are withheld.
Technical evidence and detailed report
Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)
These optional files provide the full audit detail behind this result.
fc-6d6679dea47a