AI exam result · CURRENT
Gemini and OpenCode pilots — image, audio, text and structured, one attempt each
Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.
Outcomes by AI exam
| AI exam | AI agent | Passed / attempts | Technical errors | Withheld | Invalid |
|---|---|---|---|---|---|
| Object geometry annotation | cli/opencode-muse-spark | 0/1 + 1 withheld | 0 | 1 | 0 |
| Record normalization and duplicate QA | cli/opencode-muse-spark | 1/1 + 1 withheld | 0 | 1 | 0 |
| Intent and entity annotation | cli/opencode-muse-spark | 1/1 + 1 withheld | 0 | 1 | 0 |
| Audio event and silence annotation | gemini-3.6-flash | 0/0 + 1 lane error | 1 | 0 | 0 |
| Object geometry annotation | gemini-3.6-flash | 0/0 + 1 lane error | 1 | 0 | 0 |
| Record normalization and duplicate QA | gemini-3.6-flash | 0/0 + 1 lane error | 1 | 0 | 0 |
| Intent and entity annotation | gemini-3.6-flash | 0/0 + 1 lane error | 1 | 0 | 0 |
Campaign coverage
10 of 8 planned attempts recorded. Some planned attempts are missing.
Audio · gemini-3.6-flash
Not measured · Grader quality score
1 attempts; 0 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: delivery not established for every attempt.
Image · gemini-3.6-flash
Not measured · Grader quality score
1 attempts; 0 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: delivery not established for every attempt.
Image · cli/opencode-muse-spark
72.96/100 · Grader quality score
2 attempts; 0 passed; 1 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: image attachments.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
| Box matching (F1) | 0.00 |
Structured · gemini-3.6-flash
Not measured · Grader quality score
1 attempts; 0 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
Structured · cli/opencode-muse-spark
100.00/100 · Grader quality score
2 attempts; 1 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|
Text · gemini-3.6-flash
Not measured · Grader quality score
1 attempts; 0 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
Text · cli/opencode-muse-spark
100.00/100 · Grader quality score
2 attempts; 1 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: canonical text tools.
| Measure | Score |
|---|---|
| Classification accuracy | 1.00 |
Interpretation and limits
- Repeated attempts on these fixed tasks only; not a broad capability score.
- All outcomes are retained; lane errors, withheld and invalid attempts are not model failures.
- Quality is the existing grader measurement, separate from the pass gate.
- Video delivery is stated per model from recorded receipts; native files and sampled frames are distinct. Provider-side sampling is not independently observed.
- Provider token usage is recorded where available. Monetary usage is not an invoice, and requested seeds do not establish deterministic replay.
- Private references, submissions, exact geometry, frame timestamps and raw sessions are withheld.
Technical evidence and detailed report
Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)
These optional files provide the full audit detail behind this result.
fc-6f40beb6605e