AI exam result · HISTORICAL
Gemini 3.8 Flash native-video pilot — interrupted before grading
Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.
Outcomes by AI exam
| AI exam | AI agent | Passed / attempts | Technical errors | Withheld | Invalid |
|---|---|---|---|---|---|
| Robot catch temporal and tracking annotation | gemini-3.8-flash | 0/0 + 3 lane errors | 3 | 0 | 0 |
Campaign coverage
3 of 3 planned attempts recorded. All planned attempts are recorded.
Video · gemini-3.8-flash
Not measured · Grader quality score
3 attempts; 0 passed; 0 failed. Quality measures answer accuracy; passing also requires the exam’s declared minimums.
Input: delivery not established for every attempt.
Interpretation and limits
- Repeated attempts on these fixed tasks only; not a broad capability score.
- All outcomes are retained; lane errors, withheld and invalid attempts are not model failures.
- Quality is the existing grader measurement, separate from the pass gate.
- Video delivery is stated per model from recorded receipts; native files and sampled frames are distinct. Provider-side sampling is not independently observed.
- Provider token usage is recorded where available. Monetary usage is not an invoice, and requested seeds do not establish deterministic replay.
- Private references, submissions, exact geometry, frame timestamps and raw sessions are withheld.
Technical evidence and detailed report
Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)
These optional files provide the full audit detail behind this result.
fc-6fb3c7588801