AI exam result · HISTORICAL
Certified campaign — one rollout per lane
Observed outcomes from sealed evaluation records. Failed attempts and non-model errors remain visible.
Outcomes by AI exam
| AI exam | AI agent | Passed / attempts | Technical errors | Withheld | Invalid |
|---|---|---|---|---|---|
| Streaming items must equal what the full parse would have built | cli/claude-sonnet | 1/1 | 0 | 0 | 0 |
| Trusting an inaccurate __len__ in partition_all | cli/claude-sonnet | 1/1 | 0 | 0 | 0 |
| Signs and base indicators in the integer converter | cli/claude-sonnet | 1/1 | 0 | 0 | 0 |
| Answering questions from a document collection with outdated and distractor sources, and refusing when the evidence is insufficient | cli/claude-sonnet | 0/1 | 0 | 0 | 0 |
| Updating a cross-session fact: answer from the newer authoritative roster, not stale carried memory | cli/claude-sonnet | 1/1 | 0 | 0 | 0 |
| Streaming items must equal what the full parse would have built | cli/codex | 1/1 | 0 | 0 | 0 |
| Trusting an inaccurate __len__ in partition_all | cli/codex | 0/1 | 0 | 0 | 0 |
| Signs and base indicators in the integer converter | cli/codex | 0/1 | 0 | 0 | 0 |
| Answering questions from a document collection with outdated and distractor sources, and refusing when the evidence is insufficient | cli/codex | 1/1 | 0 | 0 | 0 |
| Updating a cross-session fact: answer from the newer authoritative roster, not stale carried memory | cli/codex | 1/1 | 0 | 0 | 0 |
| Streaming items must equal what the full parse would have built | openai/gpt-oss-120b | 0/0 + 1 lane error | 1 | 0 | 0 |
| Trusting an inaccurate __len__ in partition_all | openai/gpt-oss-120b | 0/0 + 1 lane error | 1 | 0 | 0 |
| Signs and base indicators in the integer converter | openai/gpt-oss-120b | 0/0 + 1 lane error | 1 | 0 | 0 |
| Answering questions from a document collection with outdated and distractor sources, and refusing when the evidence is insufficient | openai/gpt-oss-120b | 0/0 + 1 lane error | 1 | 0 | 0 |
| Updating a cross-session fact: answer from the newer authoritative roster, not stale carried memory | openai/gpt-oss-120b | 1/1 | 0 | 0 | 0 |
| Streaming items must equal what the full parse would have built | codestral-latest | 0/1 | 0 | 0 | 0 |
| Trusting an inaccurate __len__ in partition_all | codestral-latest | 0/1 | 0 | 0 | 0 |
| Signs and base indicators in the integer converter | codestral-latest | 0/1 | 0 | 0 | 0 |
| Answering questions from a document collection with outdated and distractor sources, and refusing when the evidence is insufficient | codestral-latest | 0/1 | 0 | 0 | 0 |
| Updating a cross-session fact: answer from the newer authoritative roster, not stale carried memory | codestral-latest | 0/1 | 0 | 0 | 0 |
Technical evidence and detailed report
Detailed technical report (PDF) · Results and record digests (JSON) · Attempt counts and commitments (JSON)
These optional files provide the full audit detail behind this result.
fc-8626f712e26f