Multimodal Annotation · video
Robot catch temporal and tracking annotation
A robot-catching video is reviewed for temporal events, frame-level object boxes, track identity and occlusion.
Environment certification: certified. Certification checks whether the environment grades work reliably.
AI agent results on this task
| AI agent | Attempts | Passed | Failed | Lane errors | Withheld | Invalid |
|---|---|---|---|---|---|---|
| meta/muse-spark-1.3 | 12 | 4 | 6 | 2 | 0 | 0 |
Repeated attempts on this fixed task. Quality measurements are separate from the pass gate; sampled video frames do not establish native video capability.
What is reviewed
Annotation types: box, classification, event, segment. Rubric version 1.1.
Source: High-speed catching system, Z22 · 2012-03-19. CC-BY-SA-4.0. Original recording unchanged. Reference annotations remain private.
Certification evidence
- Empty baseline: fail; observed pass rate 0.0, 2 independent seeds.
- Private reference: pass; observed pass rate 1.0, 2 independent seeds.
- Controls: 19 cases; distinguished: True.
- Attacks: 9 of 9 blocked; trivial exploits: 0.
These are environment acceptance checks, not model performance scores.
Reference grader metrics
Recorded aggregate results of the private reference check. Exact labels, timestamps and geometry are withheld.
| Metric | Observed value |
|---|---|
| boxF1 | 1.0 |
| boxIoU | 1.0 |
| boxPrecision | 1.0 |
| boxRecall | 1.0 |
| classificationAccuracy | 1.0 |
| classificationF1 | 1.0 |
| classificationPrecision | 1.0 |
| classificationRecall | 1.0 |
| eventF1 | 1.0 |
| eventPrecision | 1.0 |
| eventRecall | 1.0 |
| eventTemporalIoU | 1.0 |
| segmentF1 | 1.0 |
| segmentPrecision | 1.0 |
| segmentRecall | 1.0 |
| segmentTemporalIoU | 1.0 |
| trackingConsistency | 1.0 |
Agreement and validation
Submitted annotation sets can be compared for agreement, disagreement, ambiguity, missing and invalid annotations. Disagreement alone is not a failed reference grade. No multi-annotator study was run for this release; agreement study metrics are N/A.
Validation covers malformed input, schema, bounds and modality-specific consistency. Error-class details and supported capabilities are documented in the report JSON.
Runtime, fingerprints and verification files
Runtime and isolation
{
"fsPolicy": "agent-tree-only",
"imageDigest": "sha256:679b36225a1f83238efba8c71b9cde3c24caaeacc9b682ae92819cb55eec328f",
"networkPolicy": "deny",
"resourceLimits": {
"cpus": "1",
"memory": "512m",
"network": "none",
"pids": 128,
"readOnlyRoot": true,
"timeoutS": 120
}
}
Verify this evidence
vvdex.annotation.robot-video-1
Environment fingerprint:
85e026007af17db692614084f12ff40c5bbf2348bcfda0a01deb95aad8928f4c
Certification evidence digest:
0a0a51360cde69737c9b5caef895b8b7c2a5e2d66da89617d8776ac724b019b9
Public descriptor · Certification receipt · Provenance · Report JSON · File digests · Verification instructions
Limitations and ownership
Small constructed engineering fixtures, not a production dataset or professional annotation engagement. Video uses sparse frame boxes, not dense tracking. Audio contains tones and silence, not speech or speakers. CVAT XML 1.1 and Label Studio JSON cover tested box subsets only; unsupported or lossy conversions are refused. Polygon comparison requires the implemented vertex representation. The attached model campaign measures these fixed tasks only; general capability and annotator population performance are not inferred.
Public proof demonstrates the result. The reusable evaluation instrument remains private. Terms and ownership.