VVDex Forge AI exams. Evidence. Independent grading.

Multimodal Annotation · video

Robot catch temporal and tracking annotation

A robot-catching video is reviewed for temporal events, frame-level object boxes, track identity and occlusion.

Environment certification: certified. Certification checks whether the environment grades work reliably.

AI agent results on this task

AI agentAttemptsPassedFailedLane errorsWithheldInvalid
meta/muse-spark-1.31246200

Repeated attempts on this fixed task. Quality measurements are separate from the pass gate; sampled video frames do not establish native video capability.

Campaign report and limitations · Measured results JSON

What is reviewed

Annotation types: box, classification, event, segment. Rubric version 1.1.

Source: High-speed catching system, Z22 · 2012-03-19. CC-BY-SA-4.0. Original recording unchanged. Reference annotations remain private.

Certification evidence

  • Empty baseline: fail; observed pass rate 0.0, 2 independent seeds.
  • Private reference: pass; observed pass rate 1.0, 2 independent seeds.
  • Controls: 19 cases; distinguished: True.
  • Attacks: 9 of 9 blocked; trivial exploits: 0.

These are environment acceptance checks, not model performance scores.

Reference grader metrics

Recorded aggregate results of the private reference check. Exact labels, timestamps and geometry are withheld.

MetricObserved value
boxF11.0
boxIoU1.0
boxPrecision1.0
boxRecall1.0
classificationAccuracy1.0
classificationF11.0
classificationPrecision1.0
classificationRecall1.0
eventF11.0
eventPrecision1.0
eventRecall1.0
eventTemporalIoU1.0
segmentF11.0
segmentPrecision1.0
segmentRecall1.0
segmentTemporalIoU1.0
trackingConsistency1.0

Agreement and validation

Submitted annotation sets can be compared for agreement, disagreement, ambiguity, missing and invalid annotations. Disagreement alone is not a failed reference grade. No multi-annotator study was run for this release; agreement study metrics are N/A.

Validation covers malformed input, schema, bounds and modality-specific consistency. Error-class details and supported capabilities are documented in the report JSON.

Runtime, fingerprints and verification files

Runtime and isolation

{
  "fsPolicy": "agent-tree-only",
  "imageDigest": "sha256:679b36225a1f83238efba8c71b9cde3c24caaeacc9b682ae92819cb55eec328f",
  "networkPolicy": "deny",
  "resourceLimits": {
    "cpus": "1",
    "memory": "512m",
    "network": "none",
    "pids": 128,
    "readOnlyRoot": true,
    "timeoutS": 120
  }
}

Verify this evidence

vvdex.annotation.robot-video-1

Environment fingerprint:

85e026007af17db692614084f12ff40c5bbf2348bcfda0a01deb95aad8928f4c

Certification evidence digest:

0a0a51360cde69737c9b5caef895b8b7c2a5e2d66da89617d8776ac724b019b9

Public descriptor · Certification receipt · Provenance · Report JSON · File digests · Verification instructions

Limitations and ownership

Small constructed engineering fixtures, not a production dataset or professional annotation engagement. Video uses sparse frame boxes, not dense tracking. Audio contains tones and silence, not speech or speakers. CVAT XML 1.1 and Label Studio JSON cover tested box subsets only; unsupported or lossy conversions are refused. Polygon comparison requires the implemented vertex representation. The attached model campaign measures these fixed tasks only; general capability and annotator population performance are not inferred.

Public proof demonstrates the result. The reusable evaluation instrument remains private. Terms and ownership.

All five modalities