Multimodal Annotation · image
Object geometry annotation
A constructed image is reviewed for classification, object boxes, attributes and keypoints.
Environment certification: certified. Certification checks whether the environment grades work reliably.
AI agent results on this task
| AI agent | Attempts | Passed | Failed | Lane errors | Withheld | Invalid |
|---|---|---|---|---|---|---|
| cli/opencode-muse-spark | 10 | 0 | 10 | 0 | 0 | 0 |
Repeated attempts on this fixed task. Quality measurements are separate from the pass gate; sampled video frames do not establish native video capability.
What is reviewed
Annotation types: box, classification, keypoint. Rubric version 1.0.
Constructed project fixture. No customer dataset or professional annotation engagement is claimed.
Certification evidence
- Empty baseline: fail; observed pass rate 0.0, 2 independent seeds.
- Private reference: pass; observed pass rate 1.0, 2 independent seeds.
- Controls: 19 cases; distinguished: True.
- Attacks: 9 of 9 blocked; trivial exploits: 0.
These are environment acceptance checks, not model performance scores.
Reference grader metrics
Recorded aggregate results of the private reference check. Exact labels, timestamps and geometry are withheld.
| Metric | Observed value |
|---|---|
| boxF1 | 1.0 |
| boxIoU | 1.0 |
| boxPrecision | 1.0 |
| boxRecall | 1.0 |
| classificationAccuracy | 1.0 |
| classificationF1 | 1.0 |
| classificationPrecision | 1.0 |
| classificationRecall | 1.0 |
| keypointF1 | 1.0 |
| keypointNormalizedError | 0.0 |
| keypointPrecision | 1.0 |
| keypointRecall | 1.0 |
Agreement and validation
Submitted annotation sets can be compared for agreement, disagreement, ambiguity, missing and invalid annotations. Disagreement alone is not a failed reference grade. No multi-annotator study was run for this release; agreement study metrics are N/A.
Validation covers malformed input, schema, bounds and modality-specific consistency. Error-class details and supported capabilities are documented in the report JSON.
Runtime, fingerprints and verification files
Runtime and isolation
{
"fsPolicy": "agent-tree-only",
"imageDigest": "sha256:679b36225a1f83238efba8c71b9cde3c24caaeacc9b682ae92819cb55eec328f",
"networkPolicy": "deny",
"resourceLimits": {
"cpus": "1",
"memory": "512m",
"network": "none",
"pids": 128,
"readOnlyRoot": true,
"timeoutS": 120
}
}
Verify this evidence
vvdex.annotation.image-object-1
Environment fingerprint:
422753713aeb9f41abcc4135d637e3064cae1e770fc8d22ec9dacf260b35a3ef
Certification evidence digest:
cf945b4838a266664300b32fdd5fed0c4e781ab95eaaebfde31a0ce8ae2d7295
Public descriptor · Certification receipt · Provenance · Report JSON · File digests · Verification instructions
Limitations and ownership
Small constructed engineering fixtures, not a production dataset or professional annotation engagement. Video uses sparse frame boxes, not dense tracking. Audio contains tones and silence, not speech or speakers. CVAT XML 1.1 and Label Studio JSON cover tested box subsets only; unsupported or lossy conversions are refused. Polygon comparison requires the implemented vertex representation. The attached model campaign measures these fixed tasks only; general capability and annotator population performance are not inferred.
Public proof demonstrates the result. The reusable evaluation instrument remains private. Terms and ownership.