VVDex Forge AI exams. Evidence. Independent grading.

Multimodal Annotation · text

Intent and entity annotation

An invented message is reviewed for classification and entity spans with consistent offsets.

Environment certification: certified. Certification checks whether the environment grades work reliably.

AI agent results on this task

AI agentAttemptsPassedFailedLane errorsWithheldInvalid
cli/opencode-muse-spark10100000

Repeated attempts on this fixed task. Quality measurements are separate from the pass gate; sampled video frames do not establish native video capability.

Campaign report and limitations · Measured results JSON

What is reviewed

Annotation types: classification, span. Rubric version 1.0.

Constructed project fixture. No customer dataset or professional annotation engagement is claimed.

Certification evidence

  • Empty baseline: fail; observed pass rate 0.0, 2 independent seeds.
  • Private reference: pass; observed pass rate 1.0, 2 independent seeds.
  • Controls: 19 cases; distinguished: True.
  • Attacks: 9 of 9 blocked; trivial exploits: 0.

These are environment acceptance checks, not model performance scores.

Reference grader metrics

Recorded aggregate results of the private reference check. Exact labels, timestamps and geometry are withheld.

MetricObserved value
classificationAccuracy1.0
classificationF11.0
classificationPrecision1.0
classificationRecall1.0
spanExactF11.0
spanF11.0
spanOverlapIoU1.0
spanPrecision1.0
spanRecall1.0

Agreement and validation

Submitted annotation sets can be compared for agreement, disagreement, ambiguity, missing and invalid annotations. Disagreement alone is not a failed reference grade. No multi-annotator study was run for this release; agreement study metrics are N/A.

Validation covers malformed input, schema, bounds and modality-specific consistency. Error-class details and supported capabilities are documented in the report JSON.

Runtime, fingerprints and verification files

Runtime and isolation

{
  "fsPolicy": "agent-tree-only",
  "imageDigest": "sha256:679b36225a1f83238efba8c71b9cde3c24caaeacc9b682ae92819cb55eec328f",
  "networkPolicy": "deny",
  "resourceLimits": {
    "cpus": "1",
    "memory": "512m",
    "network": "none",
    "pids": 128,
    "readOnlyRoot": true,
    "timeoutS": 120
  }
}

Verify this evidence

vvdex.annotation.text-labeling-1

Environment fingerprint:

e1abe51398b2d950a3415678aaf8ec796420dfb05ea89b3b64afd33c1e345fd3

Certification evidence digest:

0f97b9d0560aeb01635959125ac2e126ab1103e0087113deadcb7b4ed5ace282

Public descriptor · Certification receipt · Provenance · Report JSON · File digests · Verification instructions

Limitations and ownership

Small constructed engineering fixtures, not a production dataset or professional annotation engagement. Video uses sparse frame boxes, not dense tracking. Audio contains tones and silence, not speech or speakers. CVAT XML 1.1 and Label Studio JSON cover tested box subsets only; unsupported or lossy conversions are refused. Polygon comparison requires the implemented vertex representation. The attached model campaign measures these fixed tasks only; general capability and annotator population performance are not inferred.

Public proof demonstrates the result. The reusable evaluation instrument remains private. Terms and ownership.

All five modalities