VVDex ForgeAI exams. Evidence. Independent grading.

Connect your model · manual intake

There is no form. There is an email, and a human who answers it.

Intake is approved one request at a time. That is not a placeholder for a signup flow that is coming next month — it is how this runs today, and saying otherwise on a page with a text box that goes nowhere would be the first thing you could catch us lying about.

How to ask

Write to vamsi@vamsivenkatesh.com

Put Forge run in the subject line and answer the five questions below. A short answer to each is better than a long document, and an honest “not sure yet” is a fine answer to any of them.

  1. What is the model? Its name and version as you refer to it, and whether it is your own or one you are evaluating on someone else's behalf.
  2. How is it reached? An OpenAI-compatible endpoint is what the runner speaks. Say whether it needs a key or is keyless, and whether it is reachable from outside your network.
  3. What are you trying to find out? A regression against a previous version, a read on a capability area, or evidence for someone else — these lead to different exam selections.
  4. Which families interest you? The catalog lists all of them with what each grades.
  5. Who sees the result? Yours alone by default. Say if you intend to show it to anyone, because that changes what is worth arranging up front.

Open a message with the subject filled in →

Do not send a key by email

Say in the message that a key is needed — do not put one in it. Handing a credential over is arranged separately, once there is a run to arrange it for. A key arrives as a file the job reads directly; it is never taken from the process environment and never passed on a command line, where it would sit in a shell history and in the process list.

An empty key file is refused before any job runs, and a model that needs a key and was given neither a key nor an explicit keyless declaration is refused end to end rather than quietly attempted. How the key is kept out of artifacts — including the error text a rejecting gateway sends back — is written out in the methodology.

What happens after
  • A reply, or a no. Requests are approved one at a time, and a request that is not a fit gets told so rather than queued indefinitely.
  • Exam selection, agreed in writing. Which exams, which freeze of each, and how many samples — before anything runs, so the result cannot be re-scoped after the fact.
  • The run, against your endpoint. Exams and grading are ours; the model is yours.
  • A written report and a durable record. The report is a document you can paste into an email or an issue. The record attests to its own bytes and can be verified without our software. See ledger and verification.
  • Optionally, an anchored receipt. Off by default, and not yet built — zero receipts have been anchored. Do not plan around it.
Honest limits, before you write
  • The inventory is 25 exams across 14 families as counted on 2026-09-02, and none is listed as a runnable public exam. The new Annotation collection publishes evidence for five fixtures separately from that historical inventory.
  • Grading is deterministic by design. That is a strength on the axes it covers and a real limit elsewhere — a correct answer reached through an entry point the grader does not watch is marked wrong. The methodology says where.
  • Contamination is not solved. Exams built from public upstream fixes are public upstream. Exposure is recorded per exam.
  • Nothing about your run becomes public. A customer-private exam is never listed and a customer's result is never published. If you want a result to be provable to a third party, that is yours to disclose.
  • This is a small operation. One person builds and runs it. Turnaround is a conversation, not an SLA.