Never published
No holdout problem appears in the public benchmark, in a paper, or anywhere else. Contributors are bound by non-publication terms before they see any of it.
The private holdout
The holdout is a problem set that has never been published and never will be. It is refreshed every quarter with problems formalised from research published after the training cutoff of the models being tested. It is the product.
Each quarter, your named models are run against the current holdout set under the method. You receive the pass rate per model, the full transcripts (prompt, response, assembled file, compiler output, axiom report) for every problem, and the Antefacts certificate that the run was conducted as pre-registered. You do not receive the problems themselves in advance, and they are not reused once exposed to your models.
Over time the quarterly runs form a series. A board or a regulator can see how a model family's reasoning has moved, measured the same way each quarter, on problems none of the models had seen. A later entrant cannot reconstruct that history, and history is what makes a measurement programme useful.
No holdout problem appears in the public benchmark, in a paper, or anywhere else. Contributors are bound by non-publication terms before they see any of it.
Every problem is drawn from literature published after the training cutoff of the models under test, so the result is not in the training data even informally.
Separate servers, separate credentials, no shared problem store with the public set. Access is two-person and logged.
A problem shown to a customer's model is retired from that customer's future runs. Refresh is quarterly and continuous.
Marker problems are seeded so that any leak would surface in the next round of model outputs.
The problem set and scoring for each run are hash-committed before the run, and the run is certified by Antefacts. See certification.
| Customer | Why they buy | Shape |
|---|---|---|
| Frontier laboratories | An uncontaminated measurement across model families, ahead of release, with results they can put in front of a regulator. | Annual, multiple model families, quarterly runs. Entry by a two-quarter pilot. |
| Government safety institutes | Their mandate is oversight rather than product development. A quarterly independent measurement is what that mandate requires. | Annual, named model set, quarterly runs. |
| Regulated enterprise and defence | They deploy models they did not train and must evidence to their own regulator that the reasoning was tested independently. | Annual, deployed model set, quarterly runs. |
Custom problem sets in a specific domain are available for buyers with a particular requirement. Annual contracts, invoiced quarterly. Pricing is discussed in conversation, not published.
A lab that builds this internally has produced a number nobody outside the lab will accept, which defeats the purpose. The moment the measurement is in-house it stops being evidence and becomes marketing. Invigil is useful to a lab precisely because Invigil is not the lab.