Invigil

The private holdout

The only measurement of your models that cannot have leaked into them.

The holdout is a problem set that has never been published and never will be. It is refreshed every quarter with problems formalised from research published after the training cutoff of the models being tested. It is the product.

What you receive

A quarterly measurement, and a time series

Each quarter, your named models are run against the current holdout set under the method. You receive the pass rate per model, the full transcripts (prompt, response, assembled file, compiler output, axiom report) for every problem, and the Antefacts certificate that the run was conducted as pre-registered. You do not receive the problems themselves in advance, and they are not reused once exposed to your models.

Over time the quarterly runs form a series. A board or a regulator can see how a model family's reasoning has moved, measured the same way each quarter, on problems none of the models had seen. A later entrant cannot reconstruct that history, and history is what makes a measurement programme useful.

Why it stays clean
01

Never published

No holdout problem appears in the public benchmark, in a paper, or anywhere else. Contributors are bound by non-publication terms before they see any of it.

02

After the cutoff

Every problem is drawn from literature published after the training cutoff of the models under test, so the result is not in the training data even informally.

03

Separate infrastructure

Separate servers, separate credentials, no shared problem store with the public set. Access is two-person and logged.

04

Burned after exposure

A problem shown to a customer's model is retired from that customer's future runs. Refresh is quarterly and continuous.

05

Canaries

Marker problems are seeded so that any leak would surface in the next round of model outputs.

06

Certified

The problem set and scoring for each run are hash-committed before the run, and the run is certified by Antefacts. See certification.

Who subscribes

Three kinds of customer

CustomerWhy they buyShape
Frontier laboratoriesAn uncontaminated measurement across model families, ahead of release, with results they can put in front of a regulator.Annual, multiple model families, quarterly runs. Entry by a two-quarter pilot.
Government safety institutesTheir mandate is oversight rather than product development. A quarterly independent measurement is what that mandate requires.Annual, named model set, quarterly runs.
Regulated enterprise and defenceThey deploy models they did not train and must evidence to their own regulator that the reasoning was tested independently.Annual, deployed model set, quarterly runs.

Custom problem sets in a specific domain are available for buyers with a particular requirement. Annual contracts, invoiced quarterly. Pricing is discussed in conversation, not published.

Why not build it

Independence is the product

A lab that builds this internally has produced a number nobody outside the lab will accept, which defeats the purpose. The moment the measurement is in-house it stops being evidence and becomes marketing. Invigil is useful to a lab precisely because Invigil is not the lab.

Start a conversation.