Invigil

What we do

The score a lab cannot train against.

Invigil measures how well frontier AI models reason. The problems are drawn from recent research literature and formalised in Lean 4, so a model must produce a proof that a machine can check. Lean decides whether the proof is correct. Nobody grades, nobody argues, and nothing in the private set has ever been published.

The result is a measurement of a model's reasoning that cannot have leaked into its training data, and that anyone outside the lab can reproduce.

The model writes the proof. Lean marks it.
The problem

Every published benchmark is contaminated the day it is published

Labs report scores on benchmarks that are already in their training data. Nobody can say how much of a result is reasoning and how much is recall, and the labs cannot resolve this internally, because a lab cannot mark its own homework and be believed. Regulators, boards and customers are being asked to accept numbers with no independent basis.

Invigil exists to produce the one number a lab cannot produce for itself: an uncontaminated, machine-checked, independently reproducible measurement of how its models reason.

Two lines, one asset

A public benchmark and a private holdout

Public

The benchmark

Problems, harness, pinned toolchain and full model transcripts, published so that anyone can reproduce a score without contacting us. It exists to prove the method works and to make the numbers checkable.

Private

The holdout

A separate problem set that is never published, lives on separate infrastructure, and is refreshed every quarter with problems drawn from literature published after the models under test were trained. A lab buys the only measurement of its own systems that cannot have leaked into them.

Why it holds

The model never supplies the theorem

The model receives the preamble and the statement and returns a proof body. The harness assembles the final file from its own copy of the statement, so weakening the theorem is structurally impossible rather than something detected afterwards. Every accepted proof is checked for its axiom dependencies, which catches indirect routes to a false proof. Correctness is decided by Lean, a free and open proof checker that we do not own and cannot influence. The full mechanism is on the method page.

Who this is for

A small, identifiable buyer universe

Frontier laboratories that need a measurement of their own models across model families, ahead of release, with results they can put in front of a regulator. Government AI safety institutes whose mandate is oversight and for whom a quarterly independent measurement is exactly what that mandate requires. Regulated enterprises and defence primes that deploy models they did not train and must evidence to their own regulator that the reasoning was tested independently.

Every one of them can be named. There is no long tail.

Certified

Run under the Antefacts protocol

Each quarterly holdout run is pre-registered and hash-committed before it happens, scored by a published harness, entered in a hash-chained record, and certified by Antefacts as having been conducted exactly as registered. Our customers receive a measurement certified by a party other than Invigil. Antefacts and Invigil are under common ownership; what that means and the rules that follow from it are on the certification page.

Why now

Three things arrived together

Frontier models became good enough at formal reasoning that the measurement is informative rather than uniformly zero. Regulation moved from principle to evidence, so labs now need third-party results they can show. And contamination became the dominant criticism of every public leaderboard, which makes an uncontaminated holdout worth paying for rather than merely interesting.

The name

Invigilate

To invigilate is to sit in the room while the exam is taken, so that the answers are the candidate's own. That is the job: a set of problems the model has never seen, a checker that cannot be argued with, and a record anyone can inspect.