Model release assurance

Verifiable evidence for model releases.

Compare a changed model with its baseline and share signed evidence your customer or internal reviewer can check. Run a supported comparison or use compatible results from your existing evaluator.

Explore a release decisionRecorded results

Can we use the smaller model within our allowed quality loss?

Ministral 3 8B · BF16 → Q5_K_M GGUF

400 paired MMLU-Pro records

Policy satisfied

The quantized model meets the stated requirements on this benchmark.

Exact-match accuracy

Baseline
44.75%
Candidate
45.25%

+0.50 percentage points

95% interval: −1.12 to +2.12 percentage points.

Open this comparison’s report
Inspect the sources and replay this comparison →

Each example shows a specific model change checked against defined requirements. Your reviewer makes the release decision. Understand the results →

The handoff

From evaluation to release review

Bring supported evaluation records, or run a comparison with InvarLock. Share a signed evidence package and report. Your reviewer verifies the package against approved requirements and keeps a signed verification receipt.

  1. You prepare the evidence

    Choose the baseline, test cases and requirements. Run a supported comparison or import compatible results, then save the signed evidence package.

  2. Your reviewer verifies

    They verify the package’s signature, repeat the recorded checks and keep a signed receipt of the result.

  3. Your reviewer decides

    They decide whether the result is sufficient for the intended use and meets their organization’s release rules.

The reviewer uses approved inputs and requirements obtained separately from the submitted package. The review can happen within your team or with a customer.

Integration

Fit the review to your workflow

Start with the model change you need to review: a fine-tune, a quantized model or a prompt change. Choose test cases and acceptance requirements that reflect its intended use.

Plan your first comparison

Start with existing records

Bring compatible results for individual test cases, with the inputs and scoring measurements your workflow requires. A summary score alone is not enough. Replay checks those records without rerunning the model.

Use existing evaluation records →

Or run a new comparison

Run a supported model comparison in your own environment. Your team supplies the model files, test cases and compute; InvarLock packages the results for review.

Runtime and metric support →

Work with us

Bring a release that needs review.

Set up one evaluation workflow and use it across two real releases. The paid pilot includes integration help and reviewer support, with the effort to prepare and review evidence measured along the way.

See what you get, what your team contributes, and the planned timeline and budget.

Explore the pilot