Model release assurance
Verifiable evidence for model releases.
Compare a changed model with its baseline and share signed evidence your customer or internal reviewer can check. Run a supported comparison or use compatible results from your existing evaluator.
Can we use the smaller model within our allowed quality loss?
Ministral 3 8B · BF16 → Q5_K_M GGUF
400 paired MMLU-Pro records
Policy satisfied
The quantized model meets the stated requirements on this benchmark.
Exact-match accuracy
- Baseline
- 44.75%
- Candidate
- 45.25%
+0.50 percentage points
95% interval: −1.12 to +2.12 percentage points.
Open this comparison’s reportEach example shows a specific model change checked against defined requirements. Your reviewer makes the release decision. Understand the results →
The handoff
From evaluation to release review
Bring supported evaluation records, or run a comparison with InvarLock. Share a signed evidence package and report. Your reviewer verifies the package against approved requirements and keeps a signed verification receipt.
You prepare the evidence
Choose the baseline, test cases and requirements. Run a supported comparison or import compatible results, then save the signed evidence package.
Your reviewer verifies
They verify the package’s signature, repeat the recorded checks and keep a signed receipt of the result.
Your reviewer decides
They decide whether the result is sufficient for the intended use and meets their organization’s release rules.
The reviewer uses approved inputs and requirements obtained separately from the submitted package. The review can happen within your team or with a customer.
Integration
Fit the review to your workflow
Start with the model change you need to review: a fine-tune, a quantized model or a prompt change. Choose test cases and acceptance requirements that reflect its intended use.
Plan your first comparisonStart with existing records
Bring compatible results for individual test cases, with the inputs and scoring measurements your workflow requires. A summary score alone is not enough. Replay checks those records without rerunning the model.
Use existing evaluation records →Or run a new comparison
Run a supported model comparison in your own environment. Your team supplies the model files, test cases and compute; InvarLock packages the results for review.
Runtime and metric support →Work with us
Bring a release that needs review.
Set up one evaluation workflow and use it across two real releases. The paid pilot includes integration help and reviewer support, with the effort to prepare and review evidence measured along the way.
See what you get, what your team contributes, and the planned timeline and budget.
Explore the pilot