Assisted release-evidence pilot
Give your reviewer evidence they can independently check.
One release-review workflow. Two real releases. Start with a supported native comparison or compatible retained records. For model suppliers, fine-tuning vendors and optimization teams helping a customer or internal reviewer decide whether to accept a changed model.
Is this pilot a fit for your team?
Start with an upcoming model change, qualifying existing records or inputs ready for a supported native evaluation, and a customer or internal reviewer willing to take part. The pilot fits teams that repeatedly prepare or review release evidence and want to make that handoff easier to repeat and check. We review a sanitized sample first to confirm whether your records can be used or further evaluation work is needed. Exact match needs answers and references; likelihood needs measured reference log probabilities; bounded judging needs frozen answers, a rubric and retained judge calls.
What you get
- A workflow connected to your evaluations
- Connect one evaluation path in your environment. Agree the model changes and review criteria it will cover, and record your current preparation and review effort.
- Evidence your reviewer can inspect
- Produce signed evidence and a comparison summary, with the inputs and decision criteria recorded so the result can be checked.
- A guided, independent review
- Check the evidence against independently approved inputs, policy and signer for the agreed workflow, then retain a signed receipt. Native comparisons bind artifacts and runtimes; captured comparisons bind supplied runs; judging also binds its plan and measurements. Distinguish a valid policy rejection from incomplete records, mismatched identities or unverifiable signatures.
- A repeatable handoff
- Repeat the process on a second real release. Hand over a command or CI workflow, a runbook and a recommendation on whether to continue.
Your reviewer can be within your team. Your organization decides whether one person may perform both roles, with the required approval and key controls. The expected values come from approved records outside the submitted evidence. How the trust controls work →
What your team contributes
Plan for a 60-minute scoping session, a technical owner for setup and troubleshooting, and a reviewer for two review sessions. Reserve roughly half a day per week across those roles; we confirm the estimate after reviewing your inputs.
Bring a sanitized sample of your records or planned evaluation inputs and explain how results are scored. We work through the supported path and remaining requirements together. Your team supplies compute and infrastructure unless agreed otherwise. Any new judge collection needs an agreed budget and your provider access.
What determines the fee
The lower end suits ready evaluation inputs, a supported integration and minimal customization. More integration work or reviewer support increases the fee. New adapters, substantial data cleanup or additional environments require a revised scope and quote.
This covers one workflow in one environment. Private-platform deployment and hosted model execution are outside the pilot; private deployment can be scoped separately. Compatible hosted-service records can be reviewed, but fresh service calls remain your harness’s responsibility. Verification does not refresh old observations or establish judge accuracy. The pilot does not provide safety or regulatory certification.
How the pilot proceeds
With compatible records and agreed scoring, the first comparison can happen early. Preparing those inputs and scheduling the second release often take longer than processing the comparison. We plan the engagement around your readiness and release schedule, with delivery effort, review sessions and support agreed in the scope.
The 4–6 weeks is a planned engagement window, not a runtime estimate or dedicated full-time engineering. If new evaluation records are needed, we agree how they will be collected before confirming timing.
Confirm readiness
Inspect a sanitized sample, confirm that your cases represent the intended workload, check scoring and comparison conditions, and agree the scope and success measures with your technical lead, reviewer and budget owner.
First comparison and review
Use ready records, or collect them through the agreed evaluation path, to produce the first comparison. Help your reviewer verify the evidence, investigate the result and record preparation and review effort.
Repeat and hand over
Repeat the workflow on a second real release, measure repeat-run timing and customer effort, and hand over the runbook and a recommendation on next steps.
What would count as success?
Your reviewer can independently verify evidence for two real releases, understand the results and repeat the handoff using the runbook. We measure preparation and review effort against your starting workflow, including setup and support. We also measure the full evaluation cycle on your workload: collecting or reusing records, generating and scoring answers where needed, comparing results and reviewing the evidence. This gives your team a basis for deciding how the workflow fits your release process. We agree the comparison and improvement target up front; outcomes and savings are not guaranteed.
A valid policy rejection means a checked comparison did not meet the agreed criteria. Incomplete or unverifiable records cannot support that checked result; more evidence or corrected inputs are needed. Completing the pilot does not require a passing verdict, and a pass does not authorize deployment. Your organization retains the release decision.
Discuss your next release
Tell us what is changing, how you evaluate it today, and who needs to review the result. We’ll check whether the pilot fits your workflow and agree the scope, timing and fee before work begins.
We work with your team in your environment and agree data access and handling before setup. Start with a description of your workflow; keep private data, model weights and keys out of the contact form.
Prefer to integrate it yourself? InvarLock is open source and available without a paid pilot. Get started with the documentation.