Back to blog

InvarLock 0.17: Plan the Comparison, Explain the Result

Check judge-evaluation precision before collection, explore exact-match sensitivity, and rehearse reassessment while preserving earlier evidence.

3 min readInvarLock Team
A fan of possible paths is measured against a narrow opening before any path is chosen.

Before collecting more evaluation results, a team needs to know whether its plan can answer the review question. After the comparison, it needs to explain a borderline result and preserve what was originally assessed if new evidence arrives.

InvarLock 0.17 adds tools and worked examples for those stages. The existing evaluation, independent verification and reporting workflow remains in place. These additions help plan and interpret a comparison; they do not extend a technical verdict into approval to deploy.

Check precision before collecting judge ratings

Judge preflight now forecasts conservative interval-width bounds from the declared rating scale, independent units and statistical settings. It distinguishes plans that meet the width requirement for every possible score outcome, plans that cannot meet it, and plans without that guarantee.

The forecast assumes the planned trials finish and the units are independent. Repeating ratings within one unit does not create more independent units. Meeting a precision requirement also does not predict a passing comparison or establish judge accuracy. Use the judge preflight reference before committing collection effort.

Explain sensitivity without rewriting the result

For supported native exact-match comparisons, the Python sensitivity helper explores hypothetical changes to candidate outcomes. Its bounded search can return an exact minimum number of changes with an example, a lower bound from an incomplete search, or a result showing that no flip is possible under the fixed edit model.

This is an explanation under the selected policy. It does not edit the signed evidence, change the official verdict or estimate how likely a future rerun is to differ. The decision semantics guide covers its scope. A separate score-contribution recipe helps inspect additive score changes while preserving the declared unit weights.

Preserve the earlier assessment when evidence changes

The hosted-service guide now includes an offline example of a later assessment and a linked correction of a source-mapping error. Both retain the original signed evidence and verification receipt alongside the replacement.

These are synthetic integration examples, not a fresh service-quality study. A later rejection does not erase a historical pass, and preserving an erroneous assessment does not justify continued reliance on it. The recipient still reviews affected decisions; the example does not automatically revoke approvals.

Upgrade and check your workflow

Install the release with Python 3.12 or newer:

python -m pip install "invarlock==0.17.0"

Use the matching CPU quickstart to verify retained evidence and produce a report without running a model. Optional dependencies use the same distribution's diagnostics, vision-text and judge extras.

The release also improves failure explanations and optional numerical diagnostics. Diagnostics now cap the intermediate covariance allocation at 128 MiB by default; this is not a total process-memory limit. Wide inputs can use the smaller-Gram method. Diagnostics remain advisory and do not determine the policy verdict. The release record includes dependency and security updates.

For an assisted comparison, the paid assessment starts with one model change, agreed requirements and a fixed quote. Planning, interpretation and any further assessment are scoped around that workflow.

Sources

Website documentation explains the maintained workflow. This article describes v0.17.0; the tagged records preserve its release scope. Existing report examples retain the versions and source records that produced them.

More in Release

Explore nearby related posts.