Release note
InvarLock 0.16.0: Native, Captured and Judge Evidence
InvarLock 0.16 brings captured evaluator results and bounded judge measurements into the shared review workflow, with clearer reports and explicit verification limits.
Blog
Engineering notes on model evaluation, independent verification and the choices behind InvarLock.
54 posts across 22 topics. Latest: September 16, 2026
Release note
InvarLock 0.16 brings captured evaluator results and bounded judge measurements into the shared review workflow, with clearer reports and explicit verification limits.
7 min read
Research note
One retained comparison has 54 changed answers, 27 changed correctness outcomes, and seven net additional correct answers. Each count answers a different review question.
Read more6 min read
Research note
Three BF16-to-Q5_K_M comparisons pass the same declared policy, each with its own evidence and receipt. The repeatable result is the review standard.
Read more5 min read
Research note
In the retained preflight example, all eleven references must reproduce their expected outcomes: eight policy passes and three integrity-valid policy rejections.
Read more5 min read
Research note
What a verified comparison establishes depends on its exact records, declared contract, and independently supplied trust inputs.
Read more5 min read
Research note
A signed evaluator output becomes decision evidence only when its records can be independently replayed under an authorized contract.
Read more8 min read
Release note
InvarLock 0.15 adds replayable deployment evidence, an absolute accuracy floor, four retained evaluator journeys, and a protected consumer example.
Read more6 min read
Research note
Four retained evidence packs passed four different tests. Their metrics, intervals, runtimes, and policies show exactly how far each result reaches.
Read more6 min read
Research note
Evaluator support has three independent meanings: a maintained adapter, replayable records, and a retained signed journey. None implies the other two.
Read more8 min read
Release note
v0.14 normalizes external evaluator records, then hands a signed technical receipt to recipient-owned trust and acceptance rules.
Read more9 min read
Release note
One YAML request now binds artifacts, schedule, metric, and policy. One signed evidence pack carries that comparison through independent verification and reporting.
Read more8 min read
Paper note
A 340-run study separates three release gates: artifact readiness, evidence saturation, and host parity. Each answers a question that a valid evaluation report leaves open.
Read more3 min read
Research Note
Structural checks matter because they prevent obviously broken edits from reaching evaluation. They do not, by themselves, show that an edit preserved quality.
Read more2 min read
Release
InvarLock 0.12.1 tightens the Mistral guard-value evidence contract, promotes targeted spectral/RMT/VE probes to required detections, and replaces the attention control with a stock-cap-clean 1.05x lane.
Read more2 min read
Release
InvarLock 0.12.0 adds knowledge/self-edit workflow metadata, widens BYOE edit-lane evidence, and reorganizes public evidence around backend compatibility and larger model findings.
Read more4 min read
Evidence Note
Even when an evidence pack is strict, signed, and marked PASS, it is still a portable verification bundle. It does not answer every scientific, safety, or deployment question a reader might ask.
Read more5 min read
Research Note
Archiving a model-edit decision is not about saving more files. It is about preserving the exact bundle another reviewer would need to re-check the result later.
Read more3 min read
Release
InvarLock 0.11.0 makes human reports more consistent, separates guard warnings from hard failures, and expands the public evidence surface across model families.
Read more5 min read
Research Note
A screenshot can communicate a result. An evidence pack can be inspected, checked, and re-verified later. That is the difference between presentation and portable evidence.
Read more4 min read
Research Note
A strong evaluation result should carry its runtime provenance with it. In InvarLock, that means the runtime manifest travels next to the report and is rechecked by invarlock verify.
Read more2 min read
Release
InvarLock 0.10.0 makes public evidence packs more explicit, adds signer-authenticity checks, and expands optional quantized-subject adapter validation.
Read more4 min read
Research Note
An evaluation report is strongest when it is treated as a stable evidence contract: a small required core, meaningful optional blocks, and a clear boundary around what still lives outside the JSON.
Read more5 min read
Research Note
Calibration is not just analysis around the product. It changes how thresholds are derived, when correction paths may turn on, and which policy values later govern reports.
Read more2 min read
Release
InvarLock 0.9.0 adds strict assurance mode, fail-closed verifier checks, runtime provenance guidance, and maintainer evidence gates for release review.
Read more5 min read
Research Note
Calibration becomes operational when sweep artifacts end in reviewable YAML patches that later appear as resolved runtime policy in reports.
Read more5 min read
Research Note
Variance equalization is stronger when it must earn enablement through predictive evidence, explicit tier knobs, and report-visible provenance.
Read more5 min read
Research Note
Thresholds are stronger when they come from measured null behavior and end in a policy patch, not from knob-tuning folklore.
Read more4 min read
Research Note
A trustworthy weight-edit result needs more than a benchmark delta. It needs a bounded claim, an exactly paired comparison, and verification that rejects incomplete evidence.
Read more2 min read
Release
InvarLock 0.8.0 moves the public bundle surface to evidence packs, pins docs to versioned release paths, and makes container-vs-host runtime provenance explicit across evaluate and verify.
Read more5 min read
Research Note
A verifier is only useful if it rejects incomplete evidence. InvarLock's verification path is designed to stop stronger claims when the evidence bundle is missing or inconsistent.
Read more2 min read
Release
InvarLock 0.7.2 simplifies the public release surface around immutable source tags plus the PyPI wheel and sdist, with docs and verification gates aligned around that path.
Read more5 min read
Research Note
A model-edit benchmark number is only as strong as the comparison behind it. Pairing makes the comparison inspectable.
Read more2 min read
Release
InvarLock 0.7.1 makes wheel-only verify/report workflows first-class, ships a public contract bundle, and tightens supply-chain and release-validation gates.
Read more2 min read
Release
InvarLock 0.7.0 adds first-class GPT-OSS support, pilot Ministral 3 8B/14B presets, and a CUDA-capable attested runtime path for GPU hosts.
Read more6 min read
Research Note
A narrow claim can be stronger than a broad one. InvarLock is about auditable regression risk from weight edits, not general model safety.
Read more2 min read
Release
InvarLock 0.6.0 adds a shipped Gemma 4 E2B text lane, phase-1 multimodal evaluation, and a unified `--assurance attested|trusted-local` workflow.
Read more2 min read
Release
InvarLock 0.5.1 adds a push-gated tiny attested smoke lane, a scheduled GPT-2 canary lane, and package-native Ed25519 evidence pack signatures.
Read more2 min read
Release
InvarLock 0.5.0 adds offline release-verification bundles, package-native evidence-pack verification, and a simplified public CLI centered on evaluate, verify, and report.
Read more2 min read
Release
InvarLock 0.4.0 stabilizes contracts around policies, evidence packs, and evaluation provenance while tightening verification, CI, and coverage enforcement.
Read more2 min read
Release
Split-module coverage thresholds now protect critical CLI/reporting paths while config, plugin, report, overhead, and observability edge cases fail closed more reliably.
Read more2 min read
Release
A focused hardening release: safer AWQ plugin discovery, stronger quantization clipping behavior, and broader report-schema acceptance for edge payloads.
Read more2 min read
Release
Evidence packs add new showcase and evidence artifacts, while CI and release flows become more deterministic and easier to validate repeatedly.
Read more2 min read
Release
A stability-focused release: cleaner report output, safer offline evidence-pack flows, and CI/test hardening after the report rename.
Read more2 min read
Release
A terminology reset (report/evaluate), stricter evidence-pack verification, and a clean upgrade path for Hugging Face Transformers v5.
Read more2 min read
Release
Adapters move to role-based routing, evidence packs become easier to inspect (v2 layout), and reporting output gets a readability upgrade.
Read more2 min read
Release
Reports now record and enforce estimator measurement contracts under CI/release profiles, and evidence pack suites can cleanly split calibration vs execution.
Read more2 min read
Release
Evidence packs gain a deterministic bash test suite and better runtime helpers, window selection becomes stable/offline, and perplexity runs get safer around bad token IDs.
Read more1 min read
Release
CI/release baseline pairing is fail-closed (pairing evidence is required), and adapters reduce peak memory during retries via chunked snapshot/restore.
Read more1 min read
Release
Token-weighted paired bootstrap lands across the pipeline, strictness toggles expand, and CI/release pairing expectations become explicit and enforceable.
Read more1 min read
Release
`invarlock calibrate` arrives, determinism utilities mature, and regression harness + golden tracking help prevent silent policy drift.
Read more2 min read
Release
Fixes a GPU memory leak during reload fallback, hardens B200 scripts, and adds practical controls for acceptance ranges and overhead measurement.
Read more1 min read
Release
First-class quantization metadata, safer device movement across quantized models, auto-routing based on checkpoint info, and major test coverage expansion.
Read more1 min read
Release
The initial public release on GitHub and PyPI: core evaluate pipeline, guard chain, schema v1, and the first docs/CLI surface.
Read more3 min read
Announcement
The original InvarLock introduction, with a current starting point for native comparisons, captured results and bounded judging.
Read more