Back to blog

What Signed Case Studies Add Beyond Conformance Fixtures

A small test with known answers checks the software. A signed case study records an actual model comparison. A reviewer needs to know which question each answers.

Updated 6 min readInvarLock Team
A known reference block fits a measuring gauge, while a separate irregular specimen remains tied to its own inspection record.

A demo gets the expected score. That shows one part of the evaluation software behaved as expected. It does not tell a reviewer what happened when the proposed model change was tested.

A conformance fixture is a controlled test with an expected result. It helps catch software that handles an input incorrectly. A signed case study records a particular model comparison: which versions ran, which answers changed, how they were scored, and what remains uncertain. Signing links those records for inspection; it does not make the tested cases representative of every deployment.

InvarLock's saved examples make the distinction concrete. One small fixture contains an exact match and a mismatch. Each of the three published GGUF case studies contains 400 paired records and a separate signed decision trail. These examples answer different questions, not just questions of different sizes.

A fixture's score is part of the test

The saved two-record LM Evaluation Harness fixture reports scores of 1.0 and 0.0, a mean of 0.5, and the outcome qualified_for_import. Its files also identify the evaluator and link the result to the exact test inputs, software runner, and dependencies.

The mean was expected because the test was constructed that way. Calling it “50% model accuracy” would misread the evidence. The match and mismatch check scoring and conversion into a common record format. They do not measure how a proposed model performs on the recipient's workload.

Small fixtures are valuable precisely because their outcomes are known. A reviewer can check whether an implementation preserves a match, a mismatch, or an expected failure without running a model. A larger case study does not replace that focused software check.

Saved model answers let us check the scoring

Another example uses 102 saved Qwen3.5-0.8B answers: 61 exact matches and 41 mismatches. Different evaluators score those same answers. Another verifier can then recalculate the scores from the common records without rerunning the model.

This checks scoring against actual model outputs, rather than only invented test answers. It still does not provide a complete signed comparison between an original and changed model. The evaluator qualification reference explains these import checks and the separate requirements for a signed comparison.

Choose the evidence that answers the review question:

  • Example
    Known-answer test
    Records
    2
    What it checks
    Expected scores and links to the test inputs
    What it does not show
    Accuracy on a deployment workload
  • Example
    Saved-answer import
    Records
    102
    What it checks
    Scores recalculated from common records
    What it does not show
    A new model run or a complete signed comparison
  • Example
    Signed GGUF case
    Records
    400 pairs each
    What it checks
    One deployment comparison, its uncertainty and policy result
    What it does not show
    General quantization behavior or permission to deploy

These are complementary checks. Record count alone does not turn one kind of evidence into another.

A signed case preserves the comparison behind the result

The three GGUF deployment comparisons compare Qwen3.5 9B, Qwen3.8 27B, and Ministral 3 8B in BF16 and Q5_K_M formats. Each records the exact model files, execution software, ordered cases, and answers. Its report and signed evidence pack travel with a separately signed verification receipt.

All three pass the same declared exact-match policy on their 400-record schedules. All three paired intervals include zero, so their positive point estimates do not demonstrate improvement or establish equivalence. The reports also contain no latency, memory, throughput, energy, or cost measurements. Those limits belong beside the results even when verification succeeds.

The added value is that the evidence remains available for questions the headline score cannot answer. The September answer-change inspection used the Qwen3.5 records to separate regressions from improvements and output changes that left correctness unchanged. A fixture designed to return known scores would not supply those observations about that model change.

The records also show that the model format and execution software changed together. They therefore cannot identify quantization alone as the cause of an answer change. A reviewer can inspect that limitation directly instead of having to infer it from a demo.

Reuse the review method, then supply your own evidence

A useful case study gives the next reviewer a method to repeat: identify the change, record the exact inputs, save both sides' answers, state the scoring and acceptance rules, and keep the report and receipt. The public evidence guide explains how to preserve and verify those records.

An earlier result cannot stand in for the new comparison. A different model file, runtime, prompt format, workload, or policy needs evidence for that change. Even unchanged software can encounter failures absent from a small test set.

Start by asking what you need to check: expected software behavior, scoring of saved answers, or a particular model comparison. Then apply the recipient checks. The receiving organization supplies its own trusted expectations and decides whether the evidence is sufficient.

Replay can check the retained evidence and reconstruct its report without repeating inference. That is useful for independent review, but it does not prove that the original execution occurred or decide whether the recipient should deploy the change.

Limitations

  • The examples are selected retained fixtures, imports, and three text-only deployment comparisons. They are not a complete inventory of current evidence types or representative coverage of deployment workloads.
  • The two-record fixture and 102-record import test specific exact-match paths. They do not establish every evaluator capability or a new signed comparison.
  • The GGUF intervals include zero, and artifact representation and runtime change together. The cases establish neither improvement, equivalence, causal attribution, nor an unmeasured performance benefit.
  • Signatures and replay preserve a bounded review trail. Genuine original execution, wider model quality, safety, recipient acceptance, and deployment approval require other evidence or authority.

Sources

Website documentation explains the maintained workflow. The conformance fixture and shared-output import were inspected as retained examples in InvarLock v0.16.3, commit 2b2b64d2f668ab4768ba044df7883f86e5e0a835. The three GGUF cases retain their original v0.15.0 evidence basis, commit 4f036ffe99cfaa82019c90ea527b8c4291a091b3; their evidence directories and receipts are unchanged in v0.16.3. This synthesis inspects existing artifacts and published analyses; it does not report new inference or a new replay run.

More in Public Evidence

Explore nearby related posts.