What Signed Case Studies Add Beyond Conformance Fixtures
A small test with known answers checks the software. A signed case study records an actual model comparison. A reviewer needs to know which question each answers.
A demo gets the expected score. That shows one part of the evaluation software behaved as expected. It does not tell a reviewer what happened when the proposed model change was tested.
A conformance fixture is a controlled test with an expected result. It helps catch software that handles an input incorrectly. A signed case study records a particular model comparison: which versions ran, which answers changed, how they were scored, and what remains uncertain. Signing links those records for inspection; it does not make the tested cases representative of every deployment.
InvarLock's saved examples make the distinction concrete. One small fixture contains an exact match and a mismatch. Each of the three published GGUF case studies contains 400 paired records and a separate signed decision trail. These examples answer different questions, not just questions of different sizes.
A fixture's score is part of the test
The saved two-record LM Evaluation Harness fixture reports scores of 1.0 and 0.0, a mean of 0.5, and the outcome qualified_for_import. Its files also identify the evaluator and link the result to the exact test inputs, software runner, and dependencies.
The mean was expected because the test was constructed that way. Calling it “50% model accuracy” would misread the evidence. The match and mismatch check scoring and conversion into a common record format. They do not measure how a proposed model performs on the recipient's workload.
Small fixtures are valuable precisely because their outcomes are known. A reviewer can check whether an implementation preserves a match, a mismatch, or an expected failure without running a model. A larger case study does not replace that focused software check.
Saved model answers let us check the scoring
Another example uses 102 saved Qwen3.5-0.8B answers: 61 exact matches and 41 mismatches. Different evaluators score those same answers. Another verifier can then recalculate the scores from the common records without rerunning the model.
This checks scoring against actual model outputs, rather than only invented test answers. It still does not provide a complete signed comparison between an original and changed model. The evaluator qualification reference explains these import checks and the separate requirements for a signed comparison.
Choose the evidence that answers the review question:
- Example
- Known-answer test
- Records
- 2
- What it checks
- Expected scores and links to the test inputs
- What it does not show
- Accuracy on a deployment workload
- Example
- Saved-answer import
- Records
- 102
- What it checks
- Scores recalculated from common records
- What it does not show
- A new model run or a complete signed comparison
- Example
- Signed GGUF case
- Records
- 400 pairs each
- What it checks
- One deployment comparison, its uncertainty and policy result
- What it does not show
- General quantization behavior or permission to deploy
These are complementary checks. Record count alone does not turn one kind of evidence into another.
A signed case preserves the comparison behind the result
The three GGUF deployment comparisons compare Qwen3.5 9B, Qwen3.8 27B, and Ministral 3 8B in BF16 and Q5_K_M formats. Each records the exact model files, execution software, ordered cases, and answers. Its report and signed evidence pack travel with a separately signed verification receipt.
All three pass the same declared exact-match policy on their 400-record schedules. All three paired intervals include zero, so their positive point estimates do not demonstrate improvement or establish equivalence. The reports also contain no latency, memory, throughput, energy, or cost measurements. Those limits belong beside the results even when verification succeeds.
The added value is that the evidence remains available for questions the headline score cannot answer. The September answer-change inspection used the Qwen3.5 records to separate regressions from improvements and output changes that left correctness unchanged. A fixture designed to return known scores would not supply those observations about that model change.
The records also show that the model format and execution software changed together. They therefore cannot identify quantization alone as the cause of an answer change. A reviewer can inspect that limitation directly instead of having to infer it from a demo.
Reuse the review method, then supply your own evidence
A useful case study gives the next reviewer a method to repeat: identify the change, record the exact inputs, save both sides' answers, state the scoring and acceptance rules, and keep the report and receipt. The public evidence guide explains how to preserve and verify those records.
An earlier result cannot stand in for the new comparison. A different model file, runtime, prompt format, workload, or policy needs evidence for that change. Even unchanged software can encounter failures absent from a small test set.
Start by asking what you need to check: expected software behavior, scoring of saved answers, or a particular model comparison. Then apply the recipient checks. The receiving organization supplies its own trusted expectations and decides whether the evidence is sufficient.
Replay can check the retained evidence and reconstruct its report without repeating inference. That is useful for independent review, but it does not prove that the original execution occurred or decide whether the recipient should deploy the change.
Limitations
- The examples are selected retained fixtures, imports, and three text-only deployment comparisons. They are not a complete inventory of current evidence types or representative coverage of deployment workloads.
- The two-record fixture and 102-record import test specific exact-match paths. They do not establish every evaluator capability or a new signed comparison.
- The GGUF intervals include zero, and artifact representation and runtime change together. The cases establish neither improvement, equivalence, causal attribution, nor an unmeasured performance benefit.
- Signatures and replay preserve a bounded review trail. Genuine original execution, wider model quality, safety, recipient acceptance, and deployment approval require other evidence or authority.
Sources
Website documentation explains the maintained workflow. The conformance fixture and shared-output import were inspected as retained examples in InvarLock v0.16.3, commit 2b2b64d2f668ab4768ba044df7883f86e5e0a835. The three GGUF cases retain their original v0.15.0 evidence basis, commit 4f036ffe99cfaa82019c90ea527b8c4291a091b3; their evidence directories and receipts are unchanged in v0.16.3. This synthesis inspects existing artifacts and published analyses; it does not report new inference or a new replay run.
- Evaluator qualification
- Public evidence guide
- Reports and receipts
- Pairing and replay
- Three GGUF Deployment Comparisons Under One Review Standard
- Which Answers Changed Behind the Score?
- Five Recipient Checks Before a Technical Result Can Be Accepted
- Retained two-record qualification schedule
- Retained LM Evaluation Harness upstream output
- Retained LM Evaluation Harness export
- Retained LM Evaluation Harness qualification result
- Retained shared-output import guide
- Retained shared-output corpus
- Original GGUF deployment journey
- Original public evidence index
More in Public Evidence
Explore nearby related posts.
Research note
Five Recipient Checks Before a Technical Result Can Be Accepted
Technical verification is an input to acceptance, not the decision itself. Five checks keep identity, trust, policy, uncertainty, and organizational authority separate.
Research note
Which Answers Changed Behind the Score?
One retained comparison has 54 changed answers, 27 changed correctness outcomes, and seven net additional correct answers. Each count answers a different review question.
Research note
Three GGUF Deployment Comparisons Under One Review Standard
Three BF16-to-Q5_K_M comparisons pass the same declared policy, each with its own evidence and receipt. The repeatable result is the review standard.