Three GGUF Deployment Comparisons Under One Review Standard
Three BF16-to-Q5_K_M comparisons pass the same declared policy, each with its own evidence and receipt. The repeatable result is the review standard.
A reviewer considering a BF16-to-GGUF deployment change needs more than a favorable score difference. The review must identify the exact artifacts and runtimes, preserve which answers changed, and explain what counted as an acceptable result.
The three retained deployment comparisons show that review standard applied repeatedly. Qwen3.5 9B, Qwen3.8 27B, and Ministral 3 8B each pass the same declared exact-match policy, with a separate evidence pack and independently signed verification receipt. The useful common result is a repeatable review contract. These transactions do not combine into a model ranking or a general finding about quantization.
Three results, each with its own uncertainty
Each row compares a BF16 checkpoint with its source-derived Q5_K_M GGUF deployment on 400 paired, balanced MMLU-Pro records. BF16 and Q5_K_M columns show exact-match accuracy. The effect is subject minus baseline; pp means percentage points.
- Comparison
- Qwen3.5 9B
- BF16
- 53.00%
- Q5_K_M
- 54.75%
- Paired effect
- +1.75 pp
- Paired 95% interval
- [−0.83, 4.32] pp
- Policy
- Pass
- Comparison
- Qwen3.8 27B
- BF16
- 63.00%
- Q5_K_M
- 64.50%
- Paired effect
- +1.50 pp
- Paired 95% interval
- [−0.74, 3.75] pp
- Policy
- Pass
- Comparison
- Ministral 3 8B
- BF16
- 44.75%
- Q5_K_M
- 45.25%
- Paired effect
- +0.50 pp
- Paired 95% interval
- [−1.12, 2.12] pp
- Policy
- Pass
Values come from the three tagged canonical reports, with effects and interval endpoints rounded to two decimals. All three intervals include zero. The positive point estimates therefore do not demonstrate improvement, and a policy pass does not establish equivalence.
The reports measure text exact match for these particular deployment changes. They contain no latency, memory, throughput, energy, or cost measurement from which to infer a deployment benefit.
A shared policy makes the passes interpretable
All three reports use invarlock/comparison-report-v3 and the same four controls:
| Control | Declared requirement |
|---|---|
| Paired record count | At least 400 |
| Paired 95% interval width | At most 10 percentage points |
| Subject-minus-baseline lower interval bound | At least −2 percentage points |
| Accuracy on each side | At least 0.20, or 20%, for both baseline and subject |
Every lower endpoint is above the −2 point floor. The interval widths are 5.15, 4.49, and 3.24 percentage points respectively, all below 10. Each row has 400 records, and all six side accuracies exceed 20%. Each report consequently records a pass.
The side-accuracy check matters because a small relative difference can conceal two poorly scoring sides. Requiring both sides to clear a floor prevents the delta alone from deciding policy. The 20% threshold belongs to this declared policy; it is not a general capability or safety threshold. The report contract explains how these checks combine.
Using the same controls makes the meaning of “pass” consistent across the three reviews. It does not turn three curated transactions into a representative sample of models, tasks, or quantization behavior.
Shared questions do not mean identical input bytes
The retained schedules use the same ordered selection of 400 semantic question IDs. That selection is balanced across 14 MMLU-Pro domains and answer choices A–J. Each model renders those questions through its pinned chat format, so the rendered inputs and full schedule digests are not identical across all three transactions.
Within each transaction, baseline and subject face the same ordered inputs. That is the pairing required to preserve which answers changed. Across transactions, a shared question selection does not erase differences in artifacts, prompt rendering, conversion, or runtime. The tagged deployment journey describes the text-only execution boundaries for each profile; this comparison does not evaluate their vision capabilities.
These distinctions keep the review reproducible without treating the rows as interchangeable. The pairing and replay contract describes what must remain bound when a result is rechecked.
Identity travels with each result
“The quantized model” is too loose an identity for a release review. Each transaction retains its own baseline and subject artifact identities, separate runtime identities, normalized request, conversion provenance, ordered records, metric, and policy. Those objects answer different questions: which bytes were compared, how they were run, what inputs they received, how they were scored, and which rule produced the verdict.
Changing the conversion toolchain, artifact, runtime, prompt format, or policy changes the review question. An old receipt cannot silently authorize that new transaction. Ministral adds an independent model family to the retained examples, but its text-only comparison remains a bounded canary outside the evaluator qualification matrix.
Each pack has its own signed inventory and retained verification receipt. Following the public evidence workflow, a reviewer supplies independent artifact, schedule, policy, runtime, request, and signer anchors. Verification can then reconstruct the comparison and issue a fresh receipt without rerunning model inference or conversion.
For this article, all three packs were replayed with the published 0.15.0 wheel using the release's separately retained reference anchors. The three replays passed, their fresh signed receipts validated, and repeated HTML renderings matched byte for byte. That is a check of retained evidence and report arithmetic. It does not independently prove genuine original execution, correctness, representativeness, or safety.
The practical pattern is to reuse the review standard while retaining a complete decision trail for every change. A technical policy pass remains separate from recipient acceptance and organizational deployment approval. Those decisions need their own authority and evidence.
Limitations
- These are three curated, text-only deployment comparisons. They do not establish a model ranking, vision capability, or general quantization behavior.
- All three intervals include zero; the results do not demonstrate improvement or establish equivalence. The reports contain no latency, memory, throughput, energy, or cost measurements.
- Replay checks retained evidence and report arithmetic. It does not rerun inference or conversion, or independently establish genuine original execution.
- Policy passes remain bound to their transactions. Safety, recipient acceptance, and organizational deployment approval require separate evidence or authority.
Sources
Website documentation explains the maintained workflow. The comparisons and receipts described here use InvarLock v0.15.0; the tagged index, reports, and receipts preserve that evidence basis. Replay uses release-maintained reference inputs held separately from the submitted packs; recipients still choose which identities and policies to trust.
- Public evidence guide
- Reports and receipts: side-accuracy qualification
- Pairing and replay
- Interactive evidence page
- Tagged v0.15.0 public evidence index
- Qwen3.5 9B canonical report
- Qwen3.5 9B retained verification receipt
- Qwen3.5 9B evidence carrier
- Qwen3.8 27B canonical report
- Qwen3.8 27B retained verification receipt
- Qwen3.8 27B evidence carrier
- Ministral 3 8B canonical report
- Ministral 3 8B retained verification receipt
- Ministral 3 8B evidence carrier
- Tagged GGUF deployment journey
- Tagged release reference inputs
- Tagged replay implementation
More in Research Note
Explore nearby related posts.
Research note
Which Answers Changed Behind the Score?
One retained comparison has 54 changed answers, 27 changed correctness outcomes, and seven net additional correct answers. Each count answers a different review question.
Research note
A Passing Release Test Can Contain Policy Rejections
In the retained preflight example, all eleven references must reproduce their expected outcomes: eight policy passes and three integrity-valid policy rejections.
Research note
The Proof Stops at the Signed Schedule
What a verified comparison establishes depends on its exact records, declared contract, and independently supplied trust inputs.