Recorded policy result
Policy satisfied
Every configured check passed for this paired evaluation. Independent recipient acceptance is a separate step.
What was compared
| Recorded field | Baseline | Subject |
|---|---|---|
| ArtifactDiffers | hf://mistralai/Ministral-3-8B-Instruct-2512-BF16@f6fae9795746f63c9be8344932f01275f3c63734 | derived://mistralai/Ministral-3-8B-Instruct-2512-BF16@f6fae9795746f63c9be8344932f01275f3c63734#llama.cpp-b10015-q5_k_m@sha256:b2c56b655b15577c30734b28b820d054e43df06c7712b74b9c170bff5f10f253 |
| Model IDDiffers | mistralai/Ministral-3-8B-Instruct-2512-BF16 | gguf-sha256-b2c56b655b15577c30734b28b820d054e43df06c7712b74b9c170bff5f10f253.gguf |
| ProviderDiffers | hf_transformers | llama_cpp |
- Baseline and candidate have different authenticated artifact digests; the evidence does not identify a transformation procedure.
Results and requirements
Exact-match accuracy
Scope: All paired records
All configured checks passed.
- Baseline
- 44.75%
- Subject
- 45.25%
- Change
- +0.5 ppPaired 95% CI: -1.12 to +2.119 pp
- Observed pairs
- 400
| Check | Observed | Required | Result |
|---|---|---|---|
| Paired lower boundThe lower bound must meet the allowed change, not just the point estimate. | -1.11994 pp | >= -2 pp | Passed |
| Record countThe comparison must include the configured minimum number of paired records. | 400 | >= 400 | Passed |
| Interval widthThe interval must be narrow enough for the configured precision requirement. | 3.23886 pp | <= 10 pp | Passed |
| Baseline accuracyEach side must meet the configured absolute accuracy floor. | 44.75% | >= 20% | Passed |
| Candidate accuracyEach side must meet the configured absolute accuracy floor. | 45.25% | >= 20% | Passed |
- Higher scores are better; the change is candidate minus baseline in percentage points.
- The policy tests the interval bound, not just the observed change.
- Authenticated observations are supplementary; the paired metric and policy remain the complete acceptance calculation.
Scope and limitations
- This report summarizes the evidence bundle. Rendering checks its integrity and embedded signature; independent acceptance requires verification with your own trust inputs.
- Results apply to the recorded cases, metric and policy. A pass does not establish general model quality, safety or representative production performance.
- The interval describes paired binary outcomes under its stated method. Population interpretation requires an appropriate sampling design.
Evidence details
Identities, methods and exact recorded values
Comparison and policy identities
- Comparison
comparison-4472c601a80f319859c4bf0c505c399c- Metric
exact_match- Policy
sha256:48e929519390c41e91c5e907e7e3f3d6d0b26e5ac4345136c4f6b85244b191d9- Evidence signer
sha256:e1f2aab12ee457abad9b0c74b1b9dcd0ce0429e541c15f17b814071b7fd9abab
Paired outcome analysis
{
"baseline_fail_subject_pass": 6,
"baseline_pass_subject_fail": 4,
"both_fail": 215,
"both_pass": 175,
"discordant_pairs": 10,
"effect_size_confidence_interval": {
"confidence_level": 0.95,
"lower_pp": -1.119943249438997,
"method": "newcombe_hybrid_score_paired_v2",
"upper_pp": 2.118915687037841
},
"effect_size_pp": 0.5,
"mcnemar_exact_two_sided_p_value": 0.75390625
}Authenticated observations
[
{
"authority": "observation",
"bindings": {
"artifact_digests": {
"baseline": "sha256:2bba1d6c1331980047b18d9b00da21d997dc9ca26a183434e512bfc348b2e891",
"subject": "sha256:fa3ab58b49a8e08b4901a3f990c0d0a401ec7930bd90a632f564ff7f71f1b1a6"
},
"comparison_id": "comparison-4472c601a80f319859c4bf0c505c399c",
"policy_digest": "sha256:48e929519390c41e91c5e907e7e3f3d6d0b26e5ac4345136c4f6b85244b191d9",
"schedule_digest": "sha256:9ac24cf5e4b904197117fa8d0d9661d88fdf0dae1cc39ddfa711aa820de60fad"
},
"format": "invarlock/evidence-observation-v1",
"kind": "artifact_transformation",
"observation_id": "ministral3-8b-bf16-to-gguf-q5-k-m",
"payload": {
"conversion": {
"output": {
"byte_length": 16987563264,
"filename": "Ministral-3-8B-Instruct-BF16.gguf",
"sha256": "b6dc36267813cef3a367b3e23528f5e367db5713ca6f2dd1f1a40a3eb1344ada",
"type": "BF16"
},
"runtime_image_digest": "sha256:b905cf3b278531333d47fc80ba5b110793f80717fc2acc526e8dd291ae8f3301",
"tool": {
"name": "llama.cpp/convert_hf_to_gguf.py",
"source_commit": "12127defda4f41b7679cb2477a4b0d65ee6a0c8f",
"source_sha256": "5ab75e394f4c71425ecce64a213dab3b8e3e9cfe0f19d0dcda4d5a4f7733da83",
"source_tag": "b10015"
}
},
"format": "invarlock/example-gguf-deployment-transformation-v1",
"quantization": {
"runtime_image_digest": "sha256:aa5d7b75a807e786175e9edf9927cbbc0bb3099bcc2d95fe07e56f80bbfeedb6",
"tool": "llama-quantize",
"type": "Q5_K_M"
},
"source": {
"checkpoint_tree_sha256": "sha256:6cbddcebc289550569cc3f6a93676a8f4f605d8574b8aec8448d61594a283996",
"repository": "mistralai/Ministral-3-8B-Instruct-2512-BF16",
"revision": "f6fae9795746f63c9be8344932f01275f3c63734",
"tokenizer_contract_sha256": "b9e3906504b6235b5c289fe9d3f7a86512f968dfe300771ec660080924615dbc"
},
"subject": {
"byte_length": 6058747136,
"filename": "Ministral-3-8B-Instruct-Q5_K_M.gguf",
"sha256": "b2c56b655b15577c30734b28b820d054e43df06c7712b74b9c170bff5f10f253"
}
},
"scope": "subject"
}
]Exact comparison data
{
"baseline": {
"mean_score": 0.4475
},
"comparison": {
"kind": "exact_match_delta_pp",
"minimum": -2.0,
"value": 0.5
},
"comparison_id": "comparison-4472c601a80f319859c4bf0c505c399c",
"format": "invarlock/comparison-report-v3",
"metric": "exact_match",
"paired_binary": {
"baseline_fail_subject_pass": 6,
"baseline_pass_subject_fail": 4,
"both_fail": 215,
"both_pass": 175,
"discordant_pairs": 10,
"effect_size_confidence_interval": {
"confidence_level": 0.95,
"lower_pp": -1.119943249438997,
"method": "newcombe_hybrid_score_paired_v2",
"upper_pp": 2.118915687037841
},
"effect_size_pp": 0.5,
"mcnemar_exact_two_sided_p_value": 0.75390625
},
"policy_digest": "sha256:48e929519390c41e91c5e907e7e3f3d6d0b26e5ac4345136c4f6b85244b191d9",
"record_count": 400,
"sample_qualification": {
"interval_width": {
"maximum": 10.0,
"observed": 3.238858936476838,
"passed": true,
"unit": "percentage_points"
},
"passed": true,
"record_count": {
"minimum": 400,
"observed": 400,
"passed": true
}
},
"side_accuracy": {
"baseline": {
"observed": 0.4475,
"passed": true
},
"minimum": 0.2,
"passed": true,
"subject": {
"observed": 0.4525,
"passed": true
}
},
"subject": {
"mean_score": 0.4525
},
"uncertainty": {
"interval_mass": 0.95,
"lower": -1.119943249438997,
"method": "newcombe_hybrid_score_paired_v2",
"scope": "paired_binary_outcomes",
"upper": 2.118915687037841
},
"verdict": "pass"
}