InvarLockRuntime evaluation evidence

Recorded policy result

Policy satisfied

Every configured check passed for this paired evaluation. Independent recipient acceptance is a separate step.

What was compared

Recorded fieldBaselineSubject
ArtifactDiffershf://mistralai/Ministral-3-8B-Instruct-2512-BF16@f6fae9795746f63c9be8344932f01275f3c63734derived://mistralai/Ministral-3-8B-Instruct-2512-BF16@f6fae9795746f63c9be8344932f01275f3c63734#llama.cpp-b10015-q5_k_m@sha256:b2c56b655b15577c30734b28b820d054e43df06c7712b74b9c170bff5f10f253
Model IDDiffersmistralai/Ministral-3-8B-Instruct-2512-BF16gguf-sha256-b2c56b655b15577c30734b28b820d054e43df06c7712b74b9c170bff5f10f253.gguf
ProviderDiffershf_transformersllama_cpp
Additional recorded context (10)

Additional fields

Workflow
Runtime execution
Task
text_causal
Paired records
400
Schedule digest
sha256:9ac24cf5e4b904197117fa8d0d9661d88fdf0dae1cc39ddfa711aa820de60fad
Dataset
TIGER-Lab/MMLU-Pro/ministral3-instruct
Dataset split
test-balanced-400
Dataset source SHA-256
c3d83209d6f36023f0a5aef5ee9be895891cc66ecc1b7196e83227558a38fade
Dataset source format
jsonl
Selected records
400
Selection limit
Not specified

Results and requirements

Exact-match accuracy

Scope: All paired records

Policy satisfied

All configured checks passed.

Baseline
44.75%
Subject
45.25%
Change
+0.5 ppPaired 95% CI: -1.12 to +2.119 pp
Observed pairs
400
EstimatePaired 95% confidence interval Policy thresholdNo changeAllowed change region
Decision checks
CheckObservedRequiredResult
Paired lower boundThe lower bound must meet the allowed change, not just the point estimate.-1.11994 pp>= -2 ppPassed
Record countThe comparison must include the configured minimum number of paired records.400>= 400Passed
Interval widthThe interval must be narrow enough for the configured precision requirement.3.23886 pp<= 10 ppPassed
Baseline accuracyEach side must meet the configured absolute accuracy floor.44.75%>= 20%Passed
Candidate accuracyEach side must meet the configured absolute accuracy floor.45.25%>= 20%Passed

Next steps

  1. Review the decision checks and their requirements.
  2. Run invarlock verify with your independently supplied trust profile and receipt destination to create the signed acceptance or rejection receipt.
  3. Keep this report alongside the original immutable evidence and the separate verification receipt.
Scope and limitations
  • This report summarizes the evidence bundle. Rendering checks its integrity and embedded signature; independent acceptance requires verification with your own trust inputs.
  • Results apply to the recorded cases, metric and policy. A pass does not establish general model quality, safety or representative production performance.
  • The interval describes paired binary outcomes under its stated method. Population interpretation requires an appropriate sampling design.

Evidence details

Identities, methods and exact recorded values

Comparison and policy identities
Comparison
comparison-4472c601a80f319859c4bf0c505c399c
Metric
exact_match
Policy
sha256:48e929519390c41e91c5e907e7e3f3d6d0b26e5ac4345136c4f6b85244b191d9
Evidence signer
sha256:e1f2aab12ee457abad9b0c74b1b9dcd0ce0429e541c15f17b814071b7fd9abab
Paired outcome analysis
{
  "baseline_fail_subject_pass": 6,
  "baseline_pass_subject_fail": 4,
  "both_fail": 215,
  "both_pass": 175,
  "discordant_pairs": 10,
  "effect_size_confidence_interval": {
    "confidence_level": 0.95,
    "lower_pp": -1.119943249438997,
    "method": "newcombe_hybrid_score_paired_v2",
    "upper_pp": 2.118915687037841
  },
  "effect_size_pp": 0.5,
  "mcnemar_exact_two_sided_p_value": 0.75390625
}
Authenticated observations
[
  {
    "authority": "observation",
    "bindings": {
      "artifact_digests": {
        "baseline": "sha256:2bba1d6c1331980047b18d9b00da21d997dc9ca26a183434e512bfc348b2e891",
        "subject": "sha256:fa3ab58b49a8e08b4901a3f990c0d0a401ec7930bd90a632f564ff7f71f1b1a6"
      },
      "comparison_id": "comparison-4472c601a80f319859c4bf0c505c399c",
      "policy_digest": "sha256:48e929519390c41e91c5e907e7e3f3d6d0b26e5ac4345136c4f6b85244b191d9",
      "schedule_digest": "sha256:9ac24cf5e4b904197117fa8d0d9661d88fdf0dae1cc39ddfa711aa820de60fad"
    },
    "format": "invarlock/evidence-observation-v1",
    "kind": "artifact_transformation",
    "observation_id": "ministral3-8b-bf16-to-gguf-q5-k-m",
    "payload": {
      "conversion": {
        "output": {
          "byte_length": 16987563264,
          "filename": "Ministral-3-8B-Instruct-BF16.gguf",
          "sha256": "b6dc36267813cef3a367b3e23528f5e367db5713ca6f2dd1f1a40a3eb1344ada",
          "type": "BF16"
        },
        "runtime_image_digest": "sha256:b905cf3b278531333d47fc80ba5b110793f80717fc2acc526e8dd291ae8f3301",
        "tool": {
          "name": "llama.cpp/convert_hf_to_gguf.py",
          "source_commit": "12127defda4f41b7679cb2477a4b0d65ee6a0c8f",
          "source_sha256": "5ab75e394f4c71425ecce64a213dab3b8e3e9cfe0f19d0dcda4d5a4f7733da83",
          "source_tag": "b10015"
        }
      },
      "format": "invarlock/example-gguf-deployment-transformation-v1",
      "quantization": {
        "runtime_image_digest": "sha256:aa5d7b75a807e786175e9edf9927cbbc0bb3099bcc2d95fe07e56f80bbfeedb6",
        "tool": "llama-quantize",
        "type": "Q5_K_M"
      },
      "source": {
        "checkpoint_tree_sha256": "sha256:6cbddcebc289550569cc3f6a93676a8f4f605d8574b8aec8448d61594a283996",
        "repository": "mistralai/Ministral-3-8B-Instruct-2512-BF16",
        "revision": "f6fae9795746f63c9be8344932f01275f3c63734",
        "tokenizer_contract_sha256": "b9e3906504b6235b5c289fe9d3f7a86512f968dfe300771ec660080924615dbc"
      },
      "subject": {
        "byte_length": 6058747136,
        "filename": "Ministral-3-8B-Instruct-Q5_K_M.gguf",
        "sha256": "b2c56b655b15577c30734b28b820d054e43df06c7712b74b9c170bff5f10f253"
      }
    },
    "scope": "subject"
  }
]
Exact comparison data
{
  "baseline": {
    "mean_score": 0.4475
  },
  "comparison": {
    "kind": "exact_match_delta_pp",
    "minimum": -2.0,
    "value": 0.5
  },
  "comparison_id": "comparison-4472c601a80f319859c4bf0c505c399c",
  "format": "invarlock/comparison-report-v3",
  "metric": "exact_match",
  "paired_binary": {
    "baseline_fail_subject_pass": 6,
    "baseline_pass_subject_fail": 4,
    "both_fail": 215,
    "both_pass": 175,
    "discordant_pairs": 10,
    "effect_size_confidence_interval": {
      "confidence_level": 0.95,
      "lower_pp": -1.119943249438997,
      "method": "newcombe_hybrid_score_paired_v2",
      "upper_pp": 2.118915687037841
    },
    "effect_size_pp": 0.5,
    "mcnemar_exact_two_sided_p_value": 0.75390625
  },
  "policy_digest": "sha256:48e929519390c41e91c5e907e7e3f3d6d0b26e5ac4345136c4f6b85244b191d9",
  "record_count": 400,
  "sample_qualification": {
    "interval_width": {
      "maximum": 10.0,
      "observed": 3.238858936476838,
      "passed": true,
      "unit": "percentage_points"
    },
    "passed": true,
    "record_count": {
      "minimum": 400,
      "observed": 400,
      "passed": true
    }
  },
  "side_accuracy": {
    "baseline": {
      "observed": 0.4475,
      "passed": true
    },
    "minimum": 0.2,
    "passed": true,
    "subject": {
      "observed": 0.4525,
      "passed": true
    }
  },
  "subject": {
    "mean_score": 0.4525
  },
  "uncertainty": {
    "interval_mass": 0.95,
    "lower": -1.119943249438997,
    "method": "newcombe_hybrid_score_paired_v2",
    "scope": "paired_binary_outcomes",
    "upper": 2.118915687037841
  },
  "verdict": "pass"
}