Evidence artifacts

An invarlock/evidence-pack-v1 bundle is a closed, signed evidence directory. Its paths are fixed by the manifest schema; additional files are rejected in strict verification.

Reference

Surface: invarlock/evidence-pack-v1 directory, manifest, payload inventory, and external outputs

Stability: Versioned public artifact contract; fixed paths and closed inventory are verification requirements

Use this page when: Inspecting a bundle, implementing artifact storage, or determining which bytes carry a particular claim

This layout applies to native exact-match/NLL and deterministic-extension evidence. Captured comparisons use evaluation records and pack v2. Native and captured judge requests use the judge evidence envelope; native judge evidence also retains its bound runtime capture. A shared CLI does not make these inventories or receipt contracts interchangeable.

Diagram
Evidence dependency map
Evidence dependency mapEvidence dependency map

Directory layout

Diagram
The closed evidence directory assigns one fixed path to the manifest, signature, checksum inventory, inputs, schedule, paired records, canonical report, per-side runs, and provider material.
The closed evidence directory assigns one fixed path to the manifest, signature, checksum inventory, inputs, schedule, paired records, canonical report, per-side runs, and provider material.The closed evidence directory assigns one fixed path to the manifest, signature, checksum inventory, inputs, schedule, paired records, canonical report, per-side runs, and provider material.

inputs/*.json are identity records, not copies of model weights, the dataset, or the policy. Each carries a material digest and may carry a stable locator or media type. The exact normalized request, schedule, provider material, paired records, and canonical report are separate payloads referenced by the manifest.

Data table with columns: Path or group, Cardinality, Authority in replay
Path or groupCardinalityAuthority in replay
request.jsonOneNormalized evaluation intent; never substitutes for independent verifier anchors
inputs/*.jsonSixRole-specific identity records and material digests
schedule/*.jsonOneOrdered dataset records and expected outputs
providers/{side}/*Five per sideArtifact, observation, provider, configuration, and runtime provenance
runs/{side}/report.jsonOne per sideMinimal side-to-observation binding
records/paired-records.jsonOneVerifier-reproducible record scores; not an accepted aggregate
reports/evaluation.report.jsonOneCanonical means, comparison, paired interval, threshold, optional sample and side-accuracy qualification, and verdict reproduced by verification
observations/*.jsonZero to 64Authenticated typed context; descriptive only and excluded from policy replay
checksums.sha256OneExact payload-byte inventory
manifest.json and signatureOne eachClosed references and signer authentication

Scroll horizontally to see every column.

Manifest

manifest.json contains exactly:

  • format and deterministic comparison_id;
  • six input references: baseline, subject, dataset, both runtimes, and policy;
  • fifteen evidence-role references at their fixed paths;
  • the paired-record path, digest, and count;
  • optional typed observation references with fixed paths, digests, kinds, and scopes;
  • the checksum-file path and digest; and
  • the evidence-signing key fingerprint.

Every input reference has a digest of the identity file and a distinct material_digest for the authenticated input. Every evidence reference has a digest of the exact file bytes.

The manifest structure is equivalent to this abbreviated, non-copyable view:

{
  "format": "invarlock/evidence-pack-v1",
  "comparison_id": "...",
  "inputs": {
    "baseline": {
      "path": "inputs/baseline.json",
      "digest": "sha256:...",
      "material_digest": "sha256:..."
    },
    "subject": { ... },
    "dataset": { ... },
    "baseline_runtime": { ... },
    "subject_runtime": { ... },
    "policy": { ... }
  },
  "evidence": {
    "request": {"path": "request.json", "digest": "sha256:..."},
    "schedule": {"path": "schedule/runtime-behavioral-schedule.json", "digest": "sha256:..."},
    "evaluation_report": {"path": "reports/evaluation.report.json", "digest": "sha256:..."},
    "...": "twelve additional fixed evidence roles"
  },
  "paired_records": {
    "path": "records/paired-records.json",
    "digest": "sha256:...",
    "count": 20
  },
  "checksums_sha256": "checksums.sha256",
  "checksums_sha256_digest": "...bare sha256...",
  "signing_key_fingerprint": "sha256:..."
}

The real object has no comments, ellipses, or abbreviated roles. It has exactly six input references and fifteen evidence references at schema-fixed paths.

manifest.signature.json embeds the Ed25519 public key, its fingerprint, and a base64 signature over the exact canonical manifest.json bytes. The public key is evidence of who signed only after its fingerprint is matched to an independent expected signer.

Data table with columns: Signature field, Meaning
Signature fieldMeaning
formatinvarlock/evidence-pack-signature-v1
algorithmEd25519
Embedded public keyVerification material, not an independent trust anchor
Public-key fingerprintSHA-256 of the raw Ed25519 public key
SignatureAuthentication of the exact canonical manifest bytes

Checksums and closed inventory

checksums.sha256 covers every payload below the manifest, but not the manifest or its signature. The manifest separately binds the checksum file. Verification checks both directions:

  • every manifest payload must be covered; and
  • no checksum or directory entry may introduce an undeclared payload.

Regular-file, no-symbolic-link, size, and safe-relative-path rules apply while reading. A valid checksum does not make a file semantically valid; cross-file replay follows the integrity checks.

The manifest is outside the checksum payload because it binds the checksum file. The signature is outside both because it authenticates the canonical manifest. This avoids a self-referential digest while still closing every pack byte into one directed dependency graph.

Paired records

records/paired-records.json has exactly:

{
  "format": "invarlock/paired-records-v1",
  "metric": "exact_match",
  "schedule_sha256": "...",
  "records": [
    {
      "record_id": "example-1",
      "input_sha256": "...",
      "baseline": {
        "observation_record_digest": "sha256:...",
        "score": 1.0
      },
      "subject": {
        "observation_record_digest": "sha256:...",
        "score": 1.0
      }
    }
  ]
}

The engine derives this file from the canonical schedule and both scoring observations. It rejects failed observation records, order changes, input-digest mismatches, output-text/digest mismatches, and schedules without expected outputs. Import mode must supply byte-canonical paired records equal to that fresh derivation; imported scores are never accepted as authority.

For exact_match, each score is 1.0 when output text exactly equals expected output and 0.0 otherwise. For normalized_nll_per_utf8_byte, each score is -logprob_sum / utf8_byte_count; required counts must be positive and the result finite and non-negative.

When a scorer extension is selected, the paired-record object also binds the exact scorer configuration and each side's replay result. The explicitly authorized scorer receives only authenticated expected_output, output_text, and output_sha256 facts and returns one higher-is-better value in [0, 1] per record. The core verifies result order and digests, computes the arithmetic mean and paired percentage-point delta, and owns the interval and verdict.

Normalized NLL is teacher-forced expected-continuation likelihood regression, not a general model-quality measure. The paired-record object may also carry a verifier-derived perplexity interpretation when the authenticated tokenizer contracts and every pair's positive target-token counts match. That object has no policy, interval, or verdict authority; an unavailable interpretation does not invalidate the byte-normalized NLL comparison.

The canonical report adds the metric-specific paired interval to the two side means and point comparison. New reports use invarlock/comparison-report-v3; the verifier also replays existing v2 and invarlock/comparison-report-v1 packs under their original semantics. Exact match records paired regression and improvement counts, an exact two-sided McNemar probability, and the paired Newcombe 95% effect-size interval whose lower bound controls policy. Normalized NLL uses the upper bound of its authenticated-schedule resampling interval. An authorized scorer extension uses the lower bound of the same fixed paired-resampling method against metrics.scorer_extension.delta_min_pp.

When the policy supplies the coupled sample controls, a v2 or v3 report also contains sample_qualification. It records the minimum and observed paired count, maximum and observed interval width, units, each check's result, and the combined result. The final verdict is the conjunction of that result and the metric-bound result.

When a v3 exact-match policy supplies minimum_side_accuracy, the report also contains side_accuracy. It records the minimum, both observed means, each side's result, and the combined result. The final verdict additionally requires that combined result to pass. Historical v2 reports have no such section.

Provider sidecars

Each comparison side includes:

Data table with columns: File, Binding
FileBinding
model-artifact.identity.jsonPortable authenticated identity for the actual model artifact
runtime-scoring.observation.jsonOrdered backend facts bound to provider, artifact, and schedule
runtime-provider.receipt.jsonPlugin/backend/capability/artifact/settings/device/image provenance bound to the observation
report.jsonMinimal side report binding provider, artifact, observation, schedule, and count
run.yamlCanonical side configuration copied as exact evaluation-emitted bytes
runtime.manifest.jsonReport/config/sidecar digests and strict per-side worker-container facts

The evidence transaction copies these exact bytes. It does not regenerate sidecars in a way that could break their original runtime-manifest bindings.

Publication and immutability

Evaluation writes a staging directory, flushes files to disk, signs the manifest, and atomically renames the directory into place. It never replaces an existing destination. Published files are read-only and directories are non-writable.

Filesystem permissions are a local hardening measure, not the integrity model. Any later byte change, missing file, or extra file is detected by strict verification. Verification receipts and rendered report files must remain outside the bundle so the bundle can stay byte-identical.

Output writers pin directory descriptors and check pathname bindings around publication. They reject changes observed at those checks; they cannot prevent later changes by a process with the same filesystem permissions. On an ambiguous publication failure, a completed or competing file, or private staging directory, can remain. Inspect and authenticate any retained evidence before using it. Cleanup never recursively follows an old staging pathname.

Runtime-provider sidecars are published individually. If a later sidecar fails, earlier completed files remain. Callers that need an all-or-nothing set must use a private directory and publish that directory only after all checks succeed.

External outputs

Output
Signed verification receipt
Produced by
invarlock verify
May be written inside pack?
No
Trust meaning
Independent verifier assertion over manifest, anchors, and verdict
Output
Console report
Produced by
invarlock report
May be written inside pack?
Not a file
Trust meaning
Summary of canonical content after integrity and embedded-signature checks
Output
Self-contained HTML
Produced by
invarlock report --html
May be written inside pack?
No
Trust meaning
Unsigned presentation; not independent technical verification
Output
Markdown report
Produced by
invarlock report --markdown
May be written inside pack?
No
Trust meaning
Unsigned presentation; not independent technical verification
Output
JUnit report
Produced by
invarlock report --junit
May be written inside pack?
No
Trust meaning
CI test results reflecting the recorded comparison; not independent technical verification

Multiple verifiers can issue separate receipts for the same immutable manifest. They can use distinct verifier identities and keys. Each must independently supply the expected artifact, schedule, policy, runtime, and evidence-signer anchors, plus the normalized-request anchor required for GGUF evidence. The policy bytes must match the policy material already bound by the pack; a different policy requires a newly evaluated evidence bundle.

  • Public contracts defines the manifest schema, canonical JSON, and digest rules.
  • Reports and receipts defines the report and external receipt carried by or associated with a bundle.
  • Evaluation lifecycle explains atomic publication and failure boundaries.
  • Architecture places the artifact inventory within the evidence-signer-verifier trust model.