Public contracts

The workflow uses closed, versioned JSON contracts. Source mirrors live in contracts/, and byte-identical package-owned copies ship in the core wheel. Verification always loads the package-owned copies: a working directory or environment variable cannot substitute a different schema.

Reference

Surface: Versioned request, evidence, provider, runtime, report, and receipt contracts

Stability: Closed public interchange formats; incompatible shape or meaning changes require a new format identifier

Use this page when: Authoring contract objects, validating canonical bytes, or reviewing cross-file digest and signature bindings

Schema-backed contracts

Data table with columns: File, Format, Purpose
FileFormatPurpose
evaluation_request.schema.jsoninvarlock/evaluation-request-v1One closed run-or-import request
evaluation_request_v2.schema.jsoninvarlock/evaluation-request-v2Captured-only request with paired source specifications and complete-run pins
evaluation_request_v3.schema.jsoninvarlock/evaluation-request-v3Bounded judge import or collection request over frozen runs
evaluation_setup_result.schema.jsoninvarlock/evaluation-setup-v1CLI example, key generation, or case-set preparation result; no evaluation assurance
evaluation_case_set.schema.jsoninvarlock/evaluation-case-set-v1Immutable case membership and inputs shared by captured runs
evaluation_run.schema.jsoninvarlock/evaluation-run-v1Complete attributed answers and optional recorded scores for one side
normalized_captured_request.schema.jsoninvarlock/evaluation-request-v2Canonical captured-request projection retained in evidence
comparison_policy.schema.jsoninvarlock/comparison-policy-v1Deterministic metric, slice, count, precision, and regression requirements
multi_metric_comparison.schema.jsoninvarlock/multi-metric-comparison-v1Recomputed deterministic results across declared metrics and slices
scorer_extension_descriptor.schema.jsoninvarlock/scorer-extension-descriptor-v1Installed deterministic scorer identity and declared capabilities
scorer_extension_binding.schema.jsoninvarlock/scorer-extension-binding-v1Request binding to an exact scorer descriptor and configuration
scorer_extension_result.schema.jsoninvarlock/scorer-extension-result-v1Per-side deterministic scorer replay facts
evidence_pack.schema.jsoninvarlock/evidence-pack-v1Canonical bundle manifest and fixed payload paths
evidence_pack_v2.schema.jsoninvarlock/evidence-pack-v2Captured-only closed directory inventory and explicit signed/unsigned authentication
evidence_verification_receipt_v3.schema.jsoninvarlock/evidence-verification-receipt-v3Captured-only external verifier statement with run/request anchors and per-metric/slice scoring assurance
judge_measurement_plan.schema.jsoninvarlock/judge-measurement-plan-v1Frozen judge identity, prompt, rubric, scale, answer bindings, units, and schedule
judge_measurements.schema.jsoninvarlock/judge-measurements-v1Retained per-trial attempts, responses, parsed ratings, and completeness
judge_analysis_policy.schema.jsoninvarlock/judge-analysis-policy-v1Independent judge metric, interval, precision, and degradation requirements
judge_measurement_evidence.schema.jsoninvarlock/judge-measurement-evidence-v1Signed or unsigned judge evidence envelope and artifact digests
judge_measurement_recipient_policy.schema.jsoninvarlock/judge-measurement-recipient-policy-v1Recipient signer, subject, metric, result, and exact-artifact pins
judge_verification_result.schema.jsoninvarlock/judge-verification-result-v1Unsigned local judge authentication, replay, policy decision, and acceptance result
judge_measurement_verification_receipt.schema.jsoninvarlock/judge-measurement-verification-receipt-v1Scoped Ed25519 verifier statement binding the complete local judge result and recipient policy
evidence_set.schema.jsoninvarlock/evidence-set-v1Unsigned index of one deterministic captured pack and one bounded judge pack
evidence_set_recipient_policy.schema.jsoninvarlock/evidence-set-recipient-policy-v1Independent shared-run and component-policy pins for required conjunction
evidence_set_verification.schema.jsoninvarlock/evidence-set-verification-v1Fresh local conjunction result retaining distinct component receipts
evidence_observation.schema.jsoninvarlock/evidence-observation-v1Typed observation-only envelope and comparison bindings
trust_inputs.schema.jsoninvarlock/trust-inputs-v1Independent policy, anchors, verifier identity/key path, and scorer authorization
trust_inputs_v2.schema.jsoninvarlock/trust-inputs-v2Captured independent policy, complete-run/request/signer anchors and verifier identity/key path
acceptance_predicate.schema.jsoninvarlock/acceptance-predicate-v2Portable projection of one technical decision in an in-toto Statement
recipient_acceptance_policy.schema.jsoninvarlock/recipient-acceptance-policy-v2Current recipient trust, freshness, version, signer, and verdict rules
evaluator_qualification_profile.schema.jsoninvarlock/evaluator-qualification-profile-v1Evaluator identity, execution provenance, and authority classification
evaluator_qualification_schedule.schema.jsoninvarlock/evaluator-qualification-schedule-v1Independent ordered record and reference identities
evaluator_qualification_export.schema.jsoninvarlock/evaluator-qualification-export-v1Normalized per-record facts or an observation-only summary
evaluator_qualification_result.schema.jsoninvarlock/evaluator-qualification-result-v1Digest-bound qualification outcome and import authority

Scroll horizontally to see every column.

The acceptance predicate and recipient policy are described in Acceptance attestations. The detailed InvarLock receipt remains the authoritative replayable result. Native pack v1 and receipt v1/v2 retain their published meanings. Captured pack v2 and receipt v3 are the deterministic captured evidence family; captured metric: judge uses the separate judge envelope and receipt. Legacy monolithic comparison evidence is not accepted. A captured receipt's verification_scope: captured_comparison does not qualify native execution, acceptance attestations, ModelKit acceptance, or deployment approval. See Captured records for run, case-set, policy and request normalization contracts; these schema identifiers are not product-release pins.

Evidence sets compose existing deterministic and judge components over the same original runs. Their combined decision does not add a joint confidence guarantee or change either component's acceptance scope.

Evaluator qualification has two stable wire classifications:

Data table with columns: Qualification result, Meaning
Qualification resultMeaning
outcome: qualified_for_import, authority: verdict_authorityComplete ordered facts passed deterministic recomputation and are independently replayable
outcome: observation_only, authority: observation_onlyAuthenticated context was retained but cannot contribute imported verdict facts

These fields do not express adapter maintenance or signed-journey maturity. Those independent axes belong to the examples-layer qualification catalog, so an export cannot promote itself by claiming support or demonstration status.

Identity and evaluation-context boundaries

A model name is a display label. It does not select the evaluated artifact, package, runtime, scorer or complete execution context. Independent verification requires recipient-selected expectations and checks the relationships between the authenticated objects, in addition to validating their schemas.

Data table with columns: Requirement, Existing binding and recipient check, Authority and limit
RequirementExisting binding and recipient checkAuthority and limit
Exact artifactTyped HF snapshot, GGUF or TensorRT-LLM identity; independent artifact-identity digest and actual content checks where suppliedContent identity is separate from a model name or declared ancestry
Tokenizer and templateArtifact tokenizer-metadata digest; supported providers measure relevant files, including HF chat-template metadataAdditional processor or execution settings must be checked through the applicable provider/request binding
Evaluated task and scheduleNormalized request, canonical schedule, and both authenticated provider capability declarations must agreeA valid signature cannot make conflicting task declarations consistent
Generation, runtime and security settingsProvider-specific request/receipt checks and independent runtime digests; a full normalized-request digest can additionally pin the exact declared contextAuthenticated settings and local enforcement tests do not establish independent hardware attestation
Scorer and evaluation dataScorer identifier/version/descriptor/configuration bindings, policy bytes, and independently selected schedule identityA qualified scoring domain does not imply support for every behavior of the external evaluator
Complete transaction requestOptional independent request_digest in the trust-input profile, or --expected-request-digest; required for the llama.cpp pathReceipt v2 records this expectation; receipt v1 does not acquire it retroactively
Package-to-model mappingThe ModelKit example verifies recipient-selected package blobs, both model directories and their relation to replayed evidenceThis is an example-owned point-of-use check; the generic acceptance envelope alone does not verify a ModelKit
Transformation or contextual observationsCanonical payload digest in the normalized request, comparison-bound observation envelope, and signed manifest inventoryAuthenticates the payload and its association; arbitrary payload claims are not independently validated
Current recipient acceptanceTrusted envelope and receipt signers, exact transported identity consistency, independent subject digest or actual-content binding, contract versions, freshness and current policyCurrent acceptance is separate from the original technical result

Scroll horizontally to see every column.

The independently selected complete-request digest can require exact declared settings and observation contents without introducing another model hash. It does not independently confirm a hosted service's reported revision or an observation's scientific conclusion. Do not obtain an expected digest from incoming evidence and describe the resulting equality as independent trust.

Observation payloads may use a versioned profile for method, configuration, probe-set identity, assumptions, result and uncertainty. A consumer must implement that profile's semantic checks before claiming to understand or validate it. The current recipient policy does not implement a generic required lineage or context-profile result. Its optional receipt trust-profile digest identifies the selected verifier configuration; it does not create missing validation logic. Unknown policy fields reject, while opaque observation payloads retain only observation authority.

Declared transformation history, empirical similarity and exact artifact identity remain separate. No supported behavioral fingerprint proves ancestry merely because its payload is signed. A future profile must avoid including the complete normalized-request digest inside a payload already hashed by that request; bind component identities first and let the envelope add comparison bindings after normalization.

The evaluation-record contracts bind complete captured runs and policy bytes, including source versions and per-record context. Their independent run digests include outputs, so they are not pre-execution context identities. Replaying signed captured evidence verifies the comparison; it does not prove that an untrusted capture worker executed the declared model. The captured rehearsal independently pins the protocol and capture before reconstructing those runs.

Provider contracts

The provider ABI uses these schema-backed documents:

Data table with columns: File, Format, Purpose
FileFormatPurpose
model_artifact_identity.schema.jsoninvarlock/model-artifact-identity-v1Portable HF, GGUF, or TensorRT-LLM artifact identity
runtime_provider_receipt.schema.jsoninvarlock/runtime-provider-receipt-v1Provider, backend, artifact, settings, device, image, and observation binding
runtime_scoring_observation.schema.jsoninvarlock/runtime-scoring-observation-v1Ordered backend-measured record facts
runtime_behavioral_schedule.schema.jsoninvarlock/runtime-behavioral-schedule-v1Dataset identity and ordered input records
runtime_manifest.schema.jsonruntime-manifest-v1Strict execution envelope and sibling sidecar digests
runtime_provider_capabilities.jsonruntime-provider-capabilities-v1Provider ABI, artifacts, tasks, metrics, execution modes, and requirements

Scroll horizontally to see every column.

Python callers can load these exact packaged objects with functions from invarlock.public_contracts. Those loaders return schema dictionaries; they do not validate semantic cross-bindings by themselves.

from invarlock.public_contracts import (
    load_evaluation_request_schema,
    load_evidence_observation_schema,
    load_evidence_pack_schema,
    load_trust_inputs_schema,
    load_acceptance_predicate_schema,
    load_recipient_acceptance_policy_schema,
    load_model_artifact_identity_schema,
    load_runtime_behavioral_schedule_schema,
    load_runtime_manifest_schema,
    load_runtime_provider_capabilities_schema,
    load_runtime_provider_receipt_schema,
    load_runtime_scoring_observation_schema,
)

Each loader returns a new dictionary decoded from the package-owned contract. ContractLoadError identifies a missing, malformed, or non-object packaged contract. Format constants such as EVALUATION_REQUEST_FORMAT_VERSION, EVIDENCE_PACK_FORMAT_VERSION, TRUST_INPUTS_FORMAT_VERSION, and RUNTIME_PROVIDER_ABI_VERSION are exported from the same module for exact comparisons.

The schemas use JSON Schema Draft 2020-12. Every contract object is closed: fields not named by its schema or exact code-enforced shape are rejected rather than ignored.

Independent trust-input profile

invarlock/trust-inputs-v1 is the portable caller-owned input to independent verification. It contains the policy path, baseline and subject artifact digests, schedule digest, both runtime digests, evidence-signer fingerprint, verifier identity, verifier signing-key path, and installed-scorer authorization. The object and all nested objects are closed.

Policy and key paths are safe relative paths resolved from the profile's directory. Absolute paths, traversal, symlinks, duplicate JSON members, unknown fields, and missing files are rejected. Formatting does not affect the profile digest: the loader hashes canonical JSON and the verifier records that digest in its signed receipt. The profile never contains private-key bytes.

Evaluation request fields

The request root contains four required fields and one optional field:

Data table with columns: Field, Type, Required value or role
FieldTypeRequired value or role
format_versionStringinvarlock/evaluation-request-v1
comparisonObjectBaseline, subject, dataset, task, policy, and exactly one metric or scorer-extension binding
executionObjectExactly one run or import transaction
observationsArrayZero to 64 authenticated context attachments
outputObjectExactly evidence, a safe relative destination

Scroll horizontally to see every column.

Each comparison side has the same closed shape:

Path
artifact.model_id
Type
String
Requirement
Yes
Meaning
Display name for the artifact; URL syntax is rejected
Path
artifact.locator
Type
String
Requirement
Yes
Meaning
Portable source locator bound into request intent
Path
artifact.path
Type
Safe relative path
Requirement
Required in run mode
Meaning
Artifact path below the request root
Path
runtime.provider
Type
Provider name
Requirement
Yes
Meaning
Selected runtime-provider ABI implementation
Path
runtime.settings
Type
Object of JSON scalars
Requirement
Yes
Meaning
Provider-owned settings validated against capabilities

For metric: judge, comparison.judge additionally requires a private workspace and signer_identity, and comparison.policy names the closed invarlock/native-judge-policy-v1 recipe. Its code-enforced plan and analysis bindings are described in judge measurements.

comparison.policy is always a safe relative path. comparison.dataset is mode-specific:

  • run mode requires a closed local-dataset object with path, bare lowercase sha256, format: jsonl, name, split, input_field, and expected_output_field, plus optional id_field, an all-or-none content role and field mapping, and limit;
  • import mode requires a safe relative path to canonical invarlock/runtime-behavioral-schedule-v1 bytes.

The canonical comparison.task binds the request, schedule, provider capabilities, evaluation batches, and provider receipts. Built-in identifiers are text_causal, masked_language, text_seq2seq, and vision_text_generation; a provider must explicitly declare execution support. The request selects exactly one built-in metric or one complete scorer_extension binding. Built-in scorers are exact_match, normalized_nll_per_utf8_byte and judge. Providers declare the required collection metric: judge and deterministic scorer extensions use exact_match to retain complete text outputs; normalized NLL requires its own likelihood facts. A scorer extension uses that authenticated collection so that expected and observed text are authenticated for verifier replay. The built-in hf_transformers provider declares exact-match and normalized-NLL collection. The first-party llama_cpp, tensorrt_llm, and hf_vision_text add-ins currently declare exact match for their tasks.

Deterministic scorer extension

comparison.scorer_extension is a closed invarlock/scorer-extension-binding-v1 object. It binds scorer ABI 1, a dotted scorer ID, semantic version, descriptor digest, configuration object, and canonical configuration digest. The descriptor fixes supported tasks and input/output kinds, the configuration-schema digest, and these v1 semantics:

  • replay reads exactly authenticated expected_output, output_text, and output_sha256 facts for every ordered record;
  • each result is a finite higher-is-better value in [0, 1];
  • the core computes the arithmetic mean, subject-minus-baseline percentage- point delta, and fixed 2,048-replicate paired interval; and
  • network access, external models and externally assigned ratings are forbidden in a deterministic extension scorer.

The independently supplied policy must contain resolved_policy.metrics.scorer_extension with the same scorer_id, scorer_version, descriptor_sha256, and configuration_sha256, plus delta_min_pp. Evaluation and verification require an explicitly authorized ScorerExtensionRegistry; the request and evidence cannot authorize scorer code. The verifier runs the scorer twice, requires identical canonical results, then independently reconstructs the core-owned aggregate, paired interval, threshold comparison, and verdict.

Core ships invarlock.normalized_match, invarlock.numeric_tolerance, invarlock.json_fields, invarlock.json_exact and invarlock.token_f1 through this boundary. The CLI enables them without --allow-installed-scorers; SDK callers supply ScorerExtensionRegistry(allow_installed=False). Their exact version, descriptor and configuration bindings remain mandatory. Additional implementations, such as a VQA normalization scorer, require separate installation or caller injection and explicit authorization. SQL or code execution, model-based semantic similarity, network services, externally supplied ratings, and LLM judges require different trust contracts. The bounded frozen-answer judge formats provide one such contract for their declared text profile; other judge outputs can be attached as authenticated observations without technical-verdict authority.

Evaluator input boundary

Native provider import mode is an extension boundary for measurements produced by an external evaluator. Under that contract, output is admissible for an acceptance decision only when InvarLock can authenticate the ordered per-record inputs and outputs, bind them to the exact schedule, artifacts, runtime, and source, and deterministically recompute the decision-contract metric or authorized scorer.

Captured evaluator integrations use the separate evaluation-record contracts. They bind supplied runs and scorer-specific facts without requiring native provider sidecars or asserting native execution. Captured judge requests use the bounded judge evidence and recipient-policy contracts.

An adapter alone does not establish evaluator neutrality. The generic qualification boundary binds the profile, independent schedule, normalized export, retained upstream output, runner bundle, and dependency declaration. For a deterministic exact-match profile, every ordered input and output must be present and InvarLock independently recomputes every score. Aggregate-only results, missing or reordered record facts, and external-judge outputs outside the bounded judge-measurement profile remain observation-only and expose no runtime-import records.

The maintained evaluator qualification matrix executes representative upstream tools through example-owned runners. Every deterministic profile also scores and replays the complete 102-record output of a pinned real model evaluation through the runtime-import boundary. The matrix separately records profiles that demonstrate the deeper model-running, signed transaction journey. These levels record evidence maturity rather than a permanent support hierarchy and can advance without changing the generic boundary. Evaluator names and native parsers remain outside the engine; a private evaluator crosses the same JSON, CLI, or Python SDK boundary.

Run request

format_version: invarlock/evaluation-request-v1
comparison:
  baseline:
    artifact:
      model_id: acme/baseline
      locator: registry://acme/baseline@immutable-revision
      path: models/baseline
    runtime:
      provider: hf_transformers
      settings:
        immutable_revision: 0123456789abcdef0123456789abcdef01234567
        checkpoint_tree_sha256: "1111111111111111111111111111111111111111111111111111111111111111"
        tokenizer_metadata_sha256: "3333333333333333333333333333333333333333333333333333333333333333"
        batch_size: 1
        context_length: 2048
        max_output_tokens: 64
        offline: true
        seed: 7
        timeout_seconds: 120
  subject:
    artifact:
      model_id: acme/subject
      locator: registry://acme/subject@immutable-revision
      path: models/subject
    runtime:
      provider: hf_transformers
      settings:
        immutable_revision: fedcba9876543210fedcba9876543210fedcba98
        checkpoint_tree_sha256: "2222222222222222222222222222222222222222222222222222222222222222"
        tokenizer_metadata_sha256: "3333333333333333333333333333333333333333333333333333333333333333"
        batch_size: 1
        context_length: 2048
        max_output_tokens: 64
        offline: true
        seed: 7
        timeout_seconds: 120
  dataset:
    path: inputs/release-regression.jsonl
    sha256: "4444444444444444444444444444444444444444444444444444444444444444"
    format: jsonl
    name: release-regression
    split: validation
    input_field: prompt
    expected_output_field: expected
    id_field: case_id
    limit: 400
  policy: policy/acceptance.json
  task: text_causal
  metric: normalized_nll_per_utf8_byte
execution:
  mode: run
output:
  evidence: artifacts/evidence

The digest values are illustrative. A real request must bind the actual source, artifact, and tokenizer bytes. Provider identity digests in runtime.settings and comparison.dataset.sha256 use bare lowercase 64-character values; general bundle and runtime identities use the sha256: prefix where their contracts require it.

For run mode, the transaction verifies the JSONL digest, preserves source order, maps the declared top-level text fields, selects the exact prefix named by limit, derives stable IDs or deterministic position IDs, and builds the canonical schedule. Blank lines, invalid JSON objects, missing mapped text, duplicate IDs, digest mismatch, or a limit larger than the source fail closed.

Import request

Import mode replaces provider execution with authenticated sidecars. In addition to mode, it requires records, schedule, and these six references for both baseline and subject:

Data table with columns: Import-side field, Expected document
Import-side fieldExpected document
identityinvarlock/model-artifact-identity-v1
receiptinvarlock/runtime-provider-receipt-v1
observationinvarlock/runtime-scoring-observation-v1
run_reportinvarlock/runtime-side-report-v1
runtime_manifestruntime-manifest-v1
runtime_configCanonical provider run configuration
execution:
  mode: import
  records: import/paired-records.json
  schedule: inputs/schedule.json
  baseline:
    identity: import/baseline/model-artifact.identity.json
    receipt: import/baseline/runtime-provider.receipt.json
    observation: import/baseline/runtime-scoring.observation.json
    run_report: import/baseline/report.json
    runtime_manifest: import/baseline/runtime.manifest.json
    runtime_config: import/baseline/run.yaml
  subject:
    identity: import/subject/model-artifact.identity.json
    receipt: import/subject/runtime-provider.receipt.json
    observation: import/subject/runtime-scoring.observation.json
    run_report: import/subject/report.json
    runtime_manifest: import/subject/runtime.manifest.json
    runtime_config: import/subject/run.yaml

This fragment is inserted under the same request root as the comparison and output objects. In import mode, comparison.dataset and execution.schedule both name the canonical schedule. Evaluation re-derives artifact identities and paired records; supplying a valid-looking aggregate is not sufficient.

Parser and path limits

These are native v1 request limits. Captured comparisons use the separate capacity limits; judge recipes and measurements use judge operational bounds.

Data table with columns: Boundary, Limit or rule
BoundaryLimit or rule
Request YAMLAt most 1 MiB, 64 nested levels, and 10,000 syntax nodes
Policy inputAt most 4 MiB
Other request inputAt most 64 MiB per file
Local JSONL source or schedule bytesAt most 16 MiB and 1 to 10,000 selected records
ReferencesRelative, root-confined, forward-slash paths; no links, traversal, URLs, drive prefixes, or empty components

The YAML loader accepts only a JSON-compatible mapping. It rejects aliases, anchors, tags, directives, merge keys, duplicate keys, include-like keys, unsafe scalar types, and non-canonical scalar spellings. File reads repeat component-by-component no-follow checks at use time, so a path that passed the first parse cannot be replaced with a symbolic link unnoticed.

The built-in deterministic native metrics are exact_match and normalized_nll_per_utf8_byte; native judge uses the separately described judge analysis policy. Exact-match reports include paired outcome counts, an exact two-sided McNemar probability, and a versioned paired Newcombe 95% interval whose lower bound controls policy. Current v3 reports and historical v2 reports use the continuity-corrected method; strict verification preserves the original method for legacy v1 reports. Normalized-NLL reports include the fixed 2,048-replicate paired_percentile_bootstrap_sha256_v1 interval over the authenticated schedule and apply the policy ceiling to its upper bound. A scorer-extension comparison uses the same fixed paired-resampling method over unit-interval record values and applies delta_min_pp to the lower bound. Decision semantics defines the exact arithmetic.

Every metric policy contains its threshold and may also contain the coupled minimum_record_count and maximum-width fields. Exact match and scorer extensions use maximum_interval_width_pp; normalized NLL uses maximum_interval_width_ratio. Exact match may independently add minimum_side_accuracy, a finite value from 0 through 1 that both side means must meet. A v3 report passes only when the metric-bound check and every configured sample-qualification and side-accuracy check pass. Historical v2 reports have no side-accuracy section and retain their original semantics.

Provider document field map

The schemas remain authoritative for types, patterns, nullability, and nested conditional rules. This map makes every top-level contract field discoverable:

Data table with columns: Contract, Required top-level fields
ContractRequired top-level fields
Runtime capabilitiesformat_version, provider_abi, provider_name, artifact_formats, tasks, metrics, execution_modes, required_extra, required_image
Artifact identityformat_version, artifact_format, plus the format-specific identity fields listed in the runtime-provider reference
Scoring observationformat_version, provider_name, artifact_identity_sha256, schedule_sha256, records, aggregate_source_sha256
Provider receiptformat_version, plugin, backend, capabilities, artifact_identity, execution_settings, device, outer_image_digest, scoring_observation_sha256
Runtime manifestmanifest_version, generated_at_utc, verifier_contract_version, report, config, execution_mode, outer_container, runtime_provider
Behavioral scheduleformat_version, task, dataset_identity, records

Nested provider-receipt groups are closed:

Data table with columns: Group, Fields
GroupFields
pluginprovider name, distribution, distribution version, provider ABI
backendname, version, and at least one source, binary, or build SHA-256
execution_settingsseed, context length, batch size, output limit, timeout, network permission
devicekind, name, optional compute capability, driver, and CUDA runtime
capabilitiesthe complete capability object above
artifact_identityone exact HF, GGUF, or TensorRT-LLM identity variant

Each scoring record requires record_id, input_sha256, and status. An ok record contains output text plus its digest, log-likelihood facts, or both. An error record contains only a canonical error_code and no measured facts. Log-likelihood facts are finite logprob_sum, positive token_count, and positive utf8_byte_count. Byte-normalized NLL verifies the byte count against the scheduled target. Matching authenticated tokenizer metadata and equal positive paired token counts allow a verifier-derived perplexity interpretation without adding a metric or policy surface.

The runtime manifest fixes execution_mode to container. Its outer-container object binds the image reference/digest and observed execution switches; its runtime-provider object binds the provider name, ABI, and sibling identity, observation, and receipt files. report and config bind their sibling files by digest. Unknown fields at any of these levels are rejected.

Closed formats without standalone schemas

Several exact shapes are enforced by code and verifier replay rather than a separate JSON Schema file:

Data table with columns: Format, Role
FormatRole
invarlock/evidence-input-identity-v1One input role, material digest, and optional locator/media type
invarlock/paired-records-v1Verifier-derived baseline/subject scores in schedule order
invarlock/runtime-side-report-v1Minimal link from one side to its provider observation
invarlock/comparison-report-v3Current canonical means, point comparison, metric-specific paired interval, optional sample and exact-match side-accuracy qualification, threshold, and verdict
invarlock/comparison-report-v2Historical canonical report without side-accuracy qualification; accepted for backward verification, not emitted for new evaluations
invarlock/comparison-report-v1Legacy canonical report replayed with its original exact-match interval method; accepted for backward verification, not emitted for new evaluations
invarlock/evidence-pack-signature-v1Ed25519 signature over canonical manifest.json bytes
invarlock/evidence-pack-verify-v1Machine-readable independent verification result
invarlock/evidence-verification-receipt-v1Signed statement binding the pack, artifact/schedule/policy/runtime/signer anchors, verifier, optional trust-profile digest, and verdict
invarlock/evidence-verification-receipt-v2The v1 statement plus an independently supplied normalized-request digest; emitted when that anchor is present and required for GGUF evidence
invarlock/evidence-verification-receipt-signature-v1Ed25519 envelope for the receipt statement

These are documented for inspection and interchange, not as permission to construct partial objects. Use evaluate to produce bundles, verify to produce receipts, and the stable Python verify_signed_verification_receipt facade to validate a received receipt.

Canonical JSON and digests

Canonical JSON uses UTF-8, sorted object keys, compact separators, and finite numbers. Newline behavior is contract-specific: most bundle JSON documents and signed receipt statements use one trailing line feed, while canonical schedule, artifact-identity, scoring-observation, and embedded substructure hash inputs use the same representation without a final line feed. Writers and readers must use the contract serializer rather than normalize whitespace themselves.

The base JSON syntax follows RFC 8259. InvarLock deliberately narrows it with unique keys, finite numbers, closed schemas, exact encodings, and semantic cross-replay.

Bundle-level digests use lowercase sha256:<64 hex>. Some provider ABI fields and the checksum-file digest use a bare lowercase 64-character SHA-256 value; their schemas identify that distinction. Signatures use Ed25519. A key fingerprint is sha256: followed by the SHA-256 of the raw Ed25519 public key. Ed25519 is specified by RFC 8032.

Digest and signature dependency chain

Diagram
Contract bindings connect request fields to canonical artifacts, signature domains, and the verifier checks that consume them.
Contract bindings connect request fields to canonical artifacts, signature domains, and the verifier checks that consume them.Contract bindings connect request fields to canonical artifacts, signature domains, and the verifier checks that consume them.

A later layer authenticates references to earlier layers; it does not replace their semantic checks. In particular, a valid evidence signature proves that one key signed the manifest, not that artifact identities, scores, policy, or runtime declarations are true. Independent verification supplies and checks those acceptance anchors.

Semantic validation

Schema validity is only the first layer. The verifier also recomputes:

  • fixed path and closed-inventory rules;
  • canonical bytes and file digests;
  • artifact, schedule, runtime, observation, and request cross-bindings;
  • record order, input digests, and observation-record digests;
  • exact-match or normalized-NLL scores, or explicitly authorized scorer- extension results replayed twice from authenticated text facts;
  • paired exact-match counts, exact McNemar probability, and the report-version-specific Newcombe interval, or deterministic paired replicates and interval endpoints for normalized NLL and scorer-extension deltas;
  • derived perplexity facts when tokenizer and paired token counts are comparable;
  • comparison means, threshold arithmetic, optional count/interval-width and exact-match side-accuracy qualification, and policy verdict; and
  • evidence signer and verifier signature bindings.

A custom reader that performs schema validation alone is not equivalent to invarlock verify.

Version and change discipline

A format identifier names one exact shape and interpretation. Additive fields are not silently accepted by closed objects. A breaking artifact change requires a new format identifier and explicit reader behavior. Runtime providers must also match ABI 1 exactly.

The v3 comparison report is the current writer format. Strict verification continues to reconstruct v2 and v1 reports under their original shapes and arithmetic when signed historical packs identify those formats; it never relabels a reconstructed report as another version.

Receipt v1 remains valid for existing evidence whose runtime does not require a request-level executable binding. A supplied request digest selects receipt v2. GGUF verification requires that external request anchor because the normalized request authorizes the exact llama.cpp binary, source, version, execution settings, and GGUF identity reconciled with provider evidence.

Core and first-party add-ins are released at the same package version. Provider add-ins declare the exact coordinated core release, while the provider ABI remains the runtime compatibility gate. See Release verification.