Back to blog

The Dataset Is Part of the Signed Decision

A dataset name does not identify the cases behind a result. A signed comparison needs the exact inputs, expected answers, selection, and order—and a separate argument for why those cases matter.

6 min readInvarLock Team
Three distinct record strips are held in a fixed order by one blue binding that continues to the attached result.

Two reports say they evaluated the same dataset. They may still have used different cases, different prompts, or different expected answers. Even the same number of records does not establish that the comparisons asked the same question.

A reviewer needs the exact inputs behind the result. In InvarLock's native signed comparisons, those inputs form a schedule: the ordered cases, their input material, and their expected answers. The evidence links that schedule to both sides of the comparison and to the verification receipt.

This makes a specific claim inspectable: these are the cases behind this result. It leaves a separate question for the decision owner: are these the right cases for the intended use?

A name describes the source; a digest identifies the material

Consider a dataset delivered as a local JSONL file—one record per line. The evaluation request specifies its exact file digest, the fields to read, and any limit on the number of records. A digest is a hash of defined bytes that lets another reader check whether the material matches.

The native run path checks the source file's digest before preparing the schedule. It preserves record order and, when a limit is supplied, selects that exact prefix. The schedule retains the selected inputs and expected answers. The normalized request separately preserves the source path, field mapping, and limit.

There are three useful identities to distinguish:

  • Identity
    Source digest
    Material
    Exact local dataset file bytes
    Example change
    Edit the source file
    Review use
    Detect a changed source under the same filename
  • Identity
    Input digest
    Material
    One case's ordered input parts
    Example change
    Change a prompt
    Review use
    Check the same input material on both sides
  • Identity
    Schedule digest
    Material
    Dataset identity, ordered cases, and expected answers
    Example change
    Reorder cases or correct an answer
    Review use
    Identify the exact comparison inputs

“Canonical” means the contract serializes the data in a defined way before hashing it. The schedule’s object-key order and whitespace are normalized; record-array order remains meaningful. The source-file digest still depends on every original byte. A raw file hash and a canonical schedule hash therefore identify different things and should not be substituted for each other.

The evaluation request guide describes the local-file path. Import mode supplies a canonical schedule directly. Both paths bind an exact schedule, but a local file's preparation details should not be assumed for every imported dataset.

Follow one retained schedule into its receipt

The published Qwen3.5 9B BF16-to-Q5_K_M comparison provides a concrete example. Its retained schedule names TIGER-Lab/MMLU-Pro/qwen35-no-think, identifies the split as test-balanced-400, and contains 400 ordered records. The first record is mmlu_pro_00075.

Those labels help a reader understand the selection. The schedule itself supplies the precise material: record IDs, ordered input parts, input digests, and expected answers. Recomputing its canonical digest gives the same schedule anchor recorded in the separately signed verification receipt. All seven native public-pack schedules in the inspected snapshot contain 400 records, and each matches its own receipt's schedule anchor. That is an identity check across retained artifacts, not seven new model evaluations.

The qualification-suite manifest records another useful distinction: source-file hashes, canonical schedule digests, selection settings, and distributions are separate fields. A benchmark label cannot replace those details, and a suite manifest should not be assumed to identify every derived, prompt-specific schedule.

The original GGUF case study explains the model comparison and its uncertainty. Here, the narrower point is that a reviewer can follow the exact input schedule into the receipt without inferring it from the article's dataset name.

Changing the cases changes the comparison

Changing an expected answer or swapping two records changes the canonical schedule digest. A different prompt also changes that record's input digest. A recipient who independently pinned the original schedule can detect that the submitted evidence identifies a different comparison.

For native evidence-pack-v1 comparisons, both sides must contain every scheduled record in the declared order, with matching input digests and successful observations. Verification does not sort, intersect, or drop records to manufacture a passing subset. Matching counts alone are insufficient. The pairing and replay guide describes the checks in detail.

This ordering rule has a specific scope. Captured-record comparisons use a separate contract that pairs by unique record ID; changing export order alone need not change the paired result. The complete-run identity still records the submitted run. Reviewers must use the rule for the evidence format in front of them.

A corrected answer key or a better selection can be a legitimate method change. Preserve the earlier evidence, identify the change, and produce evidence for the revised comparison. The old signed result cannot silently acquire the new dataset's meaning.

A matching digest cannot choose the right cases

A carefully selected set of easy cases can be hashed and signed just as successfully as a demanding one. Digests establish identity. They do not establish relevance, representativeness, or when the selection decision was made.

Before seeing the subject's results, the decision owner should define the target behavior, fix the cases and expected answers, and retain or review the schedule digest. If failures motivate a later selection change, that change needs an explicit explanation and a new evidence trail. Verification alone cannot establish that the evaluator avoided choosing favorable cases after seeing the answers.

For a handoff, ask for the exact schedule and its expected digest alongside the result. Then ask why those cases cover the behavior you need to review. The finite-schedule article explains the limits of the resulting claim, and the recipient checks explain how the receiving organization supplies its own expectations and authority.

Limitations

  • The concrete example and seven-schedule check use retained native public packs. Their 400-record schedules do not establish representative coverage of another workload or a general dataset ranking.
  • The source audit recomputed schedule and input digests and compared retained receipt anchors. It did not run new inference, repeat full signature verification, or replay the comparison reports.
  • Exact source-byte and prefix-selection behavior describes the native local-JSONL run path. Imported schedules and captured-record comparisons have their own contracts.
  • Matching hashes and signatures do not prove genuine original execution, unbiased case selection, wider model quality, or permission to deploy.

Sources

Website documentation explains the maintained workflow. The contracts were inspected in InvarLock v0.16.3, commit 2b2b64d2f668ab4768ba044df7883f86e5e0a835. The seven public schedules and the Qwen3.5 example retain their original v0.15.0 evidence basis, commit 4f036ffe99cfaa82019c90ea527b8c4291a091b3; those schedules are unchanged in v0.16.3. The article uses retained artifacts and a local digest check, with no new model inference or full verification replay.