Use the study evidence and plan your own comparison
This guide accompanies GPT-6 Sol vs GPT-6.1 Sol on 200 coding tasks. Use it to inspect the published measurements, verify the saved evidence if you hold the private archive, or design a small comparison on your own software projects.
Explore the public results
The report and methods explain how the tasks were graded, how the estimates were calculated and where results are missing. The results JSON contains the numerical summaries. The outcome CSV lists all 600 allocated attempt identities, including completed attempts, missing grades and unexecuted slots.
Start with the 200 tasks that have a valid result for both models. Compare shared successes, gains, losses and shared failures; then inspect the separate 30-task repeat subset. A valid failing result belongs in the comparison. A missing result remains unknown. The report explains the CSV columns and how the repeated attempts relate to the primary comparison.
The provenance manifest records the public files' hashes and their source snapshot. Those hashes let you check that a downloaded file matches the published version. The public files are readable, unsigned summaries; full signed verification uses the separately retained private archive.
Verify the archive you have received
Recipients who already hold the complete study archive can use the verification helper with the delivery anchors. The helper is specific to this study's files and allocation. Use Python 3 on macOS or Linux, with cryptography installed.
Before running it, confirm the expected archive hash and signer fingerprint through a trusted channel independent of the archive delivery. The anchors identify the exact helper and archive expected for this study. Keep the helper, anchors and archive available locally, and choose extraction and output paths that do not already exist. The output file must be outside the extraction directory.
python3 -B verify_paired_delivery.py study-delivery.tar \
--anchors delivery-anchors.json \
--extract-to fresh-study-extraction \
--output fresh-verification-result.json
The helper checks its own identity and the archive bytes, rejects unsafe archive members, authenticates signed members, and reconstructs the saved statistical results. It uses the private archive, rather than the public CSV and JSON, and makes no model calls or repository test runs.
The retained recipient exercise authenticated 504 delivery members and 397 scientific members, reconstructed eight primary and three repeat results, and reused 36 completed archive replay receipts. This publication reports that saved exercise. A recipient can run the command against their delivered copy and inspect the resulting verification file.
Plan a focused comparison on your own tasks
- Choose the decision. State which model change you are considering and which behaviors matter. Define a pass using concrete checks. If the comparison will support a release decision, agree on the acceptance requirements beforehand.
- Fix the tasks and settings. Select tasks before seeing comparative results. Save the starting code, prompts, tests, model versions, reasoning settings, tools and time limits. Give both models the same conditions and preserve the task order and analysis method.
- Check the full evaluation path. Use a small qualification run to confirm that answers, test results, timing and token counters can be captured and retained. Include output sizes and failure conditions relevant to your workload. Confirm that the available time and storage cover the planned work and its evidence.
- Run and retain each attempt. Keep passing fixes, failing fixes and missing results separately. Set the stopping rule in advance. Preserve every attempt and declare retry or exclusion rules before they are needed. If a shared capture or grading defect appears, resolve its effect before continuing dependent work.
- Compare the same tasks. Show gains and losses alongside total pass counts. Compare efficiency on tasks both models solved, and provide all-attempt measurements separately. Account for related tasks from the same project when estimating uncertainty. If repeatability matters, define a repeat panel in advance and keep all its attempts.
- Make the result reviewable. Use InvarLock's captured-results workflow to tie saved results to the evaluation rules. Deliver the signed evidence with separately trusted verification expectations and a readable report. Have the receiving reviewer verify the delivered package and inspect the cases that matter to the decision.
The study's helper verifies this saved package; a new pilot needs its own task configuration and capture checks. The study recorded an unresolved capture limitation, documented in the methods report, so its collector should be qualified for a new workload before reuse.
What this comparison can help you decide
On the 136 tasks both models passed, GPT-6.1 used less recorded time and fewer tokens. The fifteen gains and seven losses identify cases worth inspecting when planning an upgrade. The accuracy difference and the repeat-consistency difference remained uncertain. Use these findings to focus your own comparison on the behavior and constraints that matter to your team.
Sources
Website documentation explains the maintained workflow. This guide uses the retained coding-study measurements and verification exercise. The archive command requires the separately supplied private evidence.