GPT-6 Sol vs GPT-6.1 Sol on 200 coding tasks
GPT-6.1 used less time and fewer observed tokens on tasks both models passed. A paired comparison of correctness, efficiency and repeat consistency.
OpenAI's GPT-6.1 Sol announcement reports stronger coding performance at lower cost. For someone choosing between GPT-6 Sol and GPT-6.1 Sol, a useful next question is what changes when both models tackle the same problems: how many they solve, how long they take, and whether their results hold up across repeated attempts.
We compared them on 200 coding tasks, both using High reasoning effort. GPT-6.1 passed 151 tasks, compared with 143 for GPT-6. On the 136 tasks both models passed, GPT-6.1 used about 27% less recorded time and 62% fewer recorded tokens—the units used to count text handled by a model. Less time and fewer tokens on successful fixes was the clearest finding; the accuracy difference remained uncertain.
Looking at each task also reveals what an overall score can hide: GPT-6.1 gained fifteen successes and lost seven. Keeping those results with verifiable evidence gives a team a practical basis for evaluating an upgrade.
How we compared the models
Both models received the same problems from Python software projects, with the same starting code, instructions, tools and time limits. A task counted as a pass when the model's fix satisfied the specified tests for the bug and for preserving existing behavior. Both models used High, the setting that controls how much reasoning effort they can use.
The tasks came from SWE-rebench V2 and covered developer tools, scientific computing and application software. The comparison includes 200 tasks from 152 projects with complete test results for both models. Comparing both models on each task is what makes this a paired evaluation. This is a separate test set and setup from the DeepSWE benchmark in OpenAI's announcement.
The runner saved the models' work and test results. InvarLock linked the saved results and evaluation rules in a digitally signed evidence package, giving another reviewer a way to verify which records support the comparison. We also used a separate program to reproduce the study's calculations from those saved records. The methods report provides the exact settings and verification details.
Successful fixes took less time and fewer tokens
To compare the work required for a successful fix, we focus on the 136 tasks both models passed. GPT-6.1 used about 27% less time and 62% fewer tokens on those tasks, after accounting for the mix of task types.
The 95% uncertainty ranges correspond to roughly 20–33% less time and 58–66% fewer tokens. These ranges express how precisely the study estimates each difference.
Less time and fewer tokens on 136 shared successes
- Recorded time: 27% less
- GPT-6.1 used about 73% as much time as GPT-6. The 95% uncertainty range is about 20–33% less.
- Recorded tokens: 62% fewer
- GPT-6.1 used about 38% as many tokens as GPT-6. The 95% uncertainty range is about 58–66% fewer.
Both models passed these same 136 tasks. The comparison accounts for the mix of task types; tokens come from the latest recorded counters.
GPT-6.1 also used about 35% fewer tool actions, such as reading files, searching and editing code, while requesting about 22% more tests. Successful fixes involved fewer actions overall and more requests to test the code. The difference in time spent waiting for tests was less clear.
Token totals come from the latest counters in the saved output. The full six-measure chart and report provide the detailed breakdown, including measurements across attempts that failed or lacked a complete test result.
Eight more tasks solved overall
GPT-6.1 passed 151/200 tasks (75.5%), compared with 143/200 (71.5%) for GPT-6. Looking at the same task on both sides shows how that difference arose:
The same 200 tasks, four outcomes
- Both passed: 136
- Successful fixes shared by both models.
- Only GPT-6.1 passed: 15
- Gains when moving to GPT-6.1.
- Only GPT-6 passed: 7
- Losses when moving to GPT-6.1.
- Both failed: 42
- Tasks neither model solved.
Fifteen gains minus seven losses means eight more tasks solved overall. The estimated accuracy difference remains uncertain.
After accounting for the mix of task types, the estimated accuracy gain is 4.25 percentage points—roughly four extra passes per hundred tasks. Its 95% uncertainty interval runs from −0.10 to +8.96 points. Because that range includes zero, the accuracy difference remains uncertain.
The fifteen gains and seven losses give a team specific cases to review. Some lost fixes may matter more to its work than the higher total score, making them useful checks for a model upgrade.
What happened when we repeated the tasks?
We examined three attempts per model on 30 tasks, ten from each task category. These tasks were selected after the main comparison because both models had a complete first test result, whether a pass or a failure.
| Across three attempts | GPT-6 | GPT-6.1 |
|---|---|---|
| Passed every time | 18 | 21 |
| Failed every time | 8 | 6 |
| Switched between pass and fail | 4 | 3 |
Scroll horizontally to see every column.
GPT-6 produced the same pass-or-fail result on 26/30 tasks, and GPT-6.1 on 27/30. That includes tasks that failed every time. The small difference remains uncertain, but the repeated attempts show which successes held up and which tasks produced changing outcomes. The report contains the full consistency analysis.
What this means when choosing a model
For a team choosing between these models, GPT-6.1 is a promising candidate for reducing time and token use on successful fixes. The seven tasks it lost are useful starting points for checking whether a switch preserves the behavior that matters to the team.
InvarLock makes the comparison easier to hand over for review: the saved results are tied to the evaluation rules, and verification checks that the evidence matches those records. Another person can inspect the gains and losses behind the headline and use the same approach for a focused comparison on their own software projects.
The report and methods, numerical results, task outcomes and pilot guide provide the detail for inspecting or adapting this comparison. Provenance and hashes identify the public files. Full signed verification requires the separately retained private archive.
Limitations
- The study covers selected, reused Python coding tasks with fixed tests. The repeat findings apply to the selected 30-task subset.
- Eight attempted tasks lacked complete results for both models: 11 individual results were missing, seven for GPT-6 and four for GPT-6.1. Missing results may depend on the model and workload; the methods report examines their possible effect.
- Efficiency measurements cover recorded time, token counts and tool activity on 136 shared successes. Final bills, computing resources and energy use fall outside the measurement scope.
Sources
Website documentation explains the maintained workflow. This article's numerical evidence is the retained 200-pair coding study, its 30-task repeat supplement and its signed evidence index. The methods and provenance identify the pinned implementation and recipient verification exercise; the dataset and paper describe the source inventory.
More in Research Note
Explore nearby related posts.
Research note
Five Recipient Checks Before a Technical Result Can Be Accepted
Technical verification is an input to acceptance, not the decision itself. Five checks keep identity, trust, policy, uncertainty, and organizational authority separate.
Research note
Which Answers Changed Behind the Score?
One retained comparison has 54 changed answers, 27 changed correctness outcomes, and seven net additional correct answers. Each count answers a different review question.
Research note
Three GGUF Deployment Comparisons Under One Review Standard
Three BF16-to-Q5_K_M comparisons pass the same declared policy, each with its own evidence and receipt. The repeatable result is the review standard.