Back to blog

GPT-6 Sol vs GPT-6.1 Sol on 200 coding tasks

GPT-6.1 used less time and fewer observed tokens on tasks both models passed. A paired comparison of correctness, efficiency and repeat consistency.

Updated 6 min readInvarLock Team
Two routes of different lengths reach equally emphasized inspection windows beside a shared measuring span. The drawing illustrates work and quality as separate questions, not measured distances.

OpenAI's GPT-6.1 Sol announcement reports stronger coding performance at lower cost. For someone choosing between GPT-6 Sol and GPT-6.1 Sol, a useful next question is what changes when both models tackle the same problems: how many they solve, how long they take, and whether their results hold up across repeated attempts.

We compared them on 200 coding tasks, both using High reasoning effort. GPT-6.1 passed 151 tasks, compared with 143 for GPT-6. On the 136 tasks both models passed, GPT-6.1 used about 27% less recorded time and 62% fewer recorded tokens—the units used to count text handled by a model. Less time and fewer tokens on successful fixes was the clearest finding; the accuracy difference remained uncertain.

Looking at each task also reveals what an overall score can hide: GPT-6.1 gained fifteen successes and lost seven. Keeping those results with verifiable evidence gives a team a practical basis for evaluating an upgrade.

How we compared the models

Both models received the same problems from Python software projects, with the same starting code, instructions, tools and time limits. A task counted as a pass when the model's fix satisfied the specified tests for the bug and for preserving existing behavior. Both models used High, the setting that controls how much reasoning effort they can use.

The tasks came from SWE-rebench V2 and covered developer tools, scientific computing and application software. The comparison includes 200 tasks from 152 projects with complete test results for both models. Comparing both models on each task is what makes this a paired evaluation. This is a separate test set and setup from the DeepSWE benchmark in OpenAI's announcement.

The runner saved the models' work and test results. InvarLock linked the saved results and evaluation rules in a digitally signed evidence package, giving another reviewer a way to verify which records support the comparison. We also used a separate program to reproduce the study's calculations from those saved records. The methods report provides the exact settings and verification details.

Successful fixes took less time and fewer tokens

To compare the work required for a successful fix, we focus on the 136 tasks both models passed. GPT-6.1 used about 27% less time and 62% fewer tokens on those tasks, after accounting for the mix of task types.

The 95% uncertainty ranges correspond to roughly 20–33% less time and 58–66% fewer tokens. These ranges express how precisely the study estimates each difference.

Diagram
On 136 tasks both models passed, GPT-6.1 used about 73% as much recorded time and 38% as many recorded tokens as GPT-6: reductions of 27% and 62%.

Less time and fewer tokens on 136 shared successes

Recorded time: 27% less
GPT-6.1 used about 73% as much time as GPT-6. The 95% uncertainty range is about 20–33% less.
Recorded tokens: 62% fewer
GPT-6.1 used about 38% as many tokens as GPT-6. The 95% uncertainty range is about 58–66% fewer.

Both models passed these same 136 tasks. The comparison accounts for the mix of task types; tokens come from the latest recorded counters.

What to noticeEach measure sets GPT-6's usage to 100%. GPT-6.1's shorter bars show less recorded time and fewer tokens for the same successful tasks. Black lines show the uncertainty range around each estimate.

GPT-6.1 also used about 35% fewer tool actions, such as reading files, searching and editing code, while requesting about 22% more tests. Successful fixes involved fewer actions overall and more requests to test the code. The difference in time spent waiting for tests was less clear.

Token totals come from the latest counters in the saved output. The full six-measure chart and report provide the detailed breakdown, including measurements across attempts that failed or lacked a complete test result.

Eight more tasks solved overall

GPT-6.1 passed 151/200 tasks (75.5%), compared with 143/200 (71.5%) for GPT-6. Looking at the same task on both sides shows how that difference arose:

Diagram
Outcomes across 200 tasks: both models passed 136, only GPT-6.1 passed 15, only GPT-6 passed 7, and both failed 42.

The same 200 tasks, four outcomes

Both passed: 136
Successful fixes shared by both models.
Only GPT-6.1 passed: 15
Gains when moving to GPT-6.1.
Only GPT-6 passed: 7
Losses when moving to GPT-6.1.
Both failed: 42
Tasks neither model solved.

Fifteen gains minus seven losses means eight more tasks solved overall. The estimated accuracy difference remains uncertain.

What to noticeMost tasks had the same outcome with either model. The 15 gains and 7 losses add up to eight more tasks solved by GPT-6.1; the accuracy difference remains uncertain.

After accounting for the mix of task types, the estimated accuracy gain is 4.25 percentage points—roughly four extra passes per hundred tasks. Its 95% uncertainty interval runs from −0.10 to +8.96 points. Because that range includes zero, the accuracy difference remains uncertain.

The fifteen gains and seven losses give a team specific cases to review. Some lost fixes may matter more to its work than the higher total score, making them useful checks for a model upgrade.

What happened when we repeated the tasks?

We examined three attempts per model on 30 tasks, ten from each task category. These tasks were selected after the main comparison because both models had a complete first test result, whether a pass or a failure.

Table 1: What happened when we repeated the tasks?. Data table with columns: Across three attempts, GPT-6, GPT-6.1
Across three attemptsGPT-6GPT-6.1
Passed every time1821
Failed every time86
Switched between pass and fail43

Scroll horizontally to see every column.

GPT-6 produced the same pass-or-fail result on 26/30 tasks, and GPT-6.1 on 27/30. That includes tasks that failed every time. The small difference remains uncertain, but the repeated attempts show which successes held up and which tasks produced changing outcomes. The report contains the full consistency analysis.

What this means when choosing a model

For a team choosing between these models, GPT-6.1 is a promising candidate for reducing time and token use on successful fixes. The seven tasks it lost are useful starting points for checking whether a switch preserves the behavior that matters to the team.

InvarLock makes the comparison easier to hand over for review: the saved results are tied to the evaluation rules, and verification checks that the evidence matches those records. Another person can inspect the gains and losses behind the headline and use the same approach for a focused comparison on their own software projects.

The report and methods, numerical results, task outcomes and pilot guide provide the detail for inspecting or adapting this comparison. Provenance and hashes identify the public files. Full signed verification requires the separately retained private archive.

Limitations

  • The study covers selected, reused Python coding tasks with fixed tests. The repeat findings apply to the selected 30-task subset.
  • Eight attempted tasks lacked complete results for both models: 11 individual results were missing, seven for GPT-6 and four for GPT-6.1. Missing results may depend on the model and workload; the methods report examines their possible effect.
  • Efficiency measurements cover recorded time, token counts and tool activity on 136 shared successes. Final bills, computing resources and energy use fall outside the measurement scope.

Sources

Website documentation explains the maintained workflow. This article's numerical evidence is the retained 200-pair coding study, its 30-task repeat supplement and its signed evidence index. The methods and provenance identify the pinned implementation and recipient verification exercise; the dataset and paper describe the source inventory.

More in Research Note

Explore nearby related posts.