Wide tables scroll horizontally. Focus a table and use the arrow keys to see more columns.
GPT-6 Sol vs GPT-6.1 Sol: report and methods
This report provides the methods and supporting measurements for the 200-task model comparison. Both models attempted the same Python coding tasks under the same settings. A follow-up examined three attempts per model on 30 tasks.
Start with the findings below, then use the sections on task coverage, grading, missing results and repeated attempts to examine how the comparison was calculated. The pilot guide explains how to inspect the public files, verify the separately retained archive, or plan a comparison on your own tasks.
Main findings
GPT-6 Sol passed 143/200 fully graded tasks (71.5%); GPT-6.1 Sol passed 151/200 (75.5%). The unweighted difference is 4 percentage points. Adjusting for the planned mix of task categories gives +4.246411 percentage points, with a 95% uncertainty interval of −0.099768 to +8.957760 points. Because the interval includes zero, the accuracy difference remains uncertain. The interval accounts for tasks from the same repository being related; the analysis method is described below.
Both passed 136 tasks; only GPT-6.1 passed 15; only GPT-6 passed 7; both failed 42. The 200 pairs include both passing and failing fixes, provided both models have a valid test result. Tasks with a missing result are accounted for separately below.
On the 136 tasks both passed, GPT-6.1 used about 27% less observed wall time and 62% fewer last-observed tokens. Each ratio below divides GPT-6.1's measurement by GPT-6's: 1 means equal usage, and 0.734 means about 27% less. The intervals show uncertainty for each measurement separately:
| Metric | GPT-6.1/GPT-6 | 95% interval | Paired tasks |
|---|---|---|---|
| Wall time | 0.734305 | 0.674578–0.798436 | 136 |
| Active elapsed | 0.620384 | 0.569506–0.675813 | 136 |
| Observed tokens | 0.378583 | 0.342906–0.417649 | 136 |
| Recorded tools | 0.651791 | 0.614580–0.691565 | 136 |
| Test requests | 1.223837 | 1.129942–1.329841 | 136 |
| Test wait | 1.087624 | 0.961503–1.229420 | 136 |
These ratios summarize the paired task measurements with geometric means, weighted by the planned task categories. They cover the 136 tasks both models passed and for which the measurements are positive. Test requests are included in tool requests. Active elapsed includes time used by the evaluation system; tokens are the latest recorded counters. Final bills, model compute and energy were outside the measurement scope.
Task coverage and counting rules
| Quantity | Count |
|---|---|
| Frozen tasks / repository clusters | 228/168 |
| Attempted primary tasks | 208 |
| Fully graded primary pairs / their repositories | 200/152 |
| Attempted tasks with at least one unknown grade | 8 |
| Retained attempts | 536 |
| Known individual grades / unknown individual grades | 525/11 |
| Observed usage records | 535 |
| Unattempted primary calls / calls held at setup | 36/4 |
| Unexecuted extra repeat calls | 24 |
| Full allocated calls | 600 |
The allocation contains 600 attempt identities: 536 retained attempts, 36 untouched primary calls, four calls held at setup and 24 unexecuted repeat calls. The full allocation and original 36-task repeat panel were not completed. Thus 536+36+4+24=600. The 120 follow-up attempts do not change primary pair counts. The retained dataset includes every failure; neither unknown grades nor unexecuted calls are imputed as failures or removed. The 200 fully graded tasks contain 74 developer, 66 scientific and 60 application tasks. The full 228-task empirical weights remain 81/228, 72/228 and 75/228 rather than being refitted to this subset. At most two tasks share a repository.
Design and grading
Models: gpt-6-sol and gpt-6.1-sol, both using High reasoning effort. Exact issue-derived prompts, source revisions, protected tests and evaluation software identities are bound in the signed index. Both models had a configured allowance of 600 active elapsed seconds excluding synchronous native-test wait, with a 2,200-second wall bound. The final configuration used no tool/test count quotas; metered test-byte limits and native safety bounds remained. Up to 14 attempts ran concurrently.
The bug-fix and preservation requirements were fixed before answers. Passing means satisfying those requirements, not every possible repository behavior. Grade validity, patch quality, delivery, command health, test capture, termination and cleanup are kept distinct. A confirmed stopped and captured patch can retain a valid failing grade while delivery stays unsuccessful.
Primary collection ended after the completed batch reaching 200 fully graded pairs. The stopping rule depended on grade availability, not comparative results. Tasks were not replaced and prior attempts were not retried to reach 200. The output-capture limitation described below remained unchanged.
Primary all-attempt observed measurements
There are 208 attempts per model. Means and medians include valid failures and attempts with unknown quality. Observed sums of overlapping attempt clocks are not project wall duration.
| Metric | GPT-6 coverage | Mean | Median | GPT-6.1 coverage | Mean | Median |
|---|---|---|---|---|---|---|
| Wall time (s) | 208/208 | 230.37 | 193.05 | 208/208 | 170.71 | 142.86 |
| Active elapsed (s) | 208/208 | 176.97 | 140.19 | 208/208 | 114.50 | 78.72 |
| Test wait (s) | 208/208 | 53.41 | 37.86 | 208/208 | 56.20 | 44.93 |
| Observed total tokens | 207/208 | 674,299.40 | 425,863.00 | 208/208 | 221,383.29 | 144,098.00 |
| Recorded tool requests | 208/208 | 29.83 | 25.50 | 208/208 | 19.22 | 16.00 |
| Test requests | 208/208 | 1.56 | 1.00 | 208/208 | 1.86 | 2.00 |
Observed token totals are 139,579,976 across 207 GPT-6 records and 46,047,725 across 208 GPT-6.1 records. The former sum is incomplete. The allocated outcome measurements retain observed totals and last-consumed summary totals. Component counters remain in the private signed evidence; they are not in this curated CSV. A missing measurement stays unknown. Cache/reasoning subsets are never added again to containing totals.
Expanded all-attempt observations
There are 268 retained attempts per model: 208 primary plus 60 additional repeats. This descriptive accounting includes every failure and unknown; it is not a pooled independent-task accuracy estimate.
| Metric | GPT-6 coverage | Mean | Median | GPT-6.1 coverage | Mean | Median |
|---|---|---|---|---|---|---|
| Wall time (s) | 268/268 | 229.57 | 186.58 | 268/268 | 171.69 | 142.67 |
| Active elapsed (s) | 268/268 | 173.64 | 129.85 | 268/268 | 112.69 | 75.79 |
| Test wait (s) | 268/268 | 55.93 | 38.76 | 268/268 | 58.99 | 46.38 |
| Observed total tokens | 267/268 | 664,279.49 | 428,507.00 | 268/268 | 222,567.51 | 150,567.00 |
| Recorded tools | 268/268 | 29.68 | 25.00 | 268/268 | 19.14 | 16.00 |
| Test requests | 268/268 | 1.55 | 1.00 | 268/268 | 1.88 | 2.00 |
GPT-6: 261 known grades, 7 unknowns, 183 passing grades, 266 observed terminals and 268 confirmed cleanup observations out of 268 attempts. These denominators are not substituted for primary paired accuracy.
GPT-6.1: 264 known grades, 4 unknowns, 201 passing grades, 266 observed terminals and 267 confirmed cleanup observations out of 268 attempts. These denominators are not substituted for primary paired accuracy.
Tool and native-test correspondence
Across all 536 retained attempts, GPT-6 recorded 7,953 tool requests: 3,888 reads, 1,136 searches, 2,513 patch actions and 416 test requests. GPT-6.1 recorded 5,130: 2,961 reads, 756 searches, 910 patch actions and 503 test requests. Recorded successes/errors-or-refusals are 6,312/1,641 and 4,444/686. The error-or-refusal grouping does not identify a cause; exact original error strings, refusal and truncation fields remain in the signed records.
Full raw requested counts are 7,953 and 5,132, reconstructed from all 536 authenticated call mappings. Two original extra GPT-6.1 read/patch requests retain unknown execution. All 919 recorded test requests join to saved responses and native captures: 916 test-level grades are complete and three are incomplete. These are operations, not distinct behaviors or final patch grades; no incomplete result was imputed or rerun.
Primary reliability and missingness
| Observation | GPT-6 | GPT-6.1 |
|---|---|---|
| Known grades / attempts | 201/208 | 204/208 |
| Unknown grades | 7 | 4 |
| Passed among known grades | 144/201 | 155/204 |
| Terminal observed / attempts | 207/208 | 206/208 |
| Cleanup confirmed / attempts | 208/208 | 207/208 |
The known-grade pass fractions above have different evaluability denominators and are not substituted for the paired accuracy estimate. Verified delivery is a composite of patch quality and delivery evidence; it is not the raw terminal-completion rate.
Delivery is known for 206 attempts per model: GPT-6 has 144 verified deliveries and 62 unsuccessful deliveries; GPT-6.1 has 155 and 51. Two deliveries per model remain unknown. On 206 jointly observed delivery pairs, the weighted difference is +5.64 percentage points, with an exploratory pointwise 95% interval of +0.71 to +10.63 points. Full 228-task delivery missing-outcome bounds are −4.82 to +14.47 points. This composite endpoint has a different denominator from the 200 quality pairs and does not establish accuracy superiority.
The 11 missing grades arose from provider availability, a timeout and incomplete native-test observations. Some attempts ended normally but still lacked the evidence needed for a valid grade. A capture-capacity limitation affected the study and remained in the final configuration; prior attempts were preserved without retries. These missing results remain unknown rather than being counted as failures.
Token analysis uses the latest counters present in the saved raw output. Earlier summary counters can contain less information and are retained separately in the CSV. Termination, interruption and cleanup observations are also kept distinct, so later observations do not overwrite the original attempt record.
Varying unknown outcomes in the 208 attempted tasks gives a domain-weighted difference range of +2.24 to +7.46 points for that fixed attempted set. Varying unknown and unexecuted outcomes across all 228 frozen tasks gives −7.02 to +15.35 points. These bounds are not confidence intervals. Missingness may depend on model and workload; the complete-case result is conditional.
How the results were verified
The study used InvarLock's captured-results workflow to evaluate saved records, sign the evidence and verify it against separately supplied expectations for the request, runs and signing identities. The repository runner and statistical analysis were specific to this study. The provenance manifest pins the source snapshot; the verification guide explains the recipient handoff.
A separate program reconstructed the study calculations from the saved evidence while reusing the completed verification receipts described below.
Statistical analysis and saved evidence
The fixed resampling seed is recorded in the numerical results. The weights are 81/228 developer tooling, 72/228 scientific computing and 75/228 application libraries/services. The declared 10,000 resamples draw repository clusters within each domain, moving all eligible tasks from a repository together. Intervals are pointwise and exploratory. A separate implementation independently reconstructs paired values, cluster draws, weighted aggregation and percentile interpolation. It reproduces quality, delivery and all six measured dimensions to the declared numerical tolerance. The number of completed pairs alone does not establish how small a difference the study can reliably detect.
Raw trace, native request correspondence and lifecycle receipt coverage reach all 536 retained attempts. The saved evidence was signed and independently checked. The final index's offline replay authenticates every package member, checks full allocation/condition/observation bindings and reproduces derivatives, reusing those completed archive replay receipts. It makes no model calls and runs no repository tests.
The final recipient exercise authenticated 504 delivery members and 397 scientific members, reconstructed eight primary and three conditional repeat derivatives, and reused 36 completed constituent replay receipts. No scientific/model/native run or archive lifecycle was repeated. This publication reuses that result; it does not repeat the recipient exercise.
The report and exports preserve the original model attempts and grades. The capture limitation described above still applies.
What three attempts reveal
The follow-up added 120 attempts: two more per model on 30 members of the original 36-task panel. Each selected task had a valid first grade for both models, including failing grades; ten tasks came from each domain. This selection occurred after primary collection and was based on joint evaluability, not comparative wins. Six omitted panel tasks retain 24 unexecuted repeat identities. Consistency conclusions therefore describe this selected subset, not the full original panel.
All 180 attempts in these 60 three-attempt series have valid final grades. Stable means that all three binary grades agree; it includes consistently failing tasks.
| Three-attempt result | GPT-6 | GPT-6.1 |
|---|---|---|
| Passed all three | 18 | 21 |
| Failed all three | 8 | 6 |
| Changed between pass and fail | 4 | 3 |
| Same grade all three times | 26/30 | 27/30 |
The weighted difference in the fraction of tasks with unchanged grades is +3.16 percentage points for GPT-6.1. Its pointwise 95% repository-within-domain interval is −10.66 to +17.24 points. This is uncertain consistency evidence, not an established advantage. The calculation keeps each task's complete paired series together, uses the original domain weights and 10,000 resamples, and defines the added descriptive endpoints separately from the frozen primary analysis.
The models passed 60/90 and 68/90 panel attempts respectively. These repeated attempts are dependent observations of 30 tasks; they do not add 90 independent tasks or change the 200-pair primary estimate. We retain each ordered grade and each repeat's measurements, without choosing the best attempt.
Work also varied within tasks. The median ratio of the slowest to fastest observed wall time across three attempts was 2.10 for GPT-6 and 1.76 for GPT-6.1. The corresponding observed-token ratios were 1.74 and 1.73. These descriptive spreads include all valid failures and require three positive observed values. They do not isolate model compute or establish a variability winner.
One GPT-6 repeat attempt reached the configured deadline and retained a valid failing grade on its captured patch. Its observed wall time was 716.92 seconds, including 94.67 seconds of test wait and 622.26 seconds of evaluator-inclusive active elapsed time. The configured allowance remained 600 active seconds, while the recorded active elapsed time includes evaluation-system overhead. The original failing result remains in the repeat series.
Earlier study collections used different conditions and are retained separately. Their attempts and repeat panels are not pooled with the 200-task comparison or the selected 30-task follow-up reported here.
Public files and provenance
This report is an unsigned editorial derivative of the finalized study. The article, results JSON, allocation CSV, pilot guide, scientific anchors and provenance manifest form the public reader package. The private signed archive is not included. These downloads do not supply full cryptographic replay or independent reconstruction from raw captures.
The CSV contains one row for every allocated identity, keyed by task, replica (0 for the first attempt, 1 or 2 for repeats) and model. An empty cell means unavailable or unexecuted, never zero. Quality and delivery are separate binary endpoints; unexecuted rows retain their disposition. Summary token totals and latest retained raw totals are separate columns. The 180 panel attempts include 60 first attempts already present in the primary set; do not add them again.
The JSON preserves primary comparisons, all-attempt metric summaries, the conditional repeat supplement and full-study accounting. No errors or unavailable outcomes were removed from the denominators. We omit raw operational error text, local paths, host details, source code, patches, credentials, workspaces and signing material. The complete archive is retained privately because it includes execution records and third-party source material beyond the public measurements.
Task identifiers derive from Nebius's SWE-rebench V2 dataset, whose dataset card declares CC BY 4.0. This publication attributes that inventory and publishes our measurements, not repository source, issue text, reference patches or test patches. Repository-specific rights are not replaced by the dataset's license. The tasks were reused; they are not described as unseen.
Limitations
- Complete-pair accuracy is conditional on joint evaluability; missingness may depend on model and workload. The full frozen allocation remains visible.
- Efficiency is conditional on 136 jointly successful tasks. Token counters do not establish final billing, compute or energy.
- Repeat consistency describes 30 selected tasks, including consistent failures, with uncertain comparative differences. It does not complete the 36-task panel.
- Public derivatives are unsigned; full saved-evidence verification requires the private signed archive. Study-operated signatures are not an outside auditor's endorsement or deployment approval.
Sources
Website documentation explains the maintained workflow. This report describes a separate frozen coding study and its conditional repeat supplement, with the source generation pinned in the provenance manifest.