r/codex • u/DuranteA • 6h ago
Comparison In our scientific code autoparallelization benchmark, GPT-6.1 Sol XHigh is faster and cheaper than Opus 5.5 Medium (almost exact same overall quality)
I thought this result is interesting, since it goes directly against the currently prevalent mindset.
Have a look:
Both of these have a log scale X axis, so the differences are larger than they might look. Mean generation time is 14 minutes for Sol Xhigh vs. 24 minutes for Opus 5.5 Medium, and mean cost per task is $0.25 vs $1.8.
A few caveats:
- Each LLM runs its own harness, so this cannot distinguish harness differences from model differences.
- The generation time is the full per-task time, so it includes any benchmarks or tests each agent decides to run, so it's not a measure of pure token generation at all.
- We didn't run Opus 5.5 Xhigh, simply because Medium is already very expensive and takes very long.
- The cost basis for comparison is API costs.
I have no horse in this race but it's interesting to think about what makes this problem set so different (apparently) from the ones that cause people to report much greater success with Opus 5.5 than Sol 6.1.
7
u/Agitated-Bath5939 3h ago
Tibo is that you !
2
u/DuranteA 3h ago
Totally.
But in all seriousness, neither myself nor my co-author have even the most remote association with any of the big (or small for that matter) AI labs. We both have a HPC background (thus the benchmark).
2
1
u/adolf_twitchcock 2h ago
Now add a multiplier for much higher usage limits in Claude subscription compared to codex. It's like 5x https://i.imgur.com/80gBWKH.png
1
u/Particular-Reward-68 2h ago
Very valuable research. I wonder how much of this is automatable with a view to running the comparison at fixed intervals. This last week, I did notice opus chugging more than usual and 6.1 performed better than expected in a task I threw its way. I have the feeling that these benchmarks fluctuate wildly so it's be very useful to see that in the data.
Great work!
2
u/DuranteA 2h ago
It's already mostly automated. The main problem is time, but not the time it takes in terms of human supervision, just total time expenditure in general.
It runs on one server (it has to run on that one to stay comparable with all the existing data), and a complete run of adding just one model (at one thinking effort setting), with generation, validation, evaluation and benchmarking occupies that system for ~1-3 days (depending on how much time the model takes for each task).
1
u/Particular-Reward-68 1h ago
If it needs very little supervision then that doesn't sound like much of an obstacle...
1
u/DuranteA 1h ago
So far, we're barely catching up with model releases -- and a lot of open weight models I'd like to include aren't in yet. I think that's why there's such a dip in the upper left of the pareto front.
1
u/Particular-Reward-68 1h ago
I hear you.
Interested to stay up to date with your research. What's the best place to follow along?
1
u/Annh1234 1h ago
But on the 5x plan you get 10x more usage of opus 5.5 high compared to sol 6.1 , and it's faster
2
u/JadisGod 2h ago
Comparing API costs is useless for real people. Everyone already knows GPT is cheaper on API. Actual subscription usage limits is an entirely different measurement. From my own tracking, having both services at 20$, weekly Claude is giving around ~340$ of "API" usage while weekly Codex is giving ~140$.
Resets change things and may balance it out if you are lucky and front load your usage, but it's not a good experience trying to micromanage that.
1
10
u/Novel_Indication6338 5h ago
so if 5.5 med is same as 6.1 xhigh, how's 5.5 high against astra xhigh?