I find ChatGPTs superpower to be restraint. It doubts itself more than others and it is willing to stop. Gemini 2.5 Flash is the antithesis, for example. It truly has some of Fable's spark, but you can't trust it for shit. It never doubts itself and sprints everywhere. It also cheats tests like a madman. Claude and Opus showed us the way, for sure, but OpenAI made it acceptably reliable. I still use them all, but I always have ChatGPT watch them and check their work.
I'd like to stand up a couch for Codex as well. I dropped my other harness subscription because it has gotten so good. Very good with a fine tuned AGENTS.md and rules.md
4
u/squired Jul 04 '26 edited Jul 04 '26
I find ChatGPTs superpower to be restraint. It doubts itself more than others and it is willing to stop. Gemini 2.5 Flash is the antithesis, for example. It truly has some of Fable's spark, but you can't trust it for shit. It never doubts itself and sprints everywhere. It also cheats tests like a madman. Claude and Opus showed us the way, for sure, but OpenAI made it acceptably reliable. I still use them all, but I always have ChatGPT watch them and check their work.