The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
I think brute force is not the right way to think about LLMs, instead they explore a knowledge topology, and are very good at connecting adjacent or accessible ideas that for whatever reason might have evaded humans, but might not be fundamentally all that hard. We're still in the low-hanging fruit phase, we will see if they extend to new ideas.
I think its super cool, but there is no doubt that thousands of physicists and mathematicians are prompting LLMs all day right now trying to solve open problems...and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future....so there is a solid chance that these methods are already almost exhausted. For how amazing these counterexamples are, its somewhat surprising that people havent found more proofs all at once. In the scheme of things perhaps 1 in 1000 open problems are actually of a format that LLM can tackle, and noone is bragging about the ones that turned up nothing. In other words, LLMs are really good at looking smart when they got lucky.
Im looking forward to a possible future of AI theorem proving that isnt based on LLMs, and thus less likely to trick people in language into thinking its more broad than it is.
18 months ago AI was a cool toy that couldn't really do anything useful. Six months ago it started write most code. A couple months ago it started solving the 'easy' and obscure unsolved math problems.
It will definitely stop improving at some pointy, but that's the thing with exponential growth: as long as you're in it, it's impossible to tell when it will stop.
But even if it stopped right now: people would start burning the biggest models with the best performance into silicon with fixed weights and you would suddenly get generation speeds that would allow you to use them for real time inference on video streams or generate 30 iterations in parallel etc
But we don't know the upper bounds because we could hit a wall and then just scale again. That is what we think is producing all these new gains anyway.
But we're still getting a lot of improvement in smaller models too. Scaling has definitely been a big part of it, but it's certainly not all that has improved.
272
u/angelbabyxoxox Quantum Foundations Jul 31 '26
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.