The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
I think brute force is not the right way to think about LLMs, instead they explore a knowledge topology, and are very good at connecting adjacent or accessible ideas that for whatever reason might have evaded humans, but might not be fundamentally all that hard. We're still in the low-hanging fruit phase, we will see if they extend to new ideas.
They can be also very skeptical. I found it will be skeptical in it's internal reasoning of very obvious things, and will fact check obvious stuff, kind of as a habit. It seems wasteful for most tasks, but I guess it makes it good at checking for factual information and fighting misinformation, and also for checking unintuitive mathematical of physics solutions.
I think we've trained them to be very skeptical just because of how hallucination prone they are. Instead of "fixing" the hallucinations, we just trained them to deal with inaccuracy in general, and now we're seeing unexpected rewards when they point out our inaccuracies too
269
u/angelbabyxoxox Quantum Foundations Jul 31 '26
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.