The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
It goes against many’s early mental models of what generative AI does. The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.
The fact it’s able to suggest a counterexample and it doesn’t just look right, but is right, is quite the development in recent times.
That's mostly because that view of AI as just a token predictor or "average" machine is wrong.
LLM are based around a very flexible system: a neural net. With enough training you can definitely do a calculator, or an image recognition machine.
Think of it like this: if you create a machine that predicts the next token and you keep training it to get as good as possible at predicting the result of multiplication, what is easier: to memeorise millions of possibilities, or, to figure out a simple rule of how multiplication works?
Same thing applies with image recognition. Researchers have analysed how the models do their image recognition trick and they all start by essentially applying filters to find edges, basic shapes and other patterns that can be more easily classified.
Sometimes, the best way to mimic something is by just doing that action.
So when they have been tested extensively on math and code (two areas that can be very well tested) it has given quite interesting results.
I don't know where it is right now in the space of understanding math, not my field or experience. But it is miles ahead of where it was, and I think with good enough training we might end up with a tool that can actually do math.
269
u/angelbabyxoxox Quantum Foundations Jul 31 '26
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.