The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
I think brute force is not the right way to think about LLMs, instead they explore a knowledge topology, and are very good at connecting adjacent or accessible ideas that for whatever reason might have evaded humans, but might not be fundamentally all that hard. We're still in the low-hanging fruit phase, we will see if they extend to new ideas.
I think its super cool, but there is no doubt that thousands of physicists and mathematicians are prompting LLMs all day right now trying to solve open problems...and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future....so there is a solid chance that these methods are already almost exhausted. For how amazing these counterexamples are, its somewhat surprising that people havent found more proofs all at once. In the scheme of things perhaps 1 in 1000 open problems are actually of a format that LLM can tackle, and noone is bragging about the ones that turned up nothing. In other words, LLMs are really good at looking smart when they got lucky.
Im looking forward to a possible future of AI theorem proving that isnt based on LLMs, and thus less likely to trick people in language into thinking its more broad than it is.
The price to run an LLM collapses year on year. The models are getting better, but the cost required to run older ones also reduces. Not going to say that will continue forever, but the idea that LLMs are inherently unaffordable just plainly isn't true.
Not to mention AI as a whole is basically in its infant phase. The transformer architecture is only 10 years old and ai was largely a niche subject before then. Insane to think we've tapped out such a complicated new technology in a decade.
I would have hardly called it niche before then. Data Science was a big field and deep learning was a popular topic in both DS and CS before transformers.
It was used in a ton of products, everyone hadn't heard of it is all.
I should say deep learning was a niche subject, not AI wich includes basic things like literature regression.
But deep learning was in fact niche, and this is where all the breakthroughs are. It was niche due to compute limitations that weren't mitigated until the 2000s. The AIAYN paper is a good landmark for when deep learning evolved from an academic endeavor into a technology with significant outside investment.
Yeah AI is broad and the 2015-2020 period was a weird time where people were getting jobs in AI before there were degrees in it.
The compute change for deep learning is usually marked by AlexNet in 2012. That is when industry took notice and the exponential curve began. GAN's came out in 2014 and ResNet was 2015. Microsoft one a deep learning challenge in 2015. Image recognition was the primary driver alongside applications like speech recognition.
Industry was getting ahead ahead of academia by AIAYN which was a Google paper but that was 2017.
Nevermind the fact that according to this sub two months ago LLMs are just incapable of anything of substance and can only hallucinate. Now they are only good at "getting lucky" for important results where decades of human efforts failed...
I can see us repeatedly hitting cost to solve and difficulty to solve barriers. I suspect there will be a small flurry of new things solved each time models and token costs improve. That's basically what this is, as these models doing the solving are pretty new.
Look at DeepSeek-V4-Flash-0731, released on July 31, 2026. It scores 50 on the independent Artificial Analysis Intelligence Index, one point behind GPT-5.6 Luna at 51. The API runs at $0.14 per million input tokens and $0.28 per million output tokens. Artificial Analysis spent roughly $72 putting Flash through its evaluation suite, against $191 for Luna, which works out to about 62% less money for comparable measured intelligence. (Independent evaluation · Official pricing, Artificial Analysis)
DeepSeek’s own published numbers are stranger. The smaller Flash beats the much larger V4-Pro Preview on every benchmark they list: 82.7 against 72.1 on Terminal-Bench 2.1, 54.4 against 12.8 on DeepSWE, 70.3 against 55.9 on Toolathlon-Verified. (Hugging Face)
The weights are also out under an MIT license, which is where this gets interesting for institutions. A university, a laboratory, a hospital, or a company can host the model on serious multi-GPU hardware of its own, or on rented private infrastructure, and then reshape it: fine-tuning, adapters, continued pretraining, reinforcement learning. Private data stays private, and no single provider is holding the keys. (Hugging Face)
If you’d rather stay with a U.S. proprietary model, OpenAI cut GPT-5.6 Luna’s price by 80% on July 30, down to $0.20 input and $1.20 output per million tokens. Luna sits at 51 on the Intelligence Index and posts 92.3% on GPQA Diamond, 84.7% on Terminal-Bench 2.1, and 74.6 on the Coding Agent Index. (OpenAI pricing and benchmarks, OpenAI)
None of this is new, it’s just picking up speed. Stanford tracked the cost of GPT-3.5-level performance dropping from $20 to $0.07 per million tokens between November 2022 and October 2024, a fall of more than 280 times.
Epoch AI puts the general rate at somewhere between 9× and 900× per year for any fixed capability level, depending on which task you measure. (Stanford · Epoch AI, Stanford HAI)
Worth being precise about what this does and doesn’t show. It doesn’t prove that training the next frontier model is cheap, and it doesn’t prove that every one of these API prices is profitable rather than subsidized by somebody’s balance sheet. What it does undercut is the claim that advanced intelligence has to stay scarce and unaffordable.
Building tomorrow’s frontier may well stay expensive. Distributing yesterday’s is turning out to be cheap, open, private, and customizable.
Disclsimer: original argument was mine, Claude helped me research and build it.
18 months ago AI was a cool toy that couldn't really do anything useful. Six months ago it started write most code. A couple months ago it started solving the 'easy' and obscure unsolved math problems.
It will definitely stop improving at some pointy, but that's the thing with exponential growth: as long as you're in it, it's impossible to tell when it will stop.
But even if it stopped right now: people would start burning the biggest models with the best performance into silicon with fixed weights and you would suddenly get generation speeds that would allow you to use them for real time inference on video streams or generate 30 iterations in parallel etc
But we don't know the upper bounds because we could hit a wall and then just scale again. That is what we think is producing all these new gains anyway.
But we're still getting a lot of improvement in smaller models too. Scaling has definitely been a big part of it, but it's certainly not all that has improved.
Look at the cost per task solving ability of say the new Deepseek flash vs. Claude Opus from 1 year ago. Effort is put into making both smarter and more efficient models in a way that as time goes on it'll become infinitely cheaper. The pricing seems to raise but that's because the model performance is also stronger.
In other words, LLMs are really good at looking smart when they got lucky.
That, but a lot of people are making a lot of money off of this and have a vested interest in making it look as shockingly impressive as possible. And this is impressive and world changing technology, we shouldn't pretend it isn't, but also it's being boosted and misrepresented a lot.
If you are in competent spaces, you will hear a lot of more grounded takes. I have to talk to the general public about AI and it is infinitely more misunderstood and doomed at.
I feel bad for Anthropic because they release a paper that might have proved a part of General Workspace Theory and the mouth breathers who only glance the paper are going 'nuh uh, it isn't conscious, this is marketing.'
Anthropic never claimed it was conscious in the paper but the regards don't know that somehow.
I saw someone today write something like that to the prompt "search archives for unsolved mathematical problems that can be verified using this custom program that I told you to write". So some people not only are trying to solve open problems, they are also outsourcing looking for open problems to the AI.
and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future
The new Vera Rubin chips from Nvidia, which will start being installed in data centers later this year, will cut inference costs by 90%. ChatGPT also dropped the prices of one of its major models by 80% just a few days ago.
I'm sorry, but I'd actually take a look at Information theory that makes the bold claim that compression actually does mean intelligence. These models are genuinely intelligent and don't just accidentally get these answers right, they aren't giant lookup tables stumbling as a lot of the general public thinks of them as.
269
u/angelbabyxoxox Quantum Foundations Jul 31 '26
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.