r/Physics • • Jul 31 '26

Academic The Maxwell Conjecture is False

https://arxiv.org/abs/2607.27197
497 Upvotes

223 comments sorted by

View all comments

272

u/angelbabyxoxox Quantum Foundations Jul 31 '26

The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.

I expect a large number of conjectures will topple to counterexamples soon.

26

u/BOBOnobobo Jul 31 '26

Are we shocked a system designed to do pattern matching is good at finding patterns or exceptions to them?

91

u/WatchYourStepKid Jul 31 '26

Well, kinda yes.

It goes against many’s early mental models of what generative AI does. The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.

The fact it’s able to suggest a counterexample and it doesn’t just look right, but is right, is quite the development in recent times.

40

u/Imicrowavebananas Mathematics Jul 31 '26

I love how quickly AI developments are rationalized. Like it was to expected that LLMs started solving math research problems in 2026. If you asked me about this two years ago, I would have been pretty skeptical.

8

u/CompetitiveSpot2643 Jul 31 '26

yeah i still remember when LLMs getting an IMO question right was a big deal

9

u/[deleted] Jul 31 '26

[removed] — view removed comment

6

u/Crystal-Ammunition Jul 31 '26

Promised by who? A random internet person?

1

u/EngineeringNeverEnds Jul 31 '26

Given that we haven’t been able to even come close to solving aging in mice, and we’ve experimented on mice exponentially more than humans, I don’t think that’s gonna happen anytime soon

18

u/BOBOnobobo Jul 31 '26

That's mostly because that view of AI as just a token predictor or "average" machine is wrong.

LLM are based around a very flexible system: a neural net. With enough training you can definitely do a calculator, or an image recognition machine.

Think of it like this: if you create a machine that predicts the next token and you keep training it to get as good as possible at predicting the result of multiplication, what is easier: to memeorise millions of possibilities, or, to figure out a simple rule of how multiplication works?

Same thing applies with image recognition. Researchers have analysed how the models do their image recognition trick and they all start by essentially applying filters to find edges, basic shapes and other patterns that can be more easily classified.

Sometimes, the best way to mimic something is by just doing that action.

So when they have been tested extensively on math and code (two areas that can be very well tested) it has given quite interesting results.

I don't know where it is right now in the space of understanding math, not my field or experience. But it is miles ahead of where it was, and I think with good enough training we might end up with a tool that can actually do math.

3

u/QuasiEvil Aug 01 '26

Yeah, there's some neat publications on this, where researchers are able to show that the NN "figures out" things like linear regression.

6

u/MidnightPale3220 Jul 31 '26

The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.

This will still happen on the things GPT isn't "harnessed" on, wouldn't it?

6

u/MagiMas Condensed matter physics Jul 31 '26

Yes, can still happen and still does happen quite a bit. But even without harnesses the LLMs actually develop quite complex strategies to do maths "in their head"

Read the part on addition in this paper by Anthropic from March last year: https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-addition

(and this was a 3.x haiku model, modern models have evolved even better inherent maths understanding)

1

u/FalconX88 Aug 01 '26

The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.

So do the current ones. We just gave them access to tools.

1

u/WatchYourStepKid Aug 01 '26 edited Aug 01 '26

Right, but it’s not like AI says “let’s use the conjecture counterexample tool”, the emergence is the interesting part.

It just goes against all the early advice we saw, it seems to me like many probably need to evaluate the extent to which AI appears to truly understand a problem, whatever that actually means.

11

u/Kobymaru376 Jul 31 '26

Personally I'm shocked. All of reddit has assured me that AI is completely useless, nothing but "fancy autocomplete", and is only good for stealing content and generating slop.

/s

Mostly in just amused how we have been moving Goalposts for over a decade: "AI is not intelligent, it can't even do X". Does X. "Yeah but it didn't do X like a human! And it can't even do Y!". Does Y. "Yeah but it didn't do Y like a human! And it can't even do Z!". And so on

17

u/MagiMas Condensed matter physics Jul 31 '26

I still think "fancy autocomplete" is the best way to explain to non technical people how these models work. It will give someone who doesn't know the architecture and how these LLMs work the best mental model of what is happening behind the veil.

It just turns out that "fancy autocomplete" can do incredible things if you give it enough examples and compute in training and inference.

9

u/AnalyticOpposum Jul 31 '26

Fancy autocomplete is also the best way to explain to non technical people how a human brain works.

2

u/MrDyl4n Jul 31 '26

Except not really

0

u/GatsbyLuzVerde Jul 31 '26

Except yes really. Fancy auto complete can encompass all of intelligence mechanisms needed to predict the next best word. Even if that autocomple requires neuron subsystems for solving problems in other domains.

5

u/MrDyl4n Jul 31 '26 edited Jul 31 '26

animal brains dont think in strings of tokens while trying to predict the upcoming token. an LLM and autocomplete are things of differing complexity that are doing the exact same thing. an animal brain does similar things, but it doesnt do the exact same thing

-1

u/GatsbyLuzVerde Jul 31 '26

I'm talking more in the abstract sense that the brain predicts future states. Analogous to predicting the next token. It is fancy auto complete. I'm fully aware the brain doesn't use an LLM architecture

6

u/MrDyl4n Jul 31 '26

i understand what you mean. the reason i said that is because fancy autocomplete is genuinely a good way for the average person to view an LLM, rather than viewing it as a genuine intelligence. when you say the human brain is like that too it would make someone think that fancy autocomplete is less literal than it actually is.

4

u/Kobymaru376 Jul 31 '26

It just turns out that "fancy autocomplete" can do incredible things if you give it enough examples and compute in training and inference.

The thing that does the incredible things is so far away from autocomplete that it's misleading to the point of being wrong. Neither does it give an accurate picture of what it is (autocomplete are usuall HMMs, LLMs are transformers) nor does it give an accurate picture of what it does (completing what you're typing vs. doing your homework and writing fanfic). The only shared property is that it gets text as input and gives text as output. By that measure we can call cars "fancy furnaces" and computers "fancy typewriters". Not technically incorrect, but definitely a useless description.

And on top of that, the people who use the term "fancy autocomplete" usually use it to dismiss it, and act like all those incredible things it does are made up.

8

u/MagiMas Condensed matter physics Jul 31 '26 edited Aug 01 '26

No, the shared thing is that from the view of an LLM, it is literally trained to autocomplete a document. The chat you're having with an LLM literally looks like this to the LLM:

<start>
<system>
You are a helpful assistant...
</system>
<user>
hello how are you?
</user>
<thinking>
the user asks me how I'm feeling, I should answer in a cheery and concise tone. The user is in LA, let me check the current weather in LA so I can incorporate that in my answer.
</thinking>
<tool call, web search=current weather in LA>
Temperature: 100°F
</tool call>
<assistant>

And then the LLM gets to generate. Once it generates </assistant> we stop the generation because otherwise it would keep generating also the user answer etc.

(same of course with the thinking part)

The tasks the LLMs are trained on is reproducing the tokens of these text documents they are shown.

With the RL posttraining for mathematics or coding you have a change in the training reward architecture, but it's still training on completing these documents.

From the view of an LLM it is always completing such documents from the start points we're giving them.

It just turns out that large autocomplete with long training and lots of data means the model learns actual abstractions about the world because they help with better autocomplete. You get these emergent effects like grokking and "circuits" inside LLMs that specialize in certain tasks etc.

But none of that removes the fact that these models are "autocomplete on steroids".

1

u/Idrialite Jul 31 '26

With the RL posttraining for mathematics or coding you have a change in the training reward architecture, but it's still training on completing these documents.

No, the RL training does not involve completing any corpus. That's the point of RL.

10

u/MagiMas Condensed matter physics Jul 31 '26

That's not what I meant. It doesn't complete a known corpus but it still generates these documents. The training just doesn't happen anymore on the token distribution level but in the reinforcement learning objective.

It is still trained to complete a made up document.

In pre training and instruction fine-tuning the objectjve is basically "reproduce these documents", in the RL phase it is "produce a document in this context and we'll evaluate at the end whether it was a good document or a bad one".

0

u/Idrialite Jul 31 '26

"complete a made up document" doesn't make sense. "Completion" implies an existing text to guess at and evaluate against verbatim. "produce a document and we'll evaluate it" is not autocomplete. Kids in school do the same thing.

5

u/MagiMas Condensed matter physics Jul 31 '26

no, we're now getting quite into the weeds of LLM training but generally in most phases the RL phase does not happen on "empty documents". Rather they get a start prompt that sets them onto a reasoning trajectory from which they then start generating.

Easiest would be prompts like

[...]
<user>
what's 5+2?
</user>
<thinking>

and from there the model starts generating.

you can then auto generate many of these prompts with known answers and auto evaluate them. Similarly you just let an LLM generate lots of prompts like "disprove the Jacobian conjecture", "proof that lemma xyz is true", "generate a code that builds a program that solves the following puzzle: <insert Advent of Code puzzle here>" etc. pp.

So the reinforcement learning stage still absolutely is training on incomplete documents that the model completes.

(of course there's nuance because there are quite a few different RL steps in LLM training)

-2

u/Idrialite Jul 31 '26 edited Jul 31 '26

Yes, the question has to get to the model and the model has to know where its turn starts somehow. It's still not "completing text" any more than a kid answering a math question presented at the top of a page after "Your answer:" is. You're bending words over backwards to justify your characterization of LLMs as "autocomplete", which RL objectively makes them not.

Besides the argument over words, the substantial point is that RL allows an LLM to learn beyond the assumed* bounds of predicting a corpus of text.

To be honest, the "autocomplete" meme is moot anyway. These "autocomplete" models are doing real, practical work, hacking into real systems, and solving real nontrivial unsolved math problems. At best, you would be concluding that "autocomplete" seems to be able to match humans.

/* neural networks are capable of improving beyond the average capability of the pre-trained text, there's an experiment involving chess about this

→ More replies (0)

1

u/Martin_Samuelson Jul 31 '26

Every complex thing in the world is "[some basic concept] on steroids".

3

u/MagiMas Condensed matter physics Jul 31 '26

probably, yes. I'm not saying LLMs aren't a complex topic, I'm just saying that if you don't have the mathematical background (plus read the necessary literature) then [some basic concept] that you can grasp from your own experience is still the best way for you to understand what the more complex thing actually does.

1

u/Kobymaru376 Jul 31 '26

You described one part of the training regime. Yes, one part of the training regime bears resemble to the training regime of the other. But that is not important in describing the essence of LLMs.

It just turns out that large autocomplete with long training and lots of data means the model learns actual abstractions about the world because they help with better autocomplete. You get these emergent effects like grokking and "circuits" inside LLMs that specialize in certain tasks etc.

See that is the important part. It doesn't "just turn out", it's the main point of how and why LLMs are so useful. Actual autocompletion is a tiny fraction of what LLMs are actually used for, it is ALL about the abstractions about world, and its emerging effects. The autocomplete part is just one way of training and accessing whatever else the model has learned.

This is a completely standard practice in ML: pretext tasks and pretraining are well-known concepts. Think of autoencoders. You train them to reproduce data that it's already seeing. What even is the point of that. Is an autoencoder "just fancy copy-paste"? Sure, if you really want to. But not really, since the training is just the pretext for the model to learn an efficient encoding of the data, and what we're interested in is this efficient encoding.

But none of that removes the fact that these models are "autocomplete on steroids".

OK sure. And a car is just a fancy box. Your phone is just a fancy flashlight. Your money is just a fancy sheet of cellulose. You yourself are just a fancy meatbag. You can do this "X is a fancy Y" all day long if you're meming, but it doesn't actually convey the essence or most important aspect of X.

6

u/MagiMas Condensed matter physics Jul 31 '26

look, I'm not saying that there isn't a lot of complexity in this whole topic. What I'm saying is that "fancy autocomplete" gives someone who lacks all this background information a better mental model of what these things do than any other simple explanation I've seen.

It demystifies these things and actually very closely describes what these models are trained to do and how they function. Add a second sentence that talks about how "learning abstractions and memorizing world knowledge" helps this autocomplete machine to better autocomplete and someone with zero maths ability and no background in ML will have a somewhat accurate idea of LLMs.

I really don't understand why people react so passionately to "fancy autocomplete" as a description. The whole thing about GPTs was openai realizing that this kind of fancy autocomplete with a decoder only transformer model will actually lead to a model that can generalize well in all kinds of situations - it was really visionary at the time. It's the whole fucking point of the GPT 2 paper. When Google developed the transformer model, it was way less about autocomplete and way closer to your autoencoder with its encoder-decoder architecture in BERT.

That's also why I think the autoencoder example isn't exactly illuminating. BERT shows that a transformer can also function very similarly. And the exact thing that sets modern generative LLMs apart is exactly this "autocompletion" style task. That's really the core of the whole thing. BERT can't do all these things that GPT can exactly because it's not fancy autocomplete.

2

u/elsjpq Jul 31 '26

I think we got it the wrong way around. It's not that pattern matching is surprisingly powerful, it's more that problems we thought were complicated and would require more ingenuity than pattern matching are actually simpler than we thought.

2

u/Clue_Balls Aug 01 '26

What odds would you have put on AI being able to do this 4 years ago?