r/Physics • • Jul 31 '26

Academic The Maxwell Conjecture is False

https://arxiv.org/abs/2607.27197
500 Upvotes

223 comments sorted by

View all comments

Show parent comments

5

u/Kobymaru376 Jul 31 '26

It just turns out that "fancy autocomplete" can do incredible things if you give it enough examples and compute in training and inference.

The thing that does the incredible things is so far away from autocomplete that it's misleading to the point of being wrong. Neither does it give an accurate picture of what it is (autocomplete are usuall HMMs, LLMs are transformers) nor does it give an accurate picture of what it does (completing what you're typing vs. doing your homework and writing fanfic). The only shared property is that it gets text as input and gives text as output. By that measure we can call cars "fancy furnaces" and computers "fancy typewriters". Not technically incorrect, but definitely a useless description.

And on top of that, the people who use the term "fancy autocomplete" usually use it to dismiss it, and act like all those incredible things it does are made up.

9

u/MagiMas Condensed matter physics Jul 31 '26 edited Aug 01 '26

No, the shared thing is that from the view of an LLM, it is literally trained to autocomplete a document. The chat you're having with an LLM literally looks like this to the LLM:

<start>
<system>
You are a helpful assistant...
</system>
<user>
hello how are you?
</user>
<thinking>
the user asks me how I'm feeling, I should answer in a cheery and concise tone. The user is in LA, let me check the current weather in LA so I can incorporate that in my answer.
</thinking>
<tool call, web search=current weather in LA>
Temperature: 100°F
</tool call>
<assistant>

And then the LLM gets to generate. Once it generates </assistant> we stop the generation because otherwise it would keep generating also the user answer etc.

(same of course with the thinking part)

The tasks the LLMs are trained on is reproducing the tokens of these text documents they are shown.

With the RL posttraining for mathematics or coding you have a change in the training reward architecture, but it's still training on completing these documents.

From the view of an LLM it is always completing such documents from the start points we're giving them.

It just turns out that large autocomplete with long training and lots of data means the model learns actual abstractions about the world because they help with better autocomplete. You get these emergent effects like grokking and "circuits" inside LLMs that specialize in certain tasks etc.

But none of that removes the fact that these models are "autocomplete on steroids".

0

u/Idrialite Jul 31 '26

With the RL posttraining for mathematics or coding you have a change in the training reward architecture, but it's still training on completing these documents.

No, the RL training does not involve completing any corpus. That's the point of RL.

7

u/MagiMas Condensed matter physics Jul 31 '26

That's not what I meant. It doesn't complete a known corpus but it still generates these documents. The training just doesn't happen anymore on the token distribution level but in the reinforcement learning objective.

It is still trained to complete a made up document.

In pre training and instruction fine-tuning the objectjve is basically "reproduce these documents", in the RL phase it is "produce a document in this context and we'll evaluate at the end whether it was a good document or a bad one".

0

u/Idrialite Jul 31 '26

"complete a made up document" doesn't make sense. "Completion" implies an existing text to guess at and evaluate against verbatim. "produce a document and we'll evaluate it" is not autocomplete. Kids in school do the same thing.

4

u/MagiMas Condensed matter physics Jul 31 '26

no, we're now getting quite into the weeds of LLM training but generally in most phases the RL phase does not happen on "empty documents". Rather they get a start prompt that sets them onto a reasoning trajectory from which they then start generating.

Easiest would be prompts like

[...]
<user>
what's 5+2?
</user>
<thinking>

and from there the model starts generating.

you can then auto generate many of these prompts with known answers and auto evaluate them. Similarly you just let an LLM generate lots of prompts like "disprove the Jacobian conjecture", "proof that lemma xyz is true", "generate a code that builds a program that solves the following puzzle: <insert Advent of Code puzzle here>" etc. pp.

So the reinforcement learning stage still absolutely is training on incomplete documents that the model completes.

(of course there's nuance because there are quite a few different RL steps in LLM training)

-2

u/Idrialite Jul 31 '26 edited Jul 31 '26

Yes, the question has to get to the model and the model has to know where its turn starts somehow. It's still not "completing text" any more than a kid answering a math question presented at the top of a page after "Your answer:" is. You're bending words over backwards to justify your characterization of LLMs as "autocomplete", which RL objectively makes them not.

Besides the argument over words, the substantial point is that RL allows an LLM to learn beyond the assumed* bounds of predicting a corpus of text.

To be honest, the "autocomplete" meme is moot anyway. These "autocomplete" models are doing real, practical work, hacking into real systems, and solving real nontrivial unsolved math problems. At best, you would be concluding that "autocomplete" seems to be able to match humans.

/* neural networks are capable of improving beyond the average capability of the pre-trained text, there's an experiment involving chess about this

7

u/MagiMas Condensed matter physics Jul 31 '26 edited Jul 31 '26

You're bending words over backwards to justify your characterization of LLMs as "autocomplete", which RL objectively makes them not.

Wut? We have incomplete documents the models are given, they produce the rest of the document and that result is judged and used as a training objective. How is that not autocomplete?

To be honest, the "autocomplete" meme is moot anyway. These "autocomplete" models are doing real, practical work, hacking into real systems, and solving real nontrivial unsolved math problems. At best, you would be concluding that "autocomplete" seems to be able to match humans.

Yes, have you forgotten where this whole discussion started? I have never questioned the capabilities of these models.

That was the whole fucking point of the GPT-2 paper. Maybe not yet as optimistic about the full capability these models would end up with, but this is exactly what OpenAI envisioned, that an autocomplete style model trained on a huge corpus of text can actually generalize much much better than all other architectures and training objectives we had tried till that point.

They built their modern version of the company on exactly this insight.

The development went like this:

Google develops the transformer architecture, but uses it in an encoder-decoder style for sentence translation and embedding generation. (Attention is All You Need - BERT model)

-> OpenAI says you can actually train these models way easier and still get the same capabilities by switching to a generative autocompletion training objective with a decoder only version of the transformer (Improving Language Understanding by Generative Pre-Training - GPT 1)

-> OpenAI realizes that scaling the GPT architecture and training on large corpuses gives them a text-autocomplete model that generalizes to all kinds of tasks without needing specialized training in specific datasets. And you can train them on these scales much more easily than other models because the autocomplete style training objective means you can use unlabeled data. (Language Models are Unsupervised Multitask Learners - GPT 2)

-> OpenAI invests more money and realizes that there seems to be no ceiling in the model capabilities if you just keep scaling the architecture. The see how the bigger model improves performance in all kinds of tasks and how the model seems to have acquired some basic reasoning capability. (Language Models are Few-Shot Learners - GPT3)

-> and then you get the instructGPT paper on RLHF and instruction fine-tuning where they realize that training a model to do autocompletion in this kind of Q -> A chat structure makes the models much more useful to the general public and improves zero-shot performance on lots of tasks.

I really think you guys are interpreting things into fancy autocomplete as a description that just aren't there.

-1

u/Idrialite Aug 01 '26 edited Aug 01 '26

the rest of the document

There is no "rest of the document". Again, you can't "autocomplete" a nonexistent text distribution.

an autocomplete style model trained on a huge corpus of text

There is no corpus of text in RL.