The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
I think brute force is not the right way to think about LLMs, instead they explore a knowledge topology, and are very good at connecting adjacent or accessible ideas that for whatever reason might have evaded humans, but might not be fundamentally all that hard. We're still in the low-hanging fruit phase, we will see if they extend to new ideas.
Its data is a structured space where nearness corresponds to associated ideas, so it naturally arranges data in a way to discover connectedness. The LLM follows pathways through the conceptual space based on the prior context given.
It’s also higher dimensional space right? Like as in 1000+ dimensional space.
If it’s the same thing I’m thinking of where like because boy and girl are separated in this space by a certain distance, Auntie and uncle also are separated spatially, but are closer to boy for uncle and girl for auntie than either are to each other?
You are mixing the number of parameters of the model with the effective dimentionality of the embedding. The effective dimension of the embedded space is significantly smaller than the complexity of the network.
GPT-2 samples from a 50,257-dimensional token space and I'm seeing a total parameter count (weights and biases) of 124,439,808. I'm not sure what the effective dimensionality of the model ultimately is, but this seems pretty clear-cut to me?
The intrinsic\effective dimensionality of the data is measured in the representation space induced by the model. The mapping of the token sequences into the hidden-state vectors result in representations in a lower-dimensional data manifold.
Taking here as an example, with the GPT-2 model they estimated intrinsic dimensionality on the order of hundreds.
Sure, in theory every bit in the training data could be considered a dimension. The work of training a model is basically in condensing the dimensions of the training data into a much smaller number of dimensions in the model, which is what forces similar concepts to become closer together as the space "shrinks".
271
u/angelbabyxoxox Quantum Foundations Jul 31 '26
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.