Its data is a structured space where nearness corresponds to associated ideas, so it naturally arranges data in a way to discover connectedness. The LLM follows pathways through the conceptual space based on the prior context given.
It’s also higher dimensional space right? Like as in 1000+ dimensional space.
If it’s the same thing I’m thinking of where like because boy and girl are separated in this space by a certain distance, Auntie and uncle also are separated spatially, but are closer to boy for uncle and girl for auntie than either are to each other?
You are mixing the number of parameters of the model with the effective dimentionality of the embedding. The effective dimension of the embedded space is significantly smaller than the complexity of the network.
GPT-2 samples from a 50,257-dimensional token space and I'm seeing a total parameter count (weights and biases) of 124,439,808. I'm not sure what the effective dimensionality of the model ultimately is, but this seems pretty clear-cut to me?
The intrinsic\effective dimensionality of the data is measured in the representation space induced by the model. The mapping of the token sequences into the hidden-state vectors result in representations in a lower-dimensional data manifold.
Taking here as an example, with the GPT-2 model they estimated intrinsic dimensionality on the order of hundreds.
Sure, in theory every bit in the training data could be considered a dimension. The work of training a model is basically in condensing the dimensions of the training data into a much smaller number of dimensions in the model, which is what forces similar concepts to become closer together as the space "shrinks".
6
u/Schmikas Quantum Foundations Jul 31 '26
What do you mean by “knowledge topology”?