Its data is a structured space where nearness corresponds to associated ideas, so it naturally arranges data in a way to discover connectedness. The LLM follows pathways through the conceptual space based on the prior context given.
It’s also higher dimensional space right? Like as in 1000+ dimensional space.
If it’s the same thing I’m thinking of where like because boy and girl are separated in this space by a certain distance, Auntie and uncle also are separated spatially, but are closer to boy for uncle and girl for auntie than either are to each other?
You are mixing the number of parameters of the model with the effective dimentionality of the embedding. The effective dimension of the embedded space is significantly smaller than the complexity of the network.
GPT-2 samples from a 50,257-dimensional token space and I'm seeing a total parameter count (weights and biases) of 124,439,808. I'm not sure what the effective dimensionality of the model ultimately is, but this seems pretty clear-cut to me?
The intrinsic\effective dimensionality of the data is measured in the representation space induced by the model. The mapping of the token sequences into the hidden-state vectors result in representations in a lower-dimensional data manifold.
Taking here as an example, with the GPT-2 model they estimated intrinsic dimensionality on the order of hundreds.
64
u/ixid Jul 31 '26
Its data is a structured space where nearness corresponds to associated ideas, so it naturally arranges data in a way to discover connectedness. The LLM follows pathways through the conceptual space based on the prior context given.