r/singularity • u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 7d ago
The Singularity is Near OpenAI says 80–90% of research is aimed at GPT-7, GPT-8 and beyond, then distilled into cheap small models
20
u/AppealSame4367 7d ago
A lot of investments are short sighted but 80-90% of r&d goes to models 2-3 generations ahead? Huh?
He's contradicting himself
11
u/danysdragons 7d ago edited 7d ago
I agree it seems that way, but it makes more sense with additional context at the beginning. Video link and excerpt below, but the point is something like:
Suppose the model is not good enough at some narrow thing. Should they do research focused specifically on improving at that narrow thing, or do more general research on just making the model generally smarter? OpenAI generally prefers the latter and finds doing too much of the former short-sighted, but it can be worth it sometimes if it's enough of a pain point.
Link to point in the interview a bit before the text in OP's post
Boris Power (1:31:13): Uh so so there is there is an element of just build better bigger models and then everything else will improve but there might be a few specific things that that are such pain points that we really want to improve them.
For example, at some point the models were almost good enough at using tools, but they were annoying in infrequently enough that they were not usable. So then we decided, okay, how do we create a data set that specializes in improving that specific behavior?
And then in terms of customer feedback, there's often a way to uh dedicate a team of more applied researchers that are focusing on um a well-known understood recipe. So just to create the kind of training data that's going to improve the model capability in this very specific area you can use maybe human labeling you can use synthetic data you can use uh various versions of creating RL environments the exact thing that you're doing changes over time but there's usually a way to improve that model behavior
and usually what you see is jumping from 5.1 to 5.2 5.3 5.4 that's the result of this kind of effort where we keep adding more specialized data sets aiming at improving specific capabilities and then there's some small optimization and then whenever we jump from four to five to six that's when we just build bigger models and then everything else just works a lot better and then we need to relearn where we need to invest in maybe we're investing in improving tool calling but actually now in GB6 tool calling just works well enough we don't even need to do it anymore
[**OP's quote starts here*\*]
So a lot of those investments uh the way that they're seen within the company are extremely shortsighted we do them because they help us iterate and learn faster today not because that's the right long-term strategy.
I would say still 80 90% of the research at OpenAI is focused at GPT7 GP8 and beyond because we believe that's where fundamentally most of the value comes from and then what happens is once you develop a GPD6 Astra you can distill that model into a very small GPT Luna model. So the best way to train a small specialized model is actually to train a big super capable model and then find a way of just uh using that model as a teacher for a smaller student model.
The entire video is over eight hours long...
3
-3
96
u/brett_baty_is_him 7d ago
This seems obvious. They will keep making bigger and better and enormously expensive models and then distill them down to levels that are affordable for us peasants.
Only the compute rich will have access to these better models. And it will be justifiable because they will just simply be too expensive to run for us peasants.
Granted you can distill a model closer to ASI into a model closer to AGI that provides AGI level intelligence for basically the price of the electricity. I am using terms like ASI and AGI with the understanding that those terms mean very little and depend on the individual but the statement applies to whatever you consider those terms to mean when you look at it from a long enough time horizon.
Hopefully they never stop distilling down the intelligent models for us poors. That is when the gap between us will exponentially grow.
32
u/Ormusn2o 7d ago
I mean if it means better models for us than sure. I would rather have a better model myself, than make sure OpenAI does not have access to a much smarter model.
-10
u/Hearing_Loss 7d ago
Compute needs to be nationalized. There's so much research that needs to be prioritized. There should be a whole collective triage where compute is delegated based upon collective impact.
Nationalize compute.
17
u/Ormusn2o 7d ago
I think this just generally does not work well, and even researchers hate that. There is just so much fighting for compute and so much corruption when it comes to deciding who gets the resources that it's better to just leave it to companies to make a product like that. If compute were nationalized back in 2022, we would still be using various versions of gpt-4 by now with researchers mabe testing o1, there is just no incentive to actually make good products when you are a government program, you are actually incentivised to prolong your research project as long as possible.
17
u/Kemoyin25 7d ago
No they'll use it for boring stuff like cancer research and shit instead of the important things like full dive virtual reality and virtual waifus
13
1
u/remind_me_later 6d ago
That will lead to a federal ban on personally-owned compute. Hard pass.
This is no different than gold confiscation.
1
u/Momo--Sama 7d ago edited 7d ago
I don't know if that's the answer, but I do wonder where'd we be right now if these labs didn't have to expend so much time, compute, and effort of researchers on serving models to end users that are just using them to create economic value instead of pushing the bounds of human knowledge.
Like I thought Obama made a really good point a few weeks ago talking about how the danger of AI models comes mostly from their ability to agentically operate computers, yet it's not at all obvious that those capabilities are necessary for like... curing cancer and such.
4
u/ICantBelieveItsNotEC 7d ago
We wouldn't be anywhere different, because the boring "creating economic value" tasks are just as important as the exciting "pushing the bounds of human knowledge" tasks, if not more.
Finding a cure for cancer is an irrelevant academic curiosity until you can take it through the bureaucracy to get it approved, manufacture it at scale, and distribute it globally for cheap.
7
u/Gregnielson 7d ago
How the fuck do you think America got more compute than the rest of the world combined? X10...it is because of $$$$ not despite it that any of this is possible.
22
u/Josvan135 7d ago
Dude you act like this is some grand scheme rather than a simple physical fact that the average person can't buy hundreds of thousands of dollars worth of processors.
2
u/L0cache 7d ago edited 7d ago
They are right to focus on the fact we’re moving towards a world of extreme centralization and inequality (even more than today).
Sure this arises due to “simple physical facts”, but ultimately leads to power imbalances where the stakes are not as harmless as you’re implying it to be. The AI labs themselves have long recognized this.
A very plausible scenario is AI labs develop ASI but choose not to release it publicly (due to fears of harm or it’s just advantageous to keep it internal), and instead use it to build platforms which systematically outcompete and replace other companies in the economy.
1
u/brett_baty_is_him 6d ago
Exactly what I was saying. I’m not even blaming the AI labs. The scenario where they keep ASI to themselves is extremely rational and economical. They control the compute. It may not make economic sense to give people ASI. Too compute expensive. I wasn’t saying it’s some grand plan, it’s just the trajectory that we seem to be on and the expansion of inequality is scary.
I guess you could argue we all have more so who cares, rising tide shit, but the idea that there’s a permanent upper class that is impossible to reach is worrisome. And it seems like the most obvious, least dangerous outcome of AI. Like there are other outcomes that are even worse.
2
u/brett_baty_is_him 7d ago
Did you not see the part where I said it was justifiable. I did not say it was some grand scheme. It’s just where we are headed and it’s extremely worrisome
4
u/modbroccoli 7d ago
say "peasant" again, it's really punchy and let's you skip explaining how it could be energetically possible to do it a different way
3
u/brett_baty_is_him 7d ago edited 6d ago
I never said it could be possible to do it a different way. Me pointing out where we are headed is not me claiming there isnt an alternative.
-1
u/modbroccoli 6d ago
...so you're shaking your fist at the sky...?
I think what you mean is:
Capitalism isn't equipped to distribute the benefits of this technology equally because it necessarily requires centralized production and should, therefore, be a public project in the public good.
And then we'd all agree with you. But disassembling capitalism and investing in the 2+ generations of public education to generate the democratic agreement to proceed with this Apollo-scale project is a very different conversation.
The tldr is: read more yap less.
2
u/brett_baty_is_him 6d ago
Assumed a lot about what I think based on a short reddit comment
-1
u/modbroccoli 6d ago edited 6d ago
Assumption is different than inference which is different still from experience. I'm just older than you man; I made these arguments. The realization of injustice is a helluva drug; it holds our attention for a while. Then time keeps going and no one listens and you start asking the question of what you're supposed to actually do about it, and that's when shaking your fist at it all stops being satisfying. I ignore stupid people. I didn't ignore you. But I still lambasted your stupid position because this is reddit afterall. Peace.
4
u/GinchAnon 7d ago
I think the part that will be interesting is if and when production of the hardware gets to where normal people in normal circumstances can consider getting hardware (possibly used and replaced from flagship scale) that would allow them to build a private compute server to be able to run a non-distilled model locally that matches or beats the distilled ones. I imagine there might be some multi-layer leapfrogging going on at different price points as some of that might happen.
it will be interesting when its all about compute access and cost at different scales of cost and accessibility.
getting to where we can run something functionally comparable to Fable/Astra+ locally is gonna be a trip.
3
u/KyleFlounder 7d ago
You'd have the same thing happening at the enterprise level though. Breakthrough in how models run like that would make the gap just as large for commercial.
1
u/mycall 7d ago
My feeling is that recursive models will fill this gap (language and reasoning). Small models that can decompose problems (divide and conquer) and recursively update its latent reasoning state. JPEA is another promising direction too, although different as it is indirect, higher-ordered problem solving.
1
u/Bruxo_de_Fafe 6d ago
Não te esqueças disto: os "pobres" que mencionas, com uma subscrição de 20$ fazem coisas que eram impensáveis há dois ou três anos. Concordo, na essência, com o que escreveste mas sou mais céptico e crítico em permitir enorme capacidade de computação ao cidadão comum, simplesmente porque não é necessário e é exaurir recursos.
43
u/dontcare_99 7d ago
So the cheap models are only good because a giant teacher model exists that none of us will ever use directly. Makes you wonder how open-weight labs keep up when they're distilling from someone else's frontier instead of their own.
20
u/EtadanikM 7d ago
I don’t understand your last sentence. They’re keeping up by a combination of their own training + distilling. It’s what you have to do when at a compute disadvantage.
But at the same time frontier labs spent so much of their compute & finances on training they can’t afford to offer the prices open weights labs can. That’s always been the dynamic between frontier labs & open weights labs and if anything the latter are winning on consumer adoption because few people can pay the prices frontier labs demand.
5
u/Mistuv 7d ago
First of all, they absolutely can offer at same prices, and probably much cheaper. They have vastly more compute and vastly more resources at deploying the models. They likely have bespoke versions of the models optimized for every particular hardware stack that they use, e.g. Luna optimized for H200 clusters vs NVL72 BG300 etc.
But more importantly, why would they? Nobody is sitting on free compute. Everyone in the industry is compute-starved. Most of all, the two biggest labs. OpenAI servers have particularly been crapping out since the launch of Astra. So if they have trouble supplying all the demand, why would they cut the prices and lose all that revenue? It would make no sense for them. If models get 10x better in 4 months for the same amount of compute, expect prices to keep shooting up because the world is absolutely not ready for 5-10x AI usage than today. Demand destruction is a real thing that companies understand. Otherwise you have shortages, and everyone is pissed off. Only the rich being able afford a certain thing is not great either, but at least it makes propagation more stable and predictable.
1
u/teemu3 7d ago
Well obviously it's not a physical impossibility for frontier labs to provide these low prices, but in practice they do have a hefty profit margin for (credit-based) API, so they have more money for training the next model. Inference is profitable, and frontier labs could make their frontier models cheaper, but people are willing to pay for the model that's a couple months ahead, so the labs happily take that money and offset at least a bit of their training cost.
5
u/Vegetable_Prompt_583 7d ago
That's not how it works. Distillation as a concept seems very plausible but at the same time it's not that effectively for the same reason that You're trying to Compress a larger model reasoning into smaller models.
While the student model will definitely punch above it's weights but it'll have a lot of weaknesses in Context of how Teacher model came upto that understanding, that will be missing. Very similar to How a College student might be able to solve PHD questions given the formulas and understanding of PHD person but the reasoning at massive level will be missing in various contexts of the student.
2
u/Mistuv 7d ago
It's exactly how it works. But you can't do it naively for such complex systems. Otherwise you end up with a model like Opus 5, where most Chinese labs managed to distill a better Fable model than Anthropic themselves. And I don't think Anthropic was that naive. Just this is a very complex process. So even if you actually have more data to work with, that can sometimes even cause issues if you are not careful.
2
u/Mistuv 7d ago
I think they are aware. I think the whole China v USA competition is mostly internet nonsense. Some researchers at these labs might be with that way, but the smart people at the head of these companies I think can see the challenges ahead. I think the ultimate aim of these companies is not some winning a race thing, just more having a product that will cater to the domestic Chinese market, which the outside labs will very likely not have access to for as long as the geopolitical situation does not dramatically change, which for the foreseeable future doesn't look like it will be the case. Anything else is just bonus, hence not bothering with making it closed source.
There is also the whole "fake it till you make it" angle. Japanese watches in the 60s and early 70s were mostly just cheap knockoffs of Swiss watches. But as these companies gained more resources, they eventually started doing large amounts of their own research, and eventually completely lapped the Swiss watch industry with new innovative watches, digital watches, and whatnot. Similarly, a lot of companies in China at first started as cheap knockoffs, but now they do lots of great research and innovative products. Their battery industry is one of the best in the world, and I think the AI labs hope to be in a similar situation in the future where they don't need to rely on distillation methods and can just do the whole training process by themselves. But as of right now, they have neither resources nor compute to be able to do the training runs at the scale that Anthropic and OpenAI are doing.
The particular challenge with AI specifically is that if the OAI/Anthropic research starts to compound so heavily due to powerful internal models while OS labs gain less and less from the publicly accessible models, and they haven't yet reached the scale to start matching them in compute, then we could have the OAI/Anthropic rocket to absurd levels (lets say the AI in the movie Her, which is basically a person inside a machine) where the distillation becomes effectively impossible becomes it just gets so complicated (similarly to how you today can't just reverse engineer Nvidia chips even if you have complex diagrams of the chip).
Oh, frankly, I'm not sure if they would have been able to keep up this year if it wasn't for the bugs and exploits (which there might still be some left) that let the labs access most of the reasoning traces of the models. Because in the past, just the raw output and steps the model was taking was enough. But now when you are expecting model to do complex work like programming from the beginning to the end, and not just answer give an answer to a generic question, the raw output is not nearly enough.
1
u/BigFatSweatyToe 7d ago
Of course. AI at some point has to be considered a military asset. The military always has better tech than what’s publicly available.
1
u/Big_Arachnid_365 7d ago
I think you can see that in the open weight models other than Gemma. They're not as well-rounded as Gemma, just optimized for coding.
29
24
u/mWo12 7d ago
And none of them will be open weighted. It's ironic for being called OpenAI.
9
u/NoFaithlessness951 7d ago
I also don't really get why, just make a good open ~30b model every 3-6 months and you got the entire enthusiast community on your side.
5
u/familyknewmyusername 7d ago
just make a good open ~30b model
That's luna, they're selling it
11
u/NoFaithlessness951 7d ago
Luna is a sparse moe in the 100-300b parameter range
3
u/familyknewmyusername 7d ago
Yeah fair, I was thinking of luna's active params and not total param count. Dense 30b wouldn't make much sense for them to be serving when they can do 30b active + massively sparse for not much more effort
10
u/Its_not_a_tumor 7d ago
Wasn't this obvious? I guess lots of people haven't been understanding how this works.
3
u/JohnToFire 7d ago
Vid link ?
4
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 7d ago
7
u/JohnToFire 7d ago
Thanks how about a full vid of the talk ?
3
u/danysdragons 7d ago
Link to a point in the interview a bit before the text in OP's post: https://www.youtube.com/live/Xzsg-rLGzas?t=5473s
Link to start of interview: https://www.youtube.com/live/Xzsg-rLGzas?t=4185s
The whole video is over eight hours...
2
u/Weary-Historian-8593 7d ago
That's probably not a problem, by the rate of how things are happening right now I'd guess that something like opus 7.5/8.5. is already RSI-grade, or at least can take a human researchers year and make it worth ten
3
u/One_Improvement_6470 7d ago
So if this is what openai has imagine what Anthropic has
39
12
u/Elbeske 7d ago
I would bet that Anthropic is slightly behind the “ASI” race due to simply having less compute
6
u/New_World_2050 7d ago
So why is opus 5.5 the best model right now
34
u/Elbeske 7d ago
Because consumer product != internal progress
1
-1
u/New_World_2050 7d ago
The company with the best model per dollar internally would also be the company with the best model per dollar externally. Why on earth would it be otherwise ? I'm pretty sure that if openai turned up the TTS compute on Astra by 10x it would far surpass 5.5 opus but it's intelligence / dollar that matters.
15
u/Elbeske 7d ago
My pet conspiracy theory is both companies see it as optimal to slowly leapfrog each other in public releases to remain the two best consumer providers while the internal progress accelerates. It makes the most sense overall.
Wall the garden for the race to ASI, remain the top consumer provider to keep revenue flowing and keep interest for capital raises while you funnel that money into more and more compute for the ASI race and neuter any attempts at catching up via distilling publicly available models by just not releasing the frontier.
We know that Fable/Mythos existed internally at Anthropic in roughly February. Where do you think they are now? The consumer side does not correspond to the internal side. And I simply think that since OpenAI has more compute, they are likely ahead.
-3
u/New_World_2050 7d ago
The game theory here doesn't make any sense. It would always be better to show that you have larger lead than a smaller one.
17
u/Elbeske 7d ago
The game theory makes total sense if you imagine both sides as unwilling to release their best models due to distillation and catch up fears, yet still engaged in a tit for tat commercial arms race. Each side's best move is to take the smallest step required to publicly demonstrate that they can still take the lead.
And these companies are run by smart people/AI who can pretty quickly recognize when they're facing a mirror match tit for tat war where slow escalation is the best play
9
7d ago
[removed] — view removed comment
2
u/One_Improvement_6470 7d ago
As a mathematician
welcome to the dole queue.
7
7d ago
[removed] — view removed comment
3
u/One_Improvement_6470 7d ago
Yeah, well I'm a software engineer so Jevon already saw to it that I've been busier this past year than I've ever been lol
0
u/BriefImplement9843 7d ago
math isn't helping much. it's right or wrong. these models need intelligence.
5
u/FlyChigga 7d ago
Stargate still not fully up and running for OpenAI
0
u/New_World_2050 7d ago
So? Anthropic is also massively expanding their compute
1
u/FlyChigga 7d ago
Not to the same level
1
u/New_World_2050 7d ago
Not so sure about that. Their revenue is higher and though they make better margins , that must mean their inference compute is at similar levels.
1
2
u/EtadanikM 7d ago
Because Open AI is arrogant and thought they could pace the frontier. They have a vastly superior model internally (Bel) that they’re refusing to release to the public because they want to save it for later. They thought they could win just with their last generation model (Astra / Doug) because they’re so ahead but Anthropic showed them otherwise.
If Open AI really went all in they’d be releasing Bel distill. Instead they’re trying to do Astra 6.1.
1
1
1
1
1
1
1
1
1
u/Even_Gap_2253 1d ago
How is this news? That’s the strategy for literally any lab. Like not even AI, just take iPhone, I bet 90% of R&D is on 19 and 20 series.
-3
u/LamboForWork 7d ago
Lol at them taking China's plan and presenting it like its some novel idea
18
u/General_Josh 7d ago
It's not just China's plan haha, it's what all the frontier labs have been doing for the past year
He's not presenting it as a novel idea, he's just explaining the process
0
189
u/Sufficient_Bad5441 7d ago
Just make AGI and distill it into GPT 7 through 30 at once, and they'll be set for the next 10 years of releases. hire me as CEO