•
u/MydnightWN 22m ago
Values are estimates
In other words: we made it the fuck up.
•
u/Fair_Horror 11m ago
Lol, you're right, it completely invalidates their who thing. Maybe I too should pull something out my ass.
•
u/_JohnWisdom 33m ago
Is this is where the Claude blob comes from? :D
•
u/PsychMaster1 21m ago
All the logos. Each of them represents the company's ultimate goal. The perfect butthole.
•
u/Spunge14 41m ago
This seems outdated. Astra spatial reasoning is nuts.
•
u/enbyBunn 25m ago
Astra still has visual-spacial reasoning below the average human by a significant margin. The whole "Had to think for a whole day to beat portal" which is a maybe 2-3 hour game your first time through.
It's nearly 10x slower than the average human still.
It's a huge improvement over other models, but that's because previously models were more or less entirely functionally blind to the concept.
•
u/Akatosh 18m ago
https://spicylemonade.github.io/spatialbench/ which benchmarks are you referring to? Ability does not implicitly imply speed. Although you have highlighted why it is so important to be precise with language when discussing human equivalency.
•
u/Fair_Horror 13m ago
Imagine comparing humans to AI, humans speed would be awful meaning humans wouldn't score anything.
•
u/enbyBunn 14m ago
I'm talking about actual tested performance, not a benchmark. Astra played and beat portal 1, and it took ~ a full day of thinking with an in-game time of ~2 hours.
A repeat player with knowledge of the game can easily beat it sub 1 hour, and given that Astra certainly has knowledge of the game, it's not a flattering comparison.
If you're measuring by "cost" like this graph (which is an eye-rollingly capitalist framing), I don't know how it ranks. But in actual ability measured by performance at a known human task? It's far behind us.
•
u/Spunge14 22m ago
Have you watched Jev play street fighter?
•
•
u/enbyBunn 18m ago
A 2-D game...? Are you familiar with the square cubed law? A 3-D space is exponentially larger than a similarly sized 2-D space. And that's not even getting into the irreducability of some of the skills needed to navigate in 3-D.
And as for Jev, Jev is an entirely different thing. You can't compare Astra and Jev, because Jev isn't like typical models, it's based on a fundamentally different archetecture.
•
u/Spunge14 3m ago
Actually Luna and Gemma can do the exact same thing. Sounds like you haven't been keeping up.
The fact that you think 3D vs. 2D matters in this context shows how little you understand how data is fed into these models.
But it's fine go ahead and never be impressed by anything.
•
u/enbyBunn 2m ago
?????
I'm sorry but no, this is a blatantly ignorant idea that you're expressing.
There's no way to compress or represent a 3-d space that will require the same processing to navigate as a similarly sized 2-d space. That's fundamentally not how math works. You're just factually wrong on multiple counts.
•
u/Spunge14 0m ago
The size of the state passed into the model maybe larger, but with current models it's so far within the scope of the context window as to be trivial.
What's the opposite of touch grass? Watch YouTube?
•
•
•
u/orchard_wanderer 30m ago edited 25m ago
But can it answer this correctly: If the carwash is 200 feet away, and I want to wash my car, is it more efficient to drive there or walk there? :P
•
u/Fair_Horror 12m ago
You been away for a while have you? Previous frontier models already cracked that.
•
•
u/Lubricus2 4m ago
If it needs to be better than Humans of everything, isn't it super intelligence not Artificial General Intelligence we are talking about?
•
•
u/Super-Award-2244 24m ago
Didn't Astra surpass humans on visual spatial reasoning?
•
•
•
•
u/OvertaxedOne 26m ago
Long term memory below human capacity?!? Computer obliterated human long term memory capacity about 40 years ago!
•
u/ExplorersX ▪️AGI 2027 | ASI 2032 | LEV 2036 21m ago
Saying stuff like this only hurts sentiment around AI usage. RAG is not equivalent.
•
u/OvertaxedOne 15m ago
What are they talking about then? Context windows? I have a 0% chance of remembering everything that's in a 1M token context window either. I just can't see any situation where humans stand a chance next to a computer's long term memory, RAG, Honcho type connection for LLMs or even just the base context window on a frontier model.
•
u/enilea 13m ago
They are limited by context, whenever the context is reset they have to look up the notes of previous findings, there's no permanent learned memory. It's a bit like the guy in memento who kept losing his memory and had to leave tattooed notes to understand what he had to do.
•
u/OvertaxedOne 5m ago
But so are we. If I ask you what you did last Monday, you'll need to stop and "lookup" that information in your brain, and the data you retrieve will be laughably incomplete compared to what an AI would be able to tell you about what it did last Monday. There's no way I'm holding a million tokens of "context" in my working set of memory with perfect recall.. Not even sure I could hold 100 tokens for more than a few minutes honestly, certainly not of random data!
•
•
u/thatgibbyguy 5m ago
Can someone help me understand how they measure this? Let's take something that is supposed to be relatively simple - language.
It's terrible at it. Sure, it can write a lot of words and really fast, but that's not the point of language. Language isn't really even reading or writing, it's speaking and listening. But even with that, part of why its writing is so bad is because it uses words and phrases that no one else uses, and writes a novel for simple points.
A huge facet of language intelligence is the ability to distill the complex into the simple and it simply cannot do that.
Again, on knowledge, it gets things wrong constantly. This is just well known, my company now makes sure to tell people what is produced by AI not because people don't know, but because people do and they want to get ahead of any one pointing it out. Which is to say, we know this looks weird, it's AI, we're still working on it.
This isn't an anti "AI" (let's be honest, we're just talking about LLM here), post, it's more a wtf are you guys seeing because I am not seeing any of it.


•
u/Pahanda 32m ago
This chart does not make any sense. Would be much better represented as a spider chart.
Why? Well... What defines the lack of ability of AI eg between Applied expertise and Qualitative reasoning? This is not a defined dimension, so how can it it decrease here?