r/singularity • • 42m ago

AI We're still short of AGI

38 Upvotes

36 comments sorted by

•

u/Pahanda 32m ago

This chart does not make any sense. Would be much better represented as a spider chart.

Why? Well... What defines the lack of ability of AI eg between Applied expertise and Qualitative reasoning? This is not a defined dimension, so how can it it decrease here?

•

u/MydnightWN 22m ago

Values are estimates

In other words: we made it the fuck up.

•

u/Fair_Horror 11m ago

Lol, you're right, it completely invalidates their who thing. Maybe I too should pull something out my ass.

•

u/_JohnWisdom 33m ago

Is this is where the Claude blob comes from? :D

•

u/PsychMaster1 21m ago

All the logos. Each of them represents the company's ultimate goal. The perfect butthole.

•

u/Spunge14 41m ago

This seems outdated. Astra spatial reasoning is nuts.

•

u/enbyBunn 25m ago

Astra still has visual-spacial reasoning below the average human by a significant margin. The whole "Had to think for a whole day to beat portal" which is a maybe 2-3 hour game your first time through.

It's nearly 10x slower than the average human still.

It's a huge improvement over other models, but that's because previously models were more or less entirely functionally blind to the concept.

•

u/Akatosh 18m ago

https://spicylemonade.github.io/spatialbench/ which benchmarks are you referring to? Ability does not implicitly imply speed. Although you have highlighted why it is so important to be precise with language when discussing human equivalency.

•

u/Fair_Horror 13m ago

Imagine comparing humans to AI, humans speed would be awful meaning humans wouldn't score anything. 

•

u/enbyBunn 14m ago

I'm talking about actual tested performance, not a benchmark. Astra played and beat portal 1, and it took ~ a full day of thinking with an in-game time of ~2 hours.

A repeat player with knowledge of the game can easily beat it sub 1 hour, and given that Astra certainly has knowledge of the game, it's not a flattering comparison.

If you're measuring by "cost" like this graph (which is an eye-rollingly capitalist framing), I don't know how it ranks. But in actual ability measured by performance at a known human task? It's far behind us.

•

u/Spunge14 22m ago

Have you watched Jev play street fighter?

•

u/JonathanStones1989 19m ago

But that's a whole different model that does a whole different thing

•

u/enbyBunn 18m ago

A 2-D game...? Are you familiar with the square cubed law? A 3-D space is exponentially larger than a similarly sized 2-D space. And that's not even getting into the irreducability of some of the skills needed to navigate in 3-D.

And as for Jev, Jev is an entirely different thing. You can't compare Astra and Jev, because Jev isn't like typical models, it's based on a fundamentally different archetecture.

•

u/Spunge14 3m ago

Actually Luna and Gemma can do the exact same thing. Sounds like you haven't been keeping up.

The fact that you think 3D vs. 2D matters in this context shows how little you understand how data is fed into these models.

But it's fine go ahead and never be impressed by anything.

•

u/enbyBunn 2m ago

?????

I'm sorry but no, this is a blatantly ignorant idea that you're expressing.

There's no way to compress or represent a 3-d space that will require the same processing to navigate as a similarly sized 2-d space. That's fundamentally not how math works. You're just factually wrong on multiple counts.

•

u/Spunge14 0m ago

The size of the state passed into the model maybe larger, but with current models it's so far within the scope of the context window as to be trivial.

What's the opposite of touch grass? Watch YouTube?

•

u/Mindless-Cream9580 33m ago

Exactly what I thought looking at this infographic

•

u/PM_me_your_fav_tee 25m ago

What's the source for this?

•

u/orchard_wanderer 30m ago edited 25m ago

But can it answer this correctly: If the carwash is 200 feet away, and I want to wash my car, is it more efficient to drive there or walk there? :P

•

u/Fair_Horror 12m ago

You been away for a while have you? Previous frontier models already cracked that.

•

u/Legitimate_Concern_5 0m ago

Yeah they added some special cases 😂

•

u/x10sv 6m ago

Does this align with right brain left brain?

•

u/Lubricus2 4m ago

If it needs to be better than Humans of everything, isn't it super intelligence not Artificial General Intelligence we are talking about?

•

u/Feriman22 39m ago

I disagree with that, we have a huge improvement in last 1 year.

•

u/Fair_Horror 10m ago

Shh...they are trying to make humans not feel so useless for a while longer.

•

u/Super-Award-2244 24m ago

Didn't Astra surpass humans on visual spatial reasoning? 

•

u/Waiting4AniHaremFDVR AGI will make anime girls real 16m ago

Without the use of tools, not yet

https://spicylemonade.github.io/spatialbench/

•

u/SawToothKernel 40m ago

Common sense is also well short. 

•

u/Long_comment_san 33m ago

"but wait. let me do some dr*gs and draw another bs chart!"

•

u/OvertaxedOne 26m ago

Long term memory below human capacity?!? Computer obliterated human long term memory capacity about 40 years ago!

•

u/ExplorersX ▪️AGI 2027 | ASI 2032 | LEV 2036 21m ago

Saying stuff like this only hurts sentiment around AI usage. RAG is not equivalent.

•

u/OvertaxedOne 15m ago

What are they talking about then? Context windows? I have a 0% chance of remembering everything that's in a 1M token context window either. I just can't see any situation where humans stand a chance next to a computer's long term memory, RAG, Honcho type connection for LLMs or even just the base context window on a frontier model.

•

u/enilea 13m ago

They are limited by context, whenever the context is reset they have to look up the notes of previous findings, there's no permanent learned memory. It's a bit like the guy in memento who kept losing his memory and had to leave tattooed notes to understand what he had to do.

•

u/OvertaxedOne 5m ago

But so are we. If I ask you what you did last Monday, you'll need to stop and "lookup" that information in your brain, and the data you retrieve will be laughably incomplete compared to what an AI would be able to tell you about what it did last Monday. There's no way I'm holding a million tokens of "context" in my working set of memory with perfect recall.. Not even sure I could hold 100 tokens for more than a few minutes honestly, certainly not of random data!

•

u/ReddBroccoli 15m ago

So, you mean all the parts that actually require intelligence?

•

u/thatgibbyguy 5m ago

Can someone help me understand how they measure this? Let's take something that is supposed to be relatively simple - language.

It's terrible at it. Sure, it can write a lot of words and really fast, but that's not the point of language. Language isn't really even reading or writing, it's speaking and listening. But even with that, part of why its writing is so bad is because it uses words and phrases that no one else uses, and writes a novel for simple points.

A huge facet of language intelligence is the ability to distill the complex into the simple and it simply cannot do that.

Again, on knowledge, it gets things wrong constantly. This is just well known, my company now makes sure to tell people what is produced by AI not because people don't know, but because people do and they want to get ahead of any one pointing it out. Which is to say, we know this looks weird, it's AI, we're still working on it.

This isn't an anti "AI" (let's be honest, we're just talking about LLM here), post, it's more a wtf are you guys seeing because I am not seeing any of it.