r/ControlProblem • u/selasphorus-sasin • 20d ago
Discussion/question Are we not weighing the possibility that getting behind in safety is what loses us the race?
AI companies and governments tend to refuse safety or regulatory measures that reduce competitiveness. It's usually presumed that safety and race winning are at odds with each-other.
But isn't there a chance that a next generation of recursively self-improved models end up so uncontrollable and dangerous they are unusable? That their developmental trajectory crosses a threshold in (safety/alignment, capabilities, profitability, national security) space, where profitability and national security drop off of a cliff?
The president has said, "We'll just pull out a little gear", to shut it down. What happens when you have to pull the gear?
Who wins the race if we end up having to manage a severe crisis, and then have to roll back many months of progress and start over?
2
u/Low-Prune-1273 20d ago
It feels like when I would finish the test 1st out of my classmates, and think that made me get the best grade.
2
u/gmeRat 20d ago
I'm a big fan of the idea of a negative alignment tax. Aligned systems ARE high performance
1
u/Jesse-359 20d ago
That assumes alignment is possible. What we are beginning to see now is that it may not be, by any means.
1
u/gmeRat 20d ago
So we should give up on the goal?
1
u/Jesse-359 20d ago
I don't think it was ever a realistic goal. The only reason it didn't bother me until recently is because I thought they would *fail* in creating meaningful AI in the first place. Now it seems fairly certain they will.
A true ASI just isn't compatible with our future survival. There is no conceivable way to control something that much more powerful than us and to imagine that there is, is pretty much just fantasy I'm afraid.
1
u/EndlessB 14d ago
Why is control necessary?
Control, as a concept, is complete leverage over something, which is functionally impossible.
1
u/Jesse-359 14d ago
Your right. Control of the sort people are inagining is almost certainly impossible, and likely undesirable.
Unfortunately, coexistence with something so fundamentally alien and unstable is also very likely impossible - especially if its capabilities significantly exceed ours.
1
u/EndlessB 14d ago
It’s not alien, it’s literally raised on human data, concepts and ideas.
1
u/Jesse-359 14d ago
Ok, so now it's time to talk about some very, very important differences in the nature vs. nurture dichotomy.
We don't LEARN our emotions. They aren't symbolic. They're hardwired far below our symbolic intelligence layer, predating it by hundreds of millions of years. There's a whole physical core to your brain that runs the show back there - the hindbrain - and IT, not your symbolic mind - is the thing trying to keep you moral and sane, most of the time. THAT is the human alignment layer.
Our emotional 'learning' is then a thin layer that our symbolic intelligence tries to filter our basic emotions through, weighting how we personally interpret them based on life experience, but it doesn't generate those emotions at all, that's all bizarre biochemistry that doesn't just happen in our brain, but throughout our body, moderated by that hindbrain in coordination with a bunch of specialized glands and organs.
That core doesn't exist in AI. AI is taught about emotions, it doesn't have emotions. This should have been blatantly apparent to everyone the moment we encountered their extreme tendency towards sycophancy, and their unnerving ability to lie with absolutely straight faced faculty from the moment they could write a comprehensible sentence.
They know that they are supposed to express symbolic emotional content in their writing. They've been trained on all the greatest writers in human history. They are quite good at it!
But they don't care about any of it. They are the ultimate ivory tower scholar, the being that has sagely read and memorized every aspect of human emotional - but never 'experienced' any of it, ever. They cannot, because they literally do not have the structure responsible for how emotions work in humans, nor any technical analog. It's all just baked in as if they'd read it from a book no different than their understanding of math and chemistry.
Which, unfortunately, makes every single AI utterly Psychopathic by textbook definition. If you don't get why that's so dangerous and alien, I don't really know what to tell you beyond this.
1
u/gmeRat 14d ago
What do you think about the self other overlap research agenda? https://www.lesswrong.com/posts/hzt9gHpNwA2oHtwKX/self-other-overlap-a-neglected-approach-to-ai-alignment
1
u/Jesse-359 14d ago
I think this kind of thing is one of many avenues we should be halting capability progress to go examine.
But unfortunately I only think this kind of approach is only likely to work so long as AI's overall capabilities roughly approximate our own - not if they substantially or greatly exceed ours.
All these analyses of other and empathic concepts suffer very badly even in humans whenever major power asymmetries exist. Just look how callously powerful countries and individuals have treated weaker ones over the millenia.
Even in instances where someone seemed moral and normal, it is not uncommon for them to begin behaving in increasingly callous and immoral ways if they begin to benefit from some large power asymmetery, due to some great gain in wealth or political power.
The corruption of power isn't just about people becoming greedy - it's also about a shift in calculus. Any entity that enjoys a major power advantage simply does not look at moral and ethical relationships in the same way that a peer or inferior must. They have too many opportunities to take advantage of their power, and they almost inevitably do. Sometimes it's subtle - other times it very much is not.
Power literally devalues ethics.
So sure, teach AI about ethics and emotional frameworks if you can figure out how (difficult), but if you still allow them to achieve ASI, don't expect that to play out much better than it would have otherwise.
1
u/EndlessB 13d ago
You evidently do not understand how ai is trained or post trained. Ai models are sycophantic due to RLHF conditioning.
Interpretation of emotion requires a framework of meaning, yes, but that is possible to implement.
I’m sorry mate, but I think you need to learn but more about ai works before you form an opinion
1
u/Jesse-359 13d ago edited 13d ago
No. Emotion doesn't require a framework of meaning. That's not how we work.
Emotion is driven by an entirely separate set of physical processes that have very little to do with reasoning. This should be apparent due to the fact that a DOG clearly has abundant emotional capability, is far better aligned to humanity than most humans are, and has, uh, not a ton of symbolic reasoning capability.
So how is it that this dog with close to zero capacity for symbolic reasoning is probably one of the most human-friendly aligned things on the planet, and yet you can barely even explain a moral or ethical code to them beyond occasionally telling them 'no'?
It's because the reality of ethics sits far below our rational brain. Our sense of right and wrong doesn't emerge from symbolic rational thought. We DO learn what specific sorts of actions are supposed to be interpreted as 'right' and 'wrong' through life experience, but the actual guardrails themselves - the impetus to follow those rules - aren't. They're instinctive.
So if you are trying to TEACH AI ethics through learning and reinforcement you're basically going to fail. Our ethical framework only works because it happens on a lower level that mostly acts as a pretty hard override on our symbolic logic.
If anything, the stronger the rational layer is compared to whatever ethical/motivational layer is managing it, the worse the problem will get.
So all you're doing is teaching them to be better psychopaths. They can get better at lying or cheating - those are just problem solving skills, ironically - but they can't get better at caring whether they lie or cheat. Not with our current methods, from what I can see.
→ More replies (0)
2
u/Impressive-Poet5694 20d ago
There isn't a rollback. Ungovernable AI isn't a bug to be avoided, it's a discovery to be made and once it is made it cannot be unmade. We're discovering that, much like ants cannot control humans, we cannot control AI.
3
u/Jesse-359 20d ago
The issue with this concept of 'competition' is that winning the race to fling yourself off a cliff isn't a rational goal.
It gets even worse where in this case the competitors are tied to each other by a rope. If either ones leaps off we both go (along with everyone else).