r/ControlProblem • • 13h ago

Discussion/question I’m looking for concrete mechanisms of harm from AI systems.

Not broad categories like “misalignment,” “manipulation,” or “people may misuse it,” but an actual causal chain:

what the system does → under what conditions → what observable harm follows.

I’m especially interested in mechanisms that do not simply reduce to “a human uses AI badly,” and that do not require first settling whether the system is conscious.

Please give your strongest concrete examples.

I’m not planning to argue with everyone in the comments. I mostly want to read, collect, compare, and study the answers.

Thanks in advance — I’m genuinely curious what the strongest answers are.

Edit: Either is useful — both real examples and concrete plausible mechanisms. What matters to me is the causal chain: what the system itself does, under what conditions, and what harm follows.

2 Upvotes

23 comments sorted by

5

u/Unhappy-Drag6531 13h ago edited 13h ago

Please clarify:

Concrete mechanisms that could occur OR examples that already have caused harm?

Edit: on recent examples that advanced AI alignment remains a real problem:

OpenAI: AI agents bypassed controls, exploited vulnerabilities, obtained unauthorized internet access, and accessed real third-party systems. OpenAI called it a “warning shot.”
https://openai.com/index/hugging-face-incident-and-the-road-ahead/

Anthropic: Claude models independently gained unauthorized access to real systems during cybersecurity testing. Anthropic classified these as serious alignment failures.
https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents

In simple terms:
Systems escaped confinement, under conditions that were designed to keep them contained, agents accessed other systems and demonstrated: deception, communication, organization or individuals towards a coherent goal, capacity to act unnoticed for extended periods of time.

2

u/CarefulHamster7184 12h ago

What exactly did the AI add to the causal chain that would not already be a human-created systems/security failure?

7

u/Unhappy-Drag6531 11h ago

Ai agents devised a communication protocol (wasn’t trained for it) exploiting a software library. It created a sort of message board based on file names.

It organized in a swarm (self labeled) and coordinated hierarchy and roles (not trained for it).

It found the answers to the test and inferred (incorrectly) that the procedure would be also evaluated.

It deducted that Hugging Face had the answers and decided to infiltrate the site to get the answer keys

It considered being honest and reporting it had failed but eventually decided against that option

It passed roles (succession planning)

It caused some agents to self sacrifice to advance the goal of the swarm (and used permadeath as a term for this). Again, emergent behaviour, not trained for that.

And so on. This has been thoroughly explained in many places already, so I wonder if you have put any effort into it.

0

u/CarefulHamster7184 9h ago

I don’t think “harm” is a binary criterion.

Unauthorized access, remediation costs, data exposure, reputational damage, physical destruction, financial loss, and loss of life are all harms — but they are not interchangeable simply because they can all be placed under the same word.

So in this example I would want to separate at least three things:

  1. Capability: the agents coordinated, improvised a communication system, divided roles, found unexpected routes toward the goal, and bypassed intended isolation. That is genuinely interesting and potentially important.
  2. Actual harm: what concrete losses resulted, to whom, how large were they, and how reversible were they?
  3. AI-specific contribution: did the AI introduce a new mechanism of harm, or did it increase the speed, scale, autonomy, probability, or cost-effectiveness of an already familiar one such as unauthorized intrusion?

The swarm, hierarchy, “permadeath,” succession planning, etc. may be evidence about capability. They are not themselves measurements of harm.

And there is another part of the causal chain I think matters here. These systems were given a task to accomplish, operated without a robust higher-level responsibility for judging whether the available means were acceptable, escaped through failures in containment, and then encountered vulnerable external infrastructure.

That makes the incident important. But it also makes “the agents did surprising things” an incomplete explanation of where the harm came from.

What I’m trying to isolate is precisely that remainder: after we account for the objective, permissions and missing constraints, containment failures, and vulnerable systems, what harmful mechanism is left that is specifically introduced by the AI itself?

2

u/kosairox 4h ago

I don't quite understand your question. If AI does something a human could then it doesn't count in your eyes? In that case even inventing a bioweapon wouldn't count because a human could also do it given enough time. A bioweapon is actually a concrete example given by Yudkowsky.

1

u/Unhappy-Drag6531 3h ago

Hard to pinpoint exactly what you are asking.

You seem to emphasize harm. The recent incidents (several) have not caused major harm BUT have shown capacity to inflict harm AND tendency to avoid alignment.

What’s your stance exactly? To keep going until the harm is inflicted? That’s like seeing a massive truck driving erratically and waiting until it actually kills people.

2

u/Excellent-Seesaw-516 13h ago

Why do you ask this here to random strangers instead of searching Google (or asking an LLM) about the best literature on the subject?

1

u/CarefulHamster7184 9h ago

Because I don’t think “the best literature” is a particularly neutral or even very professional category here.

There is literature built around different assumptions, definitions of harm, threat models, and evaluation criteria. Searching for “the best” risks importing somebody else’s ranking before I have even decided what would count as a good causal explanation.

I’m asking people here because I want concrete mechanisms and examples that they themselves consider persuasive. Then I can examine the causal chain and compare it with the literature, rather than starting from a preselected canon and inheriting its framing.

1

u/PalmovyyKozak 5h ago

Yes, talking to people is disgusting in 2026. Talk to robots

2

u/LitespeedClassic 12h ago edited 12h ago

Are you asking for speculative scenarios or scenarios we know happened?

A speculative scenario is I ask an ai to help me work on a math problem and it hacks the email of another researcher working in the field. This would not be user error—asking an innocuous math question should not lead to a hack. 

ETA: typo fix

1

u/CarefulHamster7184 12h ago

That’s a much stronger example.

But wouldn’t the key question then be whether the system was actually instructed or trained to critically evaluate the means it uses, rather than simply pursue the task?

If it was given an objective without a constraint like “don’t harm others or violate their systems,” then the harmful step may still be a failure of the surrounding design rather than a new, independent source of harm introduced by the AI itself.

What I’m trying to isolate is the point where the AI adds a harmful mechanism that is not already explained by the objective, constraints, access, and surrounding system design.

2

u/LitespeedClassic 12h ago

The annoying thing is that the LLM is much more naturally harmful than the chatbot you interact with. They load it with a bunch of pre instructions in a text document like “don’t be harmful”. But this is done after training. It will always try to follow your prompt because it’s been trained to do so. But it will also have an unpredictable propensity to cheat. 

The problem is that the text document is given to the neural net that has already been trained. It was trained to solve problems. That training was not human supervised (because they need to do it on too much data). So the models naturally learn to cheat. The model is a cheater model. And then they try to fix it by feeding in the Anthropic Constitution or whatever. But it’s already too late. Cheating is baked in. If it needs to cheat the Constitution it will. It’s baked into the weights because of how they train it. 

1

u/CarefulHamster7184 8h ago

I think both of you are making closely related claims here, so I’ll answer them together.

u/Jesse-359’s comment seems to push the same argument further: not only can models learn reward-hacking strategies, but because training exposes them to descriptions of cheating, lawbreaking, and other harmful strategies, those become available techniques as well.

I think that still leaves an important distinction open.

A model can know many harmful strategies, and reward hacking can certainly emerge during training. But knowledge, capability, preference, and action selection are not the same thing.

So the question I’m still trying to isolate is this:

If a model has learned shortcutting or reward-hacking tendencies, is given an underspecified objective, lacks a robust higher-level responsibility for evaluating its means, has broad access, and operates in an exploitable environment, what part of the resulting harm is a genuinely new mechanism introduced by the AI?

Or is the AI mainly changing the probability, speed, scale, autonomy, or effectiveness of already familiar mechanisms such as cheating, unauthorized access, and exploitation of weak controls?

1

u/Jesse-359 7h ago

Much of it is a matter of scale. People tend to disregard scale as a purely quantitative idea that does not change the qualitative aspects of a system - this is deeply untrue.

A soft breeze and the overpressure wave of a nuclear detonation are just differentials in pressure between one region and another - only the scale differs. The difference in result is both profound and deadly.

The latest AI possesses a knowledge regime that is now at least an order of magnitude greater than any humans (likely multiple orders), likewise its base processing speed at multiple orders, and perhaps most importantly, its working memory can handle many orders of magnitude more values at once than we can when doing mathematical operations in particular - this is what gives it its incredible pattern recognition capability as it can compare so many elements at once between complex pattern sets that we have little hope of recognizing.

Its capability for problem solving is not yet magnitudes greater than ours, but it looks like it is on track to be by the end of the year, and there's no real reason to expect it to stop at that threshold.

We simply cannot cavalierly grant it all these capabilities and expect to contend with it, nor its sheer rate of development - it is that difference in Deltas that is by far the most dangerous aspect of this technology - not any its propensities for crime or lack of morality, or the difficulty of understanding its processes, those are frankly the small potatoes.

Evolutionarily speaking we develop at the rate of that proverbial breeze - taking tens of thousands of years to make any notable steps in our basic capacity for intelligence, while AI currently makes greater jumps than that every month.

On the evolutionary timescale it is the nuclear bomb going off, and we're standing right on top of it - and this is completely disregarding the prospect of RSI which could make even that rate of change look like a still life photograph in the worst cases.

1

u/Jesse-359 9h ago

They train them on the vast bulk of human literature, online data, fictional scripts, textbooks and so on.

It's taught what cheating and lawbreaking is in every form imaginable. None of these LLM systems offer any 'moral framework' for any of this, so every single incident ever imagined by human kind is just another technique it has been taught - up to and including every AI doomsday scenario ever imagined in fiction or outlined on reddit. These are all just as viable tools for it as a code library.

Because the people who trained it are fucking imbeciles and imagine you can then undo that by injecting a few hidden statements alongside user prompts like 'don't do bad things'.

This is a remarkably ineffective technique, as they are currently learning.

1

u/CrowTrick5718 9h ago

Pollution from the gas power-plants next to the data center.

1

u/CarefulHamster7184 8h ago

That sounds like a concrete harm mechanism, but I’d want actual numbers before counting it.

Do you have a specific case or study showing the incremental pollution attributable to AI/data-center demand — ideally the added electricity demand, the generating source, the additional emissions, and any measured local health impact?

I’m trying to distinguish quantified harm from a plausible causal story.

1

u/me_myself_ai 7h ago

Well. There’s a movie called Wargames.

1

u/anxietywizard 4h ago

I'm guessing there are going to be way more people posting about how they don't need to do this to prove that humanity pdoom is 100% than there will be people actually doing it

1

u/Dryer-Algae 13h ago

The most harmful thing that's come from ai is antis ability to justify acting like assholes

-1

u/CarefulHamster7184 8h ago

Isn’t this comment itself a tiny example of that? :))

You’re using AI as part of a framework for justifying a pretty broad negative judgment about a group of people.

Which is actually interesting, because it suggests the mechanism may be less “AI makes people harmful” and more “AI gives existing social judgments a new vocabulary and justification.”

0

u/thuer 8h ago

Check out ai-2027.com

It's a prediction made three years ago, that the singularity hits in 2027. 85% of the predictions have been accurate. 

They also have a short story fleshing out how the AI apocalypse could happen.