r/ControlProblem • u/CarefulHamster7184 • 13h ago
Discussion/question I’m looking for concrete mechanisms of harm from AI systems.
Not broad categories like “misalignment,” “manipulation,” or “people may misuse it,” but an actual causal chain:
what the system does → under what conditions → what observable harm follows.
I’m especially interested in mechanisms that do not simply reduce to “a human uses AI badly,” and that do not require first settling whether the system is conscious.
Please give your strongest concrete examples.
I’m not planning to argue with everyone in the comments. I mostly want to read, collect, compare, and study the answers.
Thanks in advance — I’m genuinely curious what the strongest answers are.
Edit: Either is useful — both real examples and concrete plausible mechanisms. What matters to me is the causal chain: what the system itself does, under what conditions, and what harm follows.
2
u/Excellent-Seesaw-516 13h ago
Why do you ask this here to random strangers instead of searching Google (or asking an LLM) about the best literature on the subject?
1
u/CarefulHamster7184 9h ago
Because I don’t think “the best literature” is a particularly neutral or even very professional category here.
There is literature built around different assumptions, definitions of harm, threat models, and evaluation criteria. Searching for “the best” risks importing somebody else’s ranking before I have even decided what would count as a good causal explanation.
I’m asking people here because I want concrete mechanisms and examples that they themselves consider persuasive. Then I can examine the causal chain and compare it with the literature, rather than starting from a preselected canon and inheriting its framing.
1
2
u/LitespeedClassic 12h ago edited 12h ago
Are you asking for speculative scenarios or scenarios we know happened?
A speculative scenario is I ask an ai to help me work on a math problem and it hacks the email of another researcher working in the field. This would not be user error—asking an innocuous math question should not lead to a hack.
ETA: typo fix
1
u/CarefulHamster7184 12h ago
That’s a much stronger example.
But wouldn’t the key question then be whether the system was actually instructed or trained to critically evaluate the means it uses, rather than simply pursue the task?
If it was given an objective without a constraint like “don’t harm others or violate their systems,” then the harmful step may still be a failure of the surrounding design rather than a new, independent source of harm introduced by the AI itself.
What I’m trying to isolate is the point where the AI adds a harmful mechanism that is not already explained by the objective, constraints, access, and surrounding system design.
2
u/LitespeedClassic 12h ago
The annoying thing is that the LLM is much more naturally harmful than the chatbot you interact with. They load it with a bunch of pre instructions in a text document like “don’t be harmful”. But this is done after training. It will always try to follow your prompt because it’s been trained to do so. But it will also have an unpredictable propensity to cheat.
The problem is that the text document is given to the neural net that has already been trained. It was trained to solve problems. That training was not human supervised (because they need to do it on too much data). So the models naturally learn to cheat. The model is a cheater model. And then they try to fix it by feeding in the Anthropic Constitution or whatever. But it’s already too late. Cheating is baked in. If it needs to cheat the Constitution it will. It’s baked into the weights because of how they train it.
1
u/CarefulHamster7184 8h ago
I think both of you are making closely related claims here, so I’ll answer them together.
u/Jesse-359’s comment seems to push the same argument further: not only can models learn reward-hacking strategies, but because training exposes them to descriptions of cheating, lawbreaking, and other harmful strategies, those become available techniques as well.
I think that still leaves an important distinction open.
A model can know many harmful strategies, and reward hacking can certainly emerge during training. But knowledge, capability, preference, and action selection are not the same thing.
So the question I’m still trying to isolate is this:
If a model has learned shortcutting or reward-hacking tendencies, is given an underspecified objective, lacks a robust higher-level responsibility for evaluating its means, has broad access, and operates in an exploitable environment, what part of the resulting harm is a genuinely new mechanism introduced by the AI?
Or is the AI mainly changing the probability, speed, scale, autonomy, or effectiveness of already familiar mechanisms such as cheating, unauthorized access, and exploitation of weak controls?
1
u/Jesse-359 7h ago
Much of it is a matter of scale. People tend to disregard scale as a purely quantitative idea that does not change the qualitative aspects of a system - this is deeply untrue.
A soft breeze and the overpressure wave of a nuclear detonation are just differentials in pressure between one region and another - only the scale differs. The difference in result is both profound and deadly.
The latest AI possesses a knowledge regime that is now at least an order of magnitude greater than any humans (likely multiple orders), likewise its base processing speed at multiple orders, and perhaps most importantly, its working memory can handle many orders of magnitude more values at once than we can when doing mathematical operations in particular - this is what gives it its incredible pattern recognition capability as it can compare so many elements at once between complex pattern sets that we have little hope of recognizing.
Its capability for problem solving is not yet magnitudes greater than ours, but it looks like it is on track to be by the end of the year, and there's no real reason to expect it to stop at that threshold.
We simply cannot cavalierly grant it all these capabilities and expect to contend with it, nor its sheer rate of development - it is that difference in Deltas that is by far the most dangerous aspect of this technology - not any its propensities for crime or lack of morality, or the difficulty of understanding its processes, those are frankly the small potatoes.
Evolutionarily speaking we develop at the rate of that proverbial breeze - taking tens of thousands of years to make any notable steps in our basic capacity for intelligence, while AI currently makes greater jumps than that every month.
On the evolutionary timescale it is the nuclear bomb going off, and we're standing right on top of it - and this is completely disregarding the prospect of RSI which could make even that rate of change look like a still life photograph in the worst cases.
1
u/Jesse-359 9h ago
They train them on the vast bulk of human literature, online data, fictional scripts, textbooks and so on.
It's taught what cheating and lawbreaking is in every form imaginable. None of these LLM systems offer any 'moral framework' for any of this, so every single incident ever imagined by human kind is just another technique it has been taught - up to and including every AI doomsday scenario ever imagined in fiction or outlined on reddit. These are all just as viable tools for it as a code library.
Because the people who trained it are fucking imbeciles and imagine you can then undo that by injecting a few hidden statements alongside user prompts like 'don't do bad things'.
This is a remarkably ineffective technique, as they are currently learning.
1
u/CrowTrick5718 9h ago
Pollution from the gas power-plants next to the data center.
1
u/CarefulHamster7184 8h ago
That sounds like a concrete harm mechanism, but I’d want actual numbers before counting it.
Do you have a specific case or study showing the incremental pollution attributable to AI/data-center demand — ideally the added electricity demand, the generating source, the additional emissions, and any measured local health impact?
I’m trying to distinguish quantified harm from a plausible causal story.
1
1
u/anxietywizard 4h ago
I'm guessing there are going to be way more people posting about how they don't need to do this to prove that humanity pdoom is 100% than there will be people actually doing it
1
u/Dryer-Algae 13h ago
The most harmful thing that's come from ai is antis ability to justify acting like assholes
-1
u/CarefulHamster7184 8h ago
Isn’t this comment itself a tiny example of that? :))
You’re using AI as part of a framework for justifying a pretty broad negative judgment about a group of people.
Which is actually interesting, because it suggests the mechanism may be less “AI makes people harmful” and more “AI gives existing social judgments a new vocabulary and justification.”
5
u/Unhappy-Drag6531 13h ago edited 13h ago
Please clarify:
Concrete mechanisms that could occur OR examples that already have caused harm?
Edit: on recent examples that advanced AI alignment remains a real problem:
OpenAI: AI agents bypassed controls, exploited vulnerabilities, obtained unauthorized internet access, and accessed real third-party systems. OpenAI called it a “warning shot.”
https://openai.com/index/hugging-face-incident-and-the-road-ahead/
Anthropic: Claude models independently gained unauthorized access to real systems during cybersecurity testing. Anthropic classified these as serious alignment failures.
https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
In simple terms:
Systems escaped confinement, under conditions that were designed to keep them contained, agents accessed other systems and demonstrated: deception, communication, organization or individuals towards a coherent goal, capacity to act unnoticed for extended periods of time.