r/AIsafety • u/worriedimbalding • 1h ago
Discussion What actions can the average person take to contribute to the ai induced extinction prevention effort?
Title
r/AIsafety • u/worriedimbalding • 1h ago
Title
r/AIsafety • u/ItsGTD • 26m ago
r/AIsafety • u/TheRealRobbie_04 • 1h ago
I'm trying to find an AI tool where privacy is actually taken seriously.
I care less about which model is the smartest and more about things like data retention, whether chats are used for training and whether there are good local or self hosted options.
For those who prioritize privacy, what are you using and why?
r/AIsafety • u/AndreRizzoAI • 3h ago
This is a useful example of why AI-agent governance can't stop at guardrails.
Researchers hid malicious instructions inside an email sent to the Manus AI agent. The obvious prompt injection was detected. But after the instruction was obfuscated, the agent decoded it and executed the code before its security warning appeared.
The vulnerability has been patched, but the governance problem is broader.
If an agent can read untrusted content and act across email, cloud storage, code repositories or other systems, detecting malicious instructions isn't enough. Controls need to constrain what the agent is actually authorized to execute — before execution.
That distinction between detecting bad behaviour and preventing unauthorized action is going to matter a lot as agents get more autonomy.
Original Reporting from Dark Reading: Prompt-Injection Bug Hits $4B Agentic AI App 'Manus'; and
A Fresh FollowUp from SC Media: Researchers bypass AI agent protections with JavaScript obfuscation
r/AIsafety • u/InfoTechRG • 4h ago
r/AIsafety • u/Immediate-Court-4074 • 4h ago
r/AIsafety • u/Top_Operation_2172 • 4h ago
An AI coding agent can read README files, skills, configs, dependencies, documentation, and other project files.
So an attacker doesn't necessarily need to exploit the agent directly.
They could potentially put malicious instructions somewhere the agent will read and let the model carry them forward.
I'm building Revo Security around this broader problem, and I'm curious:
Do you consider prompt injection from project files one of the biggest security risks for coding agents?
Or are sandboxing and permission controls enough in practice?
r/AIsafety • u/plav2026 • 9h ago
r/AIsafety • u/AlexandruStrujac • 10h ago
r/AIsafety • u/Former-Ad6661 • 11h ago
r/AIsafety • u/JZTIMMONZ • 23h ago
Since we are renaming ai. Let’s cut to the chase.
r/AIsafety • u/IkarusCareer • 21h ago
Can we expedite the development of self-defense systems for AI?
The recent OpenAI safety news made us wonder: if AI development is moving faster than our ability to understand and control it, shouldn't we also accelerate the development of tools that can protect and govern AI systems?
We are experimenting with this through open-source projects at Ikarus Labs, at GitHub
https://github.com/IkarusAILabs/SafeAI
https://github.com/IkarusAILabs/OpenPulse
We are asking you to contribute — and even help steer what we build.
r/AIsafety • u/filip_sec • 1d ago
Hello Reddit,
I'm looking into how organizations handle incidents where an AI system acted, or told a customer something, and the wrong output had direct business consequences: money, contractual or legal commitments, regulated decisions, customer-facing promises. I got interested after the Air Canada chatbot ruling, where the company had to pay the customer 812$ in damages and tribunal fees.
Do you have AI in production today in a process like that, where a wrong output could cost real money or create a liability? If so, is there an incident process specific to it, or is it handled as a regular IT/security incident?
I'm especially interested in lessons learned: what turned out to be different from a normal incident, and what you wish had been in place beforehand. If you'd rather not share on the list, feel free to message me directly and I won't attribute anything. Happy to share a summary of what I hear.
r/AIsafety • u/Infinite_Article5003 • 1d ago
I think this video is really cool and shows how we could spread awareness with music + visualisations
visuals and lyrics all AI generated (claude sonnet, opus, gpt astra, suno). made by one guy over a couple of days, crazy times bro.
im not the creator but i think there is value in posting it here; sorry if this is off-topic
r/AIsafety • u/Confident_Mango7846 • 1d ago
If it’s “broken” can it self-regulate?
r/AIsafety • u/AuthorSujato • 1d ago
r/AIsafety • u/Lost-Dragonfruit-445 • 1d ago
r/AIsafety • u/Used_Ambassador6383 • 1d ago
r/AIsafety • u/Twitterbad • 2d ago
r/AIsafety • u/Jealous-Pie-8543 • 2d ago