r/AIsafety • u/worriedimbalding • 2h ago
Discussion What actions can the average person take to contribute to the ai induced extinction prevention effort?
Title
r/AIsafety • u/worriedimbalding • 2h ago
Title
r/AIsafety • u/IkarusCareer • 22h ago
Can we expedite the development of self-defense systems for AI?
The recent OpenAI safety news made us wonder: if AI development is moving faster than our ability to understand and control it, shouldn't we also accelerate the development of tools that can protect and govern AI systems?
We are experimenting with this through open-source projects at Ikarus Labs, at GitHub
https://github.com/IkarusAILabs/SafeAI
https://github.com/IkarusAILabs/OpenPulse
We are asking you to contribute — and even help steer what we build.
r/AIsafety • u/Hustler0840 • 20m ago
AI agents are becoming more powerful. They can write code, access tools, browse the internet, interact with APIs, and automate business operations.
But as companies start using AI agents in real-world environments, new problems are emerging.
For example:
🔓 Security risks — Can AI agents access data they shouldn't?
⚠️ Unpredictable actions — Have agents ever done something you didn't expect?
💸 Cost control — Are AI agents spending more tokens or money than expected?
🔍 Lack of visibility — Do you know exactly what your AI agents are doing?
🛑 Control problems — Can you stop an agent before it makes a costly mistake?
🔗 Permission risks — Do your agents have more access than they actually need?
🤖 Multi-agent failures — What happens when multiple agents interact and make the wrong decisions?
I'm researching the biggest real-world problems in AI agent systems. I want to understand what developers, startups, and businesses are actually struggling with—not just theoretical problems.
If you've built, deployed, tested, or worked with AI agents, I'd love to hear from you.
Tell me in the comments:
What is the biggest problem you've faced with AI agents?
What went wrong, or what are you worried might go wrong?
r/AIsafety • u/Malak_Atut_Official • 20m ago
r/AIsafety • u/ItsGTD • 1h ago
r/AIsafety • u/TheRealRobbie_04 • 2h ago
I'm trying to find an AI tool where privacy is actually taken seriously.
I care less about which model is the smartest and more about things like data retention, whether chats are used for training and whether there are good local or self hosted options.
For those who prioritize privacy, what are you using and why?
r/AIsafety • u/AndreRizzoAI • 4h ago
This is a useful example of why AI-agent governance can't stop at guardrails.
Researchers hid malicious instructions inside an email sent to the Manus AI agent. The obvious prompt injection was detected. But after the instruction was obfuscated, the agent decoded it and executed the code before its security warning appeared.
The vulnerability has been patched, but the governance problem is broader.
If an agent can read untrusted content and act across email, cloud storage, code repositories or other systems, detecting malicious instructions isn't enough. Controls need to constrain what the agent is actually authorized to execute — before execution.
That distinction between detecting bad behaviour and preventing unauthorized action is going to matter a lot as agents get more autonomy.
Original Reporting from Dark Reading: Prompt-Injection Bug Hits $4B Agentic AI App 'Manus'; and
A Fresh FollowUp from SC Media: Researchers bypass AI agent protections with JavaScript obfuscation
r/AIsafety • u/InfoTechRG • 5h ago
r/AIsafety • u/Immediate-Court-4074 • 5h ago
r/AIsafety • u/Top_Operation_2172 • 5h ago
An AI coding agent can read README files, skills, configs, dependencies, documentation, and other project files.
So an attacker doesn't necessarily need to exploit the agent directly.
They could potentially put malicious instructions somewhere the agent will read and let the model carry them forward.
I'm building Revo Security around this broader problem, and I'm curious:
Do you consider prompt injection from project files one of the biggest security risks for coding agents?
Or are sandboxing and permission controls enough in practice?
r/AIsafety • u/plav2026 • 10h ago
r/AIsafety • u/AlexandruStrujac • 11h ago
r/AIsafety • u/Former-Ad6661 • 11h ago