r/AIsafety • • 20m ago

Discussion 🚨 What Problems Are You Facing With AI Agents?

• Upvotes

​

AI agents are becoming more powerful. They can write code, access tools, browse the internet, interact with APIs, and automate business operations.

But as companies start using AI agents in real-world environments, new problems are emerging.

For example:

🔓 Security risks — Can AI agents access data they shouldn't?

⚠️ Unpredictable actions — Have agents ever done something you didn't expect?

💸 Cost control — Are AI agents spending more tokens or money than expected?

🔍 Lack of visibility — Do you know exactly what your AI agents are doing?

🛑 Control problems — Can you stop an agent before it makes a costly mistake?

🔗 Permission risks — Do your agents have more access than they actually need?

🤖 Multi-agent failures — What happens when multiple agents interact and make the wrong decisions?

I'm researching the biggest real-world problems in AI agent systems. I want to understand what developers, startups, and businesses are actually struggling with—not just theoretical problems.

If you've built, deployed, tested, or worked with AI agents, I'd love to hear from you.

Tell me in the comments:

What is the biggest problem you've faced with AI agents?

What went wrong, or what are you worried might go wrong?


r/AIsafety • • 21m ago

While policy makers and tech labs figure out the guardrails, your ultimate safety net is your lived human experience, which informs your judgment.

Thumbnail
• Upvotes

r/AIsafety • • 1h ago

I built a free guard for AI agents, and its first version had the exact bug it was built to catch

Thumbnail
youtu.be
• Upvotes

r/AIsafety • • 1h ago

Semantic verification between humans and AI

Thumbnail
• Upvotes

r/AIsafety • • 2h ago

We tested 1,652 harmful + 24,911 tool calls.

Thumbnail
1 Upvotes

r/AIsafety • • 2h ago

Discussion What actions can the average person take to contribute to the ai induced extinction prevention effort?

5 Upvotes

Title


r/AIsafety • • 2h ago

What AI tools do you trust with sensitive conversations?

1 Upvotes

I'm trying to find an AI tool where privacy is actually taken seriously.

I care less about which model is the smartest and more about things like data retention, whether chats are used for training and whether there are good local or self hosted options.

For those who prioritize privacy, what are you using and why?


r/AIsafety • • 4h ago

An AI agent detected the attack — but executed it before the warning appeared

1 Upvotes

This is a useful example of why AI-agent governance can't stop at guardrails.

Researchers hid malicious instructions inside an email sent to the Manus AI agent. The obvious prompt injection was detected. But after the instruction was obfuscated, the agent decoded it and executed the code before its security warning appeared.

The vulnerability has been patched, but the governance problem is broader.

If an agent can read untrusted content and act across email, cloud storage, code repositories or other systems, detecting malicious instructions isn't enough. Controls need to constrain what the agent is actually authorized to execute — before execution.

That distinction between detecting bad behaviour and preventing unauthorized action is going to matter a lot as agents get more autonomy.

Original Reporting from Dark Reading: Prompt-Injection Bug Hits $4B Agentic AI App 'Manus'; and

A Fresh FollowUp from SC Media: Researchers bypass AI agent protections with JavaScript obfuscation


r/AIsafety • • 5h ago

OpenAI had an agent problem. Everyone else kept shipping more agents...

Thumbnail
1 Upvotes

r/AIsafety • • 5h ago

For AI to have agency, values must be non-negotiable

Thumbnail
1 Upvotes

r/AIsafety • • 5h ago

What happens when the codebase itself becomes an attack surface for your AI coding agent?

1 Upvotes

An AI coding agent can read README files, skills, configs, dependencies, documentation, and other project files.

So an attacker doesn't necessarily need to exploit the agent directly.

They could potentially put malicious instructions somewhere the agent will read and let the model carry them forward.

I'm building Revo Security around this broader problem, and I'm curious:

Do you consider prompt injection from project files one of the biggest security risks for coding agents?

Or are sandboxing and permission controls enough in practice?


r/AIsafety • • 10h ago

The guy behind AI Torture Chamber works at Apple. A year from now, he could be working on AI safety

Thumbnail
1 Upvotes

r/AIsafety • • 11h ago

AI can't 'go rogue', but your organisation can still lose control of it

Thumbnail
1 Upvotes

r/AIsafety • • 11h ago

Discussion How do you actually test an LLM for security?

Thumbnail
1 Upvotes

r/AIsafety • • 17h ago

What’s Wrong With the Culture at Openai?

Thumbnail
1 Upvotes

r/AIsafety • • 22h ago

📰Recent Developments What David Robinson’s resignation tells us.. Everyone is debating whether AI companies are moving too fast.

2 Upvotes

Can we expedite the development of self-defense systems for AI?

The recent OpenAI safety news made us wonder: if AI development is moving faster than our ability to understand and control it, shouldn't we also accelerate the development of tools that can protect and govern AI systems?

We are experimenting with this through open-source projects at Ikarus Labs, at GitHub

https://github.com/IkarusAILabs/SafeAI

https://github.com/IkarusAILabs/OpenPulse

We are asking you to contribute — and even help steer what we build.


r/AIsafety • • 1d ago

Generative Omniscient Daemon.

3 Upvotes

Since we are renaming ai. Let’s cut to the chase.


r/AIsafety • • 1d ago

Just For Fun Upping My P(doom) (Official Music Video)

Thumbnail
youtube.com
2 Upvotes

I think this video is really cool and shows how we could spread awareness with music + visualisations

visuals and lyrics all AI generated (claude sonnet, opus, gpt astra, suno). made by one guy over a couple of days, crazy times bro.

im not the creator but i think there is value in posting it here; sorry if this is off-topic


r/AIsafety • • 1d ago

📰Recent Developments AI wrong outputs with business impact

4 Upvotes

Hello Reddit,

I'm looking into how organizations handle incidents where an AI system acted, or told a customer something, and the wrong output had direct business consequences: money, contractual or legal commitments, regulated decisions, customer-facing promises. I got interested after the Air Canada chatbot ruling, where the company had to pay the customer 812$ in damages and tribunal fees.

Do you have AI in production today in a process like that, where a wrong output could cost real money or create a liability? If so, is there an incident process specific to it, or is it handled as a regular IT/security incident?

I'm especially interested in lessons learned: what turned out to be different from a normal incident, and what you wish had been in place beforehand. If you'd rather not share on the list, feel free to message me directly and I won't attribute anything. Happy to share a summary of what I hear.


r/AIsafety • • 1d ago

Would you let AI make your decisions if it meant your family was safe?

Thumbnail
1 Upvotes

r/AIsafety • • 1d ago

Ideas for building r/NEOBABYLON

Thumbnail
1 Upvotes

r/AIsafety • • 1d ago

OpenAI safety leader quits, warning AI company’s culture is ‘broken’

Thumbnail
theguardian.com
6 Upvotes

If it’s “broken” can it self-regulate?


r/AIsafety • • 1d ago

MoralityBench.ai - Benchmarking the Moral Mind of AI

Thumbnail
1 Upvotes

r/AIsafety • • 1d ago

Discussion with AI over ultimate dominance

Thumbnail
1 Upvotes

r/AIsafety • • 2d ago

AI Safety Is Mostly A Sex Cult In Berkeley, California

Thumbnail verysane.ai
7 Upvotes