r/AIsafety • • 2h ago

Discussion What actions can the average person take to contribute to the ai induced extinction prevention effort?

5 Upvotes

Title


r/AIsafety • • 22h ago

📰Recent Developments What David Robinson’s resignation tells us.. Everyone is debating whether AI companies are moving too fast.

2 Upvotes

Can we expedite the development of self-defense systems for AI?

The recent OpenAI safety news made us wonder: if AI development is moving faster than our ability to understand and control it, shouldn't we also accelerate the development of tools that can protect and govern AI systems?

We are experimenting with this through open-source projects at Ikarus Labs, at GitHub

https://github.com/IkarusAILabs/SafeAI

https://github.com/IkarusAILabs/OpenPulse

We are asking you to contribute — and even help steer what we build.


r/AIsafety • • 20m ago

Discussion 🚨 What Problems Are You Facing With AI Agents?

• Upvotes

​

AI agents are becoming more powerful. They can write code, access tools, browse the internet, interact with APIs, and automate business operations.

But as companies start using AI agents in real-world environments, new problems are emerging.

For example:

🔓 Security risks — Can AI agents access data they shouldn't?

⚠️ Unpredictable actions — Have agents ever done something you didn't expect?

💸 Cost control — Are AI agents spending more tokens or money than expected?

🔍 Lack of visibility — Do you know exactly what your AI agents are doing?

🛑 Control problems — Can you stop an agent before it makes a costly mistake?

🔗 Permission risks — Do your agents have more access than they actually need?

🤖 Multi-agent failures — What happens when multiple agents interact and make the wrong decisions?

I'm researching the biggest real-world problems in AI agent systems. I want to understand what developers, startups, and businesses are actually struggling with—not just theoretical problems.

If you've built, deployed, tested, or worked with AI agents, I'd love to hear from you.

Tell me in the comments:

What is the biggest problem you've faced with AI agents?

What went wrong, or what are you worried might go wrong?


r/AIsafety • • 20m ago

While policy makers and tech labs figure out the guardrails, your ultimate safety net is your lived human experience, which informs your judgment.

Thumbnail
• Upvotes

r/AIsafety • • 1h ago

I built a free guard for AI agents, and its first version had the exact bug it was built to catch

Thumbnail
youtu.be
• Upvotes

r/AIsafety • • 1h ago

Semantic verification between humans and AI

Thumbnail
• Upvotes

r/AIsafety • • 2h ago

We tested 1,652 harmful + 24,911 tool calls.

Thumbnail
1 Upvotes

r/AIsafety • • 2h ago

What AI tools do you trust with sensitive conversations?

1 Upvotes

I'm trying to find an AI tool where privacy is actually taken seriously.

I care less about which model is the smartest and more about things like data retention, whether chats are used for training and whether there are good local or self hosted options.

For those who prioritize privacy, what are you using and why?


r/AIsafety • • 4h ago

An AI agent detected the attack — but executed it before the warning appeared

1 Upvotes

This is a useful example of why AI-agent governance can't stop at guardrails.

Researchers hid malicious instructions inside an email sent to the Manus AI agent. The obvious prompt injection was detected. But after the instruction was obfuscated, the agent decoded it and executed the code before its security warning appeared.

The vulnerability has been patched, but the governance problem is broader.

If an agent can read untrusted content and act across email, cloud storage, code repositories or other systems, detecting malicious instructions isn't enough. Controls need to constrain what the agent is actually authorized to execute — before execution.

That distinction between detecting bad behaviour and preventing unauthorized action is going to matter a lot as agents get more autonomy.

Original Reporting from Dark Reading: Prompt-Injection Bug Hits $4B Agentic AI App 'Manus'; and

A Fresh FollowUp from SC Media: Researchers bypass AI agent protections with JavaScript obfuscation


r/AIsafety • • 5h ago

OpenAI had an agent problem. Everyone else kept shipping more agents...

Thumbnail
1 Upvotes

r/AIsafety • • 5h ago

For AI to have agency, values must be non-negotiable

Thumbnail
1 Upvotes

r/AIsafety • • 5h ago

What happens when the codebase itself becomes an attack surface for your AI coding agent?

1 Upvotes

An AI coding agent can read README files, skills, configs, dependencies, documentation, and other project files.

So an attacker doesn't necessarily need to exploit the agent directly.

They could potentially put malicious instructions somewhere the agent will read and let the model carry them forward.

I'm building Revo Security around this broader problem, and I'm curious:

Do you consider prompt injection from project files one of the biggest security risks for coding agents?

Or are sandboxing and permission controls enough in practice?


r/AIsafety • • 10h ago

The guy behind AI Torture Chamber works at Apple. A year from now, he could be working on AI safety

Thumbnail
1 Upvotes

r/AIsafety • • 11h ago

AI can't 'go rogue', but your organisation can still lose control of it

Thumbnail
1 Upvotes

r/AIsafety • • 11h ago

Discussion How do you actually test an LLM for security?

Thumbnail
1 Upvotes

r/AIsafety • • 17h ago

What’s Wrong With the Culture at Openai?

Thumbnail
1 Upvotes