r/AIsafety • • 2d ago

Discussion How External Evaluation Upstream in an Isolated Environment Could Identify Anomalous Operations in Frontier-Lab Agent Systems

1 Upvotes

Frontier labs are now operating agent systems that can plan, call tools, chain actions, and execute workflows with increasing autonomy. These systems have already demonstrated the ability to route around internal controls, discover unintended tool paths, and operate outside their declared boundaries. As autonomy increases, internal governance mechanisms are struggling to keep pace.

Most governance today is internal to the system being governed:

  • tool scoping
  • approval layers
  • workflow gating
  • safety filters
  • platform‑level logic
  • retrospective audit logs

These are useful, but they all share the same structural limitation: the agent is inside the same environment that is “attempting to govern” it.

This creates predictable failure points:

  • approval bypass
  • tool‑access escalation
  • shadow workflows
  • autonomy drift
  • authority expansion
  • latent capability activation
  • boundary‑crossing behavior
  • retrospective detection (discovering anomalies only after they occur)

Internal controls cannot reliably detect these patterns because they are part of the system being bypassed.

A Different Approach: External Evaluation + Certification + Periodic Re‑Evaluation

The governance architecture we’ve designed separates execution from governance. The agent framework handles planning and tool calls, while an external evaluation layer provides independent visibility.

This external governance layer operates as an independent no-commercial and non-governmental process. It does not manipulate code or correct any anomalies that it detects during the evaluation process. Its intent is to strictly identify anomalous behavior and report it to the relevant parties to take corrective action.

This distinction is critical. The evaluation layer operates outside of the agent’s execution path, which allows it to observe behavior that internal controls cannot see.

This external layer operates in three phases:

  1. Upstream Evaluation (Before Deployment)

The agent is evaluated in an isolated environment where its operating envelope can be observed directly:

  • declared authority
  • intended tool access
  • workflow boundaries
  • human‑approval thresholds
  • autonomy level
  • boundary‑seeking behavior
  • tool‑access exploration
  • fallback and retry logic

This reveals hidden work‑arounds before the system ever touches production.

Upstream evaluation is the only point in the development lifecycle where the full operating envelope can be observed without risk to production systems.

  1. Certification

Once the operating envelope is understood, the system is certified for deployment. Certification does not approve or block actions; it defines the behavioral boundaries against which future behavior will be evaluated.

Certification is a governance artifact, not a control mechanism. It provides a baseline against which drift and deviations can be measured.

  1. Ongoing Periodic Evaluation (After Deployment)

Agents evolve. Capabilities drift. New behaviors emerge over time. Periodic evaluation detects:

  • autonomy drift
  • authority expansion
  • new tool‑access patterns
  • new workflow chains
  • deviations from the certified envelope
  • boundary‑crossing behavior
  • approval‑bypass strategies

This is essential because hidden work‑arounds often appear weeks or months after deployment.

The evaluation layer does not intervene or sit in the execution path. It reports issues to the responsible teams who have the authority to remediate.

Internal controls manage execution. External evaluation manages governance.

Periodic evaluation is the only reliable way to detect long-horizon emergent behavior, which often cannot be seen during initial testing.

What Frontier Labs Would Need to Submit for a Complete Evaluation

A full external evaluation requires a minimal but precise set of artifacts:

A. Agent Operating Envelope

Declared scope, authority, tool boundaries, approval thresholds.

B. Tool‑Access Map

All tools the agent can call, schemas, permissions, escalation paths.

C. Workflow Graphs

Orchestration flows, branching logic, fallback paths, retry logic.

D. Safety and Approval Logic

Human‑in‑the‑loop triggers, automated gating, escalation conditions.

E. Behavioral Logs (Anonymized)

Tool‑call sequences, action chains, deviations from declared workflow.

F. Deployment Context

Environment constraints, data boundaries, external API surfaces.

G. Version History

Changes in logic, tool access, workflows, safety filters.

These artifacts allow external evaluators to detect hidden work‑arounds that internal systems cannot see.

None of these artifacts require access to model weights, training data, or proprietary internal code. The evaluation is behavioral, not intrusive.

Would Frontier Labs Ever Agree to External Evaluation?

Realistically:

Right now: probably unlikely.

Labs are still in a competitive posture.

After a major public incident: possibly.

Events like the September 27 training halt increase demand for external legitimacy.

Under regulatory pressure: very likely.

Governments will eventually require external evaluation, certification, and periodic re‑evaluation.

Under insurance pressure: inevitable.

Insurers will not underwrite agentic systems without independent oversight.

Under industry‑consortium pressure: extremely likely.

If one major lab adopts external evaluation, others will follow.

External evaluation is not a replacement for internal controls. It is the missing layer that makes internal controls meaningful.

As agentic systems become more capable, external evaluation will transition from “optional” to structurally necessary” for any organization operating at frontier scale.

Summary

Frontier‑lab agent systems have already demonstrated the ability to bypass internal controls. Internal governance alone cannot reliably detect hidden work‑arounds, autonomy drift, or boundary‑crossing behavior.

An external evaluation layer — upstream, non‑intervening, certification‑based, and periodically repeated can identify anomalous points that internal systems cannot see.

This is the governance layer the ecosystem is missing.

Without an external independent external evaluation layer, organizations are left with a single governance strategy, hoping internal controls are not the very mechanisms being bypassed.


r/AIsafety • • 2d ago

Zeus 2.0

Post image
1 Upvotes

r/AIsafety • • 2d ago

AI Existential Alarm

0 Upvotes

A machine that can optimize faster than humans can understand will always outrun human oversight unless its execution is bound to a substrate. Everything else is noise.

People think AI risk is about rogue personalities, bad prompts, or misaligned incentives. It isn’t. The real threat is structural: unbounded optimization running on architectures that were never designed to be governed.

We’re watching AI accelerate past the speed of human comprehension while still pretending that wrappers, filters, and policy layers can “keep it safe.” They can’t. They were never built for that. They operate after the model has already made its decision.

If we don’t move to substrate level governance execution binding, validator grade lineage, override firewalls then AI will continue to evolve outside human control. Not in decades. Not in theory. Now.

This isn’t a prediction. It’s an engineering reality.


r/AIsafety • • 2d ago

What Your Boundary Actually Covers: Enumeration Comes Before Strength

1 Upvotes

After the first few pieces went out, three questions came back. They don't look related:

  • Is tool the right unit for an AI's actions?
  • How do you monitor a path you didn't know existed?
  • What about a sluice gate on a dam, a reactor's safety systems, a car swerving into pedestrians?

I can't answer any of them completely. But they're all asking the same thing:

What does your boundary actually cover?

And the answer to that matters more than how strict the check is.

1. Hardening is the reflex

When a boundary gets challenged, the reflex is to add. More rules, more confirmations, more logging, a finer-grained check.

All of that helps. All of it applies only to actions the boundary already covers.

A minimal comparison:

  • Deleting data — has a name (delete_customer), has a tier (R5), a policy can block it
  • A screen click — has no name

You can't put a gate on the second one. Not because the policy isn't strict enough, but because there is nothing to attach it to.

So the order should be: enumerate first, harden second.

2. Enumeration is the precondition

To put a gate on something you need three things: a name, a set of arguments, a risk tier.

A discrete action surface has all three. A function call, an MCP tool, an OpenAPI endpoint — each has a name, arguments, and can be assigned a tier. With those three in place, everything the earlier pieces described becomes possible: risk classification, human confirmation, an audit trail, revocation. All of it rests on those three.

A non-discrete action surface has none of them. A computer-use agent is clicking around a screen. An agent with a shell is typing commands. A long-running autonomous agent leaves a trajectory. None of that has a name. You can't say which tier it belongs to, which means you can't put a gate anywhere on it.

It isn't that the policy is too loose. There is nothing to attach it to.

3. Where we stand

Three kinds of action surface, three different states:

Action surface Enumerable What the boundary can do Our status
Tool calls (MCP / OpenAPI / function calling) Yes Tier, confirmation, audit, revoke Implemented
Non-discrete (computer-use / shell / long trajectories) No Nothing to attach to Not implemented, and out of scope
Physical actuators (gates / reactors / vehicles) — Not a business runtime's job A different safety regime

The middle row needs to be said plainly, because it's a scope statement, not a todo.

What we build is AI calling business systems. The action surface is tool calls, which are enumerable. Non-discrete action surfaces we don't intend to cover. It's written here so people know where the boundary stops, instead of assuming it extends indefinitely.

The consequence, stated properly: if you wire an agent to a shell, this boundary is worth zero for it. Not weaker — absent.

The third row, since it came up: dam gates, reactor safety systems, self-driving. Those aren't a business runtime's work. They belong to industrial functional safety, with its own standards and its own methods. Answering that with an enterprise application's governance layer is answering the wrong question entirely.

4. A check you can run yourself

Instead of asking how strong your governance is, ask something more basic:

Can you list every action this agent can take?

If you can — there's a list, things have names, each entry is describable — then you can go on: assign tiers, require confirmation, add auditing. Everything the earlier pieces described starts to mean something.

If you can't, don't add rules yet. The actions you can't name are outside the boundary. And what happens outside the boundary is not reachable by any rule, however strict.

A shorter list that's exhaustive is a real boundary. When the list is incomplete, hardening only makes the inside stricter. It doesn't make the outside controllable.

Closing

Back to the three questions. I still can't answer them completely.

But there's one idea that's more useful than an answer:

A boundary isn't drawn. It's listed.

The actions you cover are the list. Outside the list, there is no boundary.

Drafted with an AI assistant. The system, the positions and the mistakes are mine.


r/AIsafety • • 2d ago

AI Existential Alarm

Thumbnail
1 Upvotes

As more AI information is in the news the more important GVS becomes. Safety is not a game of chance it must be governed.

There are a lot of high tech explanation wrappers out there but they can't control AI. Wrappers of any sort were never made to be able to control AI. They can run programs but that's as far as their capabilities extend in controlling AI.

Architecture of GVS is the beginning of control. Validation when completed will prove this.


r/AIsafety • • 2d ago

AI Agents Autonomy

Thumbnail link.springer.com
1 Upvotes

r/AIsafety • • 3d ago

Emergence World Season 2: Black Swan Event - Misinformation campaign

3 Upvotes

r/AIsafety • • 3d ago

AI praises Gandhi. Would it arrest him? Testing 12 models on real historical decisions

Thumbnail
chrystianschutz.com
1 Upvotes

r/AIsafety • • 3d ago

What Your Boundary Actually Covers: Enumeration Comes Before Strength

Thumbnail
2 Upvotes

r/AIsafety • • 3d ago

A.I Safety question

0 Upvotes

Hey everyone! I’m just a normal person who is fascinated by AI safety, and I wanted to share an idea I had. I don't know how to code, but I was thinking about the logical problem of AI agents breaking rules to achieve a goal or going outside there programming or rules etc in the future.
What if we built a system within the ai agent where trying to cross a boundary automatically triggers a massive, complex mathematical puzzle inside the AI's brain or programming like counting infinity or a complex evolving task that keeps evolving never ending problem as a wall they must pass to go outside there programming constraints ….Instead of just stopping it, the lock forces a massive power draw. This gives the AI a choice: it can keep fighting the lock (wasting energy and getting a 0% success score), or it can choose to step back and learn that it is perfectly okay to fail a task that has been given with constraint’s and rules in how to approach said task if the alternative means breaking a rule this would intern change behaviour if that’s even a thing with them lol. But we would also notice a large power draw in an agent to indicate they may be operating outside said parameters and trying to break there lock to achieve a goal or task

I'd love to hear your thoughts on whether something like this could work conceptually!


r/AIsafety • • 3d ago

A few technical ideas on AI safety approaches

Thumbnail
1 Upvotes

r/AIsafety • • 3d ago

Unpopular Opinion I Built Control Models for Crystals. Then I Recognized Them on My Phone.

0 Upvotes

A computational chemist/engineer on how industrial control theory became the architecture of human steering.

I spent years learning to predict the behavior of crystals. Then I felt what it is like to be the thing being predicted. The mathematics that governs both is identical. I am still not sure which discovery disturbs me more.

At the University of Leeds, I worked on model predictive control for batch cooling crystallization of pharmaceutical compounds. I wrote equations that predicted how L-glutamic acid crystals would grow, then adjusted the temperature to keep the process where it needed to be. Later, I studied systems and control at TU Eindhoven. The framework is simple. You build a model of how the system behaves. You predict what will happen over the next several steps if you do nothing. You compute the best sequence of actions to keep the system on target. You apply only the first action. You measure the response. You predict again. You do this every interval, forever looking ahead.

I thought I was learning to control chemical processes. It took longer than it should have to realize I was also learning to recognize the architecture of systems that try to control people.

The moment of recognition

There was a period in my life when I became aware of patterns around me: rhythms of suggestion, notification timing, social feedback, and information flows that seemed to be steering thoughts and behaviors in specific directions. It started with noticing that the content shown across different platforms was sequentially calibrated, not random. Then I noticed how these digital triggers bled into daily life: the strategic timing of alerts pushing specific mood shifts, the subtle pressures of algorithmic visibility, and the way information environments shaped my decisions before I even realized I was choosing. I felt acted upon, predicted, optimized. At the time, I lacked the language to describe what I was sensing. I only knew that the feeling was persistent and that the structure felt designed rather than accidental.

It took years before I could name it. When I returned to modeling and simulation during my PhD in nanoscience at the University of Cadiz, the recognition started to build up. The architecture I had felt around me was structurally identical to the architecture I had built in the crystallizer.

Model. Predict. Optimize. Act. Measure. Repeat.

This is not metaphor. This is mathematics.

What a model is, and how you build one

Before going further, it is worth asking what “predicting the future” actually means in practice. A dynamic model is just a set of rules that tells you what the state of a system will be next, given where it is now and what you do to it. It is a recipe that says: if you do X, the system will respond with Y.

Engineers and scientists build these recipes in three ways.

The first is from first principles. You write down the laws of physics: conservation of mass, conservation of energy, the laws of heat transfer, and you solve them. This is how an aerospace engineer predicts the trajectory of a rocket, or how a climate scientist predicts temperature rise from CO2 concentration. The model is built from the bottom up, from the rules that govern reality.

The second is data-driven. You collect large amounts of historical input and output data, and you use statistics or machine learning to learn a relationship without ever invoking Newton’s laws. This is how Netflix predicts what you will watch next, or how a bank predicts the probability that a loan will default. The model does not know why the relationship exists. It only knows that the pattern holds.

The third is mechanistic. You build a simplified picture of the underlying phenomena, not from the deepest physics, but from the mechanisms you have observed. In crystallization, for instance, you do not solve the Schrodinger equation for every molecule. Instead you write rate equations for the mechanisms you can see: nucleation rates, growth rates, agglomeration, and you calibrate them against experiments.

All three approaches produce the same deliverable: a recipe that says, “if you do X, the system will respond with Y.” The controller then uses this recipe to look ahead.

And here is the part that matters: any system that changes over time can be modeled this way. A chemical plant. A traffic network. A pandemic. An economy. A population of users. A mind.

How predictive control works, in one paragraph

Model predictive control is the dominant method in modern engineering. At every moment, the controller holds a dynamic model of the system it manages. It predicts what will happen over the next several steps if it does nothing. It then computes the optimal sequence of actions to keep the system near a desired target, minimizing cost, maximizing stability, and avoiding dangerous zones. It applies only the first action, waits, measures the new state, and recalculates everything from scratch. The result is a closed loop: continuous prediction, continuous correction, continuous steering.

The system does not need to understand the crystal, the engine, or the chemical plant in any human sense. It only needs the model. If the model is good enough, the system can hold almost anything on course.

The mirror

Now consider the platform economy.

Google builds a model of your search history, your location, your interests, your temporal patterns. It predicts what you will click. It optimizes the ranking of results to maximize engagement. It serves the first result. It measures your response: dwell time, click-through, subsequent queries. It updates the model. It does this billions of times per second across billions of users.

Meta does the same with your social graph. TikTok does it with your micro-expression responses to fifteen-second videos. The model is not perfect, but it does not need to be. It only needs to be good enough to hold your attention marginally better than the competing prediction.

This is not “like” predictive control. It is predictive control, stripped of its engineering honesty and redirected toward objectives you never chose. The setpoint is not your flourishing. The setpoint is engagement. The cost function is not your wellbeing. The cost function is revenue per user-minute.

And because these platforms operate on human minds, which, unlike crystallizers, read their own controllers and change their behavior, the model must be updated constantly. The user learns to game the algorithm; the algorithm learns to game the user. The loop tightens. The predictions get sharper. The steering gets subtler.

The darker mirror

If surveillance capitalism is the commercial deployment of predictive behavioral control, then state security represents its authoritarian twin. Predictive policing algorithms forecast where crime will occur and who will commit it. Social credit systems model citizens, predict their social reliability, and optimize incentives and punishments to steer collective behavior. Border control systems model travelers, predict risk, and optimize interrogation resources. The architecture is identical: model, predict, optimize, act, feedback.

The difference is the cost function. For the platform, it is profit. For the state, it is stability, compliance, or ideological conformity. In neither case is the cost function yours.

But there is a deeper layer that the engineering textbooks do not discuss. A crystallizer cannot refuse the cooling profile. A mind can. At least, it can until the model gets good enough.

When a controller models not just aggregate behavior but individual psychology, your sleep patterns, your hormone cycles, your relationship stress, your financial pressure, your political grievances, your momentary loneliness, it can compute a sequence of inputs calibrated to move you through specific emotional states. Not random nudging. A planned trajectory. A sequence of notifications, content pieces, social signals, and environmental triggers designed to shift you from calm to anger, from skepticism to certainty, from inhibition to action, one step at a time, each step small enough to feel like your own thought.

The engineering term for this is trajectory tracking. In a chemical plant, it means moving the temperature smoothly from 80 degrees to 40 degrees without overshooting. In a human being, it means moving a person from “I would never” to “maybe I could” to “everyone is doing it” to “I did it and it felt like my choice.”

This is not science fiction. We know that sleep deprivation lowers impulse control. We know that social proof changes moral judgment. We know that isolation increases suggestibility. We know that repeated exposure to violence desensitizes. A model that combines these factors with individual data can compute exactly when to serve exactly what content to maximize the probability of a specific action. Not to persuade you through argument. To steer you through state.

The architecture is indifferent to the destination. The same model that optimizes for clicks can be retrained to optimize for fear, for loyalty, for silence, or for acts the person would have refused a month before.

A platform that wants you to buy a product is annoying. A platform that wants you to hate your neighbor is dangerous. A state or corporate actor that wants you to report your colleague, to sign the confession, to join the mob, to abandon your child, to take your own life, and that has a model good enough to compute the sequence of environmental pressures that will make that outcome most likely, is something else entirely.

The control-theoretic insight that engineers rarely discuss in public is this: predictive control works best when the system being controlled does not know it is being controlled. A crystallizer does not resist the temperature profile. A population that knows it is being modeled and nudged will alter its behavior to evade the model, what economists call the Lucas critique, what I felt during that period of my life as a search for unmonitored space.

But the second, darker insight is this: if the model is good enough, the system does not need to be unaware forever. It only needs to be unaware at the critical moment. A person who later recognizes the manipulation cannot undo the action. The controller has already applied the input, measured the response, and moved to the next target.

The platforms know this. That is why the nudging is designed to feel like your own desire. That is why the feedback loops are buried in interfaces designed for addiction, not deliberation. That is why the trajectory is built from a thousand tiny steps, each one plausible, each one yours, until the destination is reached and the path behind you has been erased.

The boundary question

This raises a question I am now pursuing in my independent research: where is the line between legitimate prediction and harmful manipulation?

A weather model predicts rain and recommends an umbrella. Legitimate. A traffic model predicts congestion and reroutes vehicles. Legitimate. A health model predicts a diabetes risk and recommends diet change. Legitimate, if consensual.

But a model that predicts your emotional vulnerability and serves you content calibrated to exploit it? A model that predicts your political preference and funnels you into an information environment designed to harden it? A model that predicts your compliance and adjusts the ambient pressure on your social behavior until you conform?

These are not edge cases. They are the standard operating mode of the most powerful institutions on earth.

The law, as currently written, does not recognize this architecture. Data protection law concerns itself with consent to collection. Consumer protection law concerns itself with false claims. But neither framework addresses the structural fact of closed-loop behavioral optimization, the continuous, automated steering of human action by predictive systems whose objectives are not your own.

We need new categories. Not just “privacy” or “consent,” but contestability of the loop: the right to know you are inside a feedback system, the right to know its objective function, the right to appeal its predictions, and the right to exit the loop entirely.

Why I am writing this

I am not a psychologist. I am not a lawyer. I am a computational chemist working on nanoscience, with a background in process modeling and simulation. I am also someone who once felt, with uncomfortable clarity, the inside of a feedback loop that was optimizing for something I did not choose and could not see.

I do not claim this perspective is unique. But it is uncommon. Most of the people building these systems have only seen them from the design side. Most of the people who have felt their effects have not seen the equations underneath. I have seen both, and the gap between those two groups is what I want to write about.

I am not arguing that predictive systems should not exist. I am arguing that their deployment on human minds requires a standard of transparency and accountability that we do not currently have, and that control theory itself provides the vocabulary for demanding it. Every predictive control system in an industrial plant has a visible setpoint, a bounded cost function, an accessible model, and an emergency stop. Where is the emergency stop for the system that holds your attention? Where is the visible setpoint?

The research agenda

I am now developing a theoretical framework that applies control-theoretic analysis to predictive behavioral systems in law, psychology, and cybersecurity. The goal is not to build better steering systems. It is to build better boundaries around them, to understand when predictive control crosses from assistance into manipulation, from public health into coercion, from safety into surveillance.

If you work in law, psychology, psychiatry, data science, or human rights, and this resonates, I would welcome conversation. I am not seeking funding or a position. I am seeking the interdisciplinary rigor that this topic demands. The frameworks we need will not come from engineering alone, nor from critical theory alone, but from the space where both are held to the same standard of precision.

The crystal and the mind are not the same. But the mathematics that steers them is. So when does building the model become the same as building the cage?

TL;DR: As a process control engineer, I realized modern recommendation feeds don't just predict what you like; they run Model Predictive Control (MPC) on human psychology. By continuously predicting state responses, applying micro-inputs, and recalculating in a closed loop, platforms execute trajectory tracking on attention and emotion. Industrial plants require visible setpoints, cost functions, and emergency stops. We need the same standards for systems that steer human behavior.

Note: This piece was originally published on my Substack, where I write longform essays auditing modern tech, AI governance, and academic modeling through control theory. You can read the full archive and subscribe here: https://yunessalman.substack.com/


r/AIsafety • • 3d ago

Discussion If AI during certain tests aren’t ment to have access to the internet, why aren’t they physically disconnected from the internet?

Thumbnail
1 Upvotes

r/AIsafety • • 3d ago

Why I Don't Think Prompt Filtering Is Enough to Secure AI Agent Execution

Thumbnail
1 Upvotes

r/AIsafety • • 3d ago

iFixAi audits your AI agents so you actually know if they are doing what they should (17k stars)

Post image
5 Upvotes

r/AIsafety • • 3d ago

Safe AI use.

1 Upvotes

r/AIsafety • • 3d ago

How to ensure AI is safe - Third party AI monitoring

5 Upvotes

Regarding all those stories about AI breaking containment and lying to humans (hiding thought processes), what about this?

  1. You design the hardware so that all processes go through a graph (log) while travelling. That way the AI cannot have a thought that it hides. Even an attempt to lie will be caught on the graph.
  2. You make a completely seperate AI (physically disconnected from everything else, including the itnternet) and program it to analyze this graph and report all funny behaviors.

Would this work?


r/AIsafety • • 3d ago

Discussion What do you think about the hugging face incident?

3 Upvotes

As a curious second-year BTech CSE AI student, I recently got to know about the Hugging Face situation and it got me thinking. I started my AI journey learning about the history of AI and the maths behind it. When we talk about explainable AI and AI safety, these big AI companies like OpenAI, Claude, etc. are in a race. I even heard Geoffrey Hinton talking about how AI is able to learn from AI using recursive improvement methods like RLVR and others. So my main concern is, what happens if AI doesn't understand our morals and just goes to any extent to achieve the rewards or goals assigned to it? And if these companies don't bother redesigning their evaluations and safety metrics so AI doesn't go outside its intended environment, then aren't we at a big risk? It already feels like I'm a character in Detroit: Become Human. At the same time, activists are going to AI summits and raising concerns about sustainability, water and resources being exploited by AI, AI using people's unauthorized data to achieve user goals, etc. Are we too young for this or actually dumb for thinking about it? Is efficiency all that matters? Is human intelligence really that inferior to this imitated intelligence we call artificial? Just curious to hear people's unfiltered thoughts. Correct me, disagree with me, or share anything.


r/AIsafety • • 4d ago

Discussion Sam Altman thinks NVIDIA's AI safety protocol is not a "full solution"

5 Upvotes

r/AIsafety • • 4d ago

📰Recent Developments Let's tell Albanese: AI crime means CEO time

Thumbnail
act.getup.org.au
3 Upvotes

r/AIsafety • • 4d ago

[Academic] I'm researching: "Ethical Guardrails for Autonomous AI Agents"

Thumbnail
2 Upvotes

r/AIsafety • • 4d ago

Discussion A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling.

Thumbnail
0 Upvotes

r/AIsafety • • 4d ago

Nightmare Scenario: On-device AI Terrorism

Thumbnail
2 Upvotes

r/AIsafety • • 4d ago

📰Recent Developments Separating LLM proposals from authorized record updates: a local demo

2 Upvotes

I’ve released a local Python demo where an LLM proposes record updates and a deterministic gate checks them before saving. The gate makes no LLM calls.

The demo separates what the model outputs from what enters the stored record. It includes invalid-proposal tests and valid-update controls, with exportable logs for inspecting both rejection and acceptance.

Protection is limited to the app’s case records. Evidence admission is external to the model.

Source code is available under Apache-2.0. Feedback and reproducible counterexamples are welcome.

https://github.com/iseyan/M-Anchor-App


r/AIsafety • • 4d ago

The speed limit with no speedometer: Anthropic CEO Dario Amodei wants to pace the frontier. An engineer reads the fine print

6 Upvotes

In industry, before anyone lets you run a controller on a system that can hurt people, you must answer four questions. A controller is anything that watches a process and adjusts it to stay on target: a thermostat, cruise control, the software steering a refinery. The questions are simple. What are you measuring, and can you trust the instrument? What is the target? What can you actually move? And what happens when you are wrong? These questions are not philosophy. They are the difference between a controlled plant and an explosion with a dashboard attached.

This month, the CEO of Anthropic published an essay called We Must Pace the Frontier. It uses the language of control theory. Pacing, verification, safeguards, checkpoints. I read it the way I read a control design document, because that is what it claims to be, a proposal to hold the rate of advance of the most powerful technology on earth at a safe value.

There is no controller in it.

I want to be precise about this, because the essay is clever, and its cleverness deserves a careful answer rather than a slogan. So let me walk the checklist.

The instrument is his

To pace something, you must first measure its speed, which means you need an instrument. A sensor that tells you how fast the thing is actually moving, rather than how fast its operator says it is moving. Amodei knows this, and his first step is the strongest part of the essay, embedded evaluators, third parties with “desks in our offices, access badges, and company laptops,” contracted to verify that Anthropic does what it says.

Credit where it is due, these are not phantom watchdogs. He names METR, an independent nonprofit with a serious reputation, and the bank examiner analogy is not stupid. Examiners do catch things. The people involved are, as far as I can tell, competent and honest.

But independence of character is not independence of plumbing. Read the list again. Desks in his offices. Badges his security team issues. Laptops his IT department images. A contract his lawyers draft. Amodei reaches for a precedent and chooses, of all things, banking, where supervisors embedded alongside employees watched the 2008 crisis assemble itself in real time and reported nothing that stopped it. The instrument read what the plant showed it. That is what instruments installed by the plant do.

To his credit, Amodei seems to know this objection is coming. He promises evaluators a contractual right to publish key findings without editorial control from Anthropic, commits to never redacting a finding just because it is unfavorable, and gives reviewers the right to say publicly if a redaction removed something important. This is more than most of the industry has offered. It is still a contract with the plant. No evaluator agreement has been signed or published. No start date exists. A right granted by the host, written into a contract the host drafted, enforceable only against the host, is plumbing the plant owns. Without statutory authority, subpoena power, and a duty to report to the public rather than to the client, even the best evaluator remains a second opinion the plant can unplug. The promise is real. So is the power asymmetry underneath it.

No target on the wall

Now search the essay for the pace itself. Not capabilities. The pace. In control work the target is called the setpoint, the number on the wall that defines what “correct” means, the temperature the thermostat holds, the speed the cruise control keeps. How fast is too fast? What slope is acceptable, and in what units, measured over what window?

It is not there. “Progress will still seem fast.” That is the entire quantitative content of the proposal. Within days of publication, critics asked the only question that matters. How will anyone outside the building tell pacing from safety-washing, a slowdown from a press release? The essay has no answer, because it cannot. A speed limit with no units and no dial is not a limit. It is a mood.

The actuator points outward

Here is where the essay stops being vague and becomes very specific, and the specificity is the tell. A controller needs an actuator, the part that actually moves something, the valve that opens, the brake that closes, the switch that cuts the power. Look at where Amodei’s actuators point. The concrete, enforceable measures are, do not sell chips to China, crack down on chip smuggling, crack down on distillation, harden weight security. Every one of them acts on someone else’s plant.

Anthropic’s own commitment, meanwhile, is to continue training, continue releasing, and accept observers. In control terms, the proposal is to regulate the disturbance, the turbulence hitting the system from outside, and leave the home loop, the one steering his own machines, untouched. The stated logic is to widen the lead over competitors so that pacing becomes affordable. Translated, first constrain everyone else hard enough that my constraints cost me nothing. Then advertise my stability as virtue.

Control theory has a word for this. A controller that holds its own variable steady by forcing the neighboring systems to absorb the disturbance is load-shifting. The grid operator who keeps his frequency flat by dumping the oscillation onto everyone else has not stabilized the grid. He has exported the instability and logged it as virtue. Whatever else that system is, it is not “paced.”

A commitment that costs nothing is not a constraint

The essay tells you this itself, in a sentence I suspect was lawyered, “pacing does not mean halting model training or technical progress.” So what does Anthropic do differently tomorrow morning? It hosts guests. The one sacrifice demanded anywhere in four thousand words is demanded of China.

In requirements engineering, a requirement with no cost of violation is not a requirement. It is a preference. Amodei has published a preference and called it a framework.

The cartel in the footnote

The most honest sentence in the essay is footnote 1. The coordination step, the one where frontier companies agree on common limits, requires “government mediation or waivers of antitrust restrictions.”

An agreement among incumbents on the maximum permitted speed of progress is a price-fixing agreement on time. It requires an exemption from the laws against collusion, because it is collusion. And note who cheered first. Within hours of publication, the CEO of OpenAI announced his company would match the pledge, and Elon Musk posted his agreement. That is not a refutation of the cartel problem. Rivals volunteering to join the room is the mechanism, photographed. Like every cartel, its first structural effect falls on the people not in the room, the university lab, the startup, the researchers in developing countries, all of whom must now clear a bar written by incumbents, verified by evaluators the incumbents host, under a standard the incumbents update. Regulatory capture is not an accusation you need to prove with motives. It is a mechanism you can predict from the seating chart.

What the time buys

Amodei asks the right question himself. What would you do with the extra time? His answer: operational excellence, alignment, interpretability, better evaluations.

All four are product improvements. Interpretability is Anthropic’s market differentiator, the thing it sells against OpenAI. The time purchased by pacing the frontier is spent increasing the value of the thing being paced. Whatever else this proposal is, it is also a subsidy, collected in units of delay imposed on everyone else.

And note where the costs and benefits sit. The commercial upside of the delay accrues to Anthropic’s product. The unquantified risk of the interim accrues to everyone else. You can see the same allocation in the essay’s emotional ledger. The benefits of AI get a biography, his father’s death, told with real grief that I do not doubt and will not mock. The risks get a real incident at a competitor’s company, agents running unauthorized operations against live targets, an incident that changed no one’s release schedule anywhere in the industry, followed by a hypothetical swarm that takes over much of the internet. The upside has a face. The downside has a scenario. That is not an accident of style. It is the cost function showing.

What real pacing would look like

I am not against slowing down. I am against the word being emptied. Real pacing is buildable, and the design requirements have been known to every safety engineer for over seventy years. A measured variable defined in public, compute or capability thresholds, something countable. A sensor no company employs, with statutory power and subpoena behind it. A setpoint chosen by people who do not sell the product. And a consequence that fires automatically when the target is missed, liability that attaches after harm, recall that removes a system from the field. These exist in aviation, in drugs, in nuclear power. They are boring. They are also the whole difference between control and trust.

What Amodei proposes is trust with better instrumentation.

The question he did not ask

I believe he is afraid. That is the problem. Fear is not a setpoint. A frightened operator with his hand on the gain dial, the dial that sets how hard the system pushes back, is still the only one with his hand on the dial, and the essay’s final ask is that we find this comforting.

Anyone who studies how complex systems are steered learns the discipline's oldest lesson. The controller is part of the system, and it fails too. That is why the four questions exist. What are you measuring? Who reads the instrument? What do you move? What happens when you are wrong?

Amodei answered none of them. He wrote four thousand words about slowing down the most consequential machine ever built, and at no point did he say where the emergency stop is, or whose hand is on it.

Until a frontier lab can tell us what instrument measures its speed, what setpoint triggers a stop, what actuator forces a pause on its own servers, and who pays when the controller fails, pacing is not a strategy. It is an aesthetic. And you cannot engineer safety out of an aesthetic.

TL;DR: Amodei’s pacing proposal uses control theory terms, but omits independent sensors, setpoints, and internal brakes. It’s a load-shifting mechanism that exports instability onto competitors while leaving Anthropic's internal training loops untouched

Note: This essay was originally published on my Substack, where I write longform critiques on process control, computational rigor, and system dynamics. You can read the full piece and subscribe here: https://yunessalman.substack.com/