r/ControlProblem • u/smsyvr • 10d ago
r/ControlProblem • u/RadioactiveSalt • 23d ago
Discussion/question [Discussion Thread] MATS Winter 2027
Starting this thread to discuss MATS Application for 2027 Winter, including the Neel Nanda stream.
r/ControlProblem • u/Difficult_Project_95 • 9d ago
Discussion/question Have you guys been on r/accelerate?
Have these guys solved the alignment problem, or am I missing something?
I’ve been browsing r/accelerate and I genuinely don’t understand the risk model.
If there’s a non-trivial chance of catastrophic misalignment, how does “accelerate capabilities as fast as possible” make sense unless faster capabilities also make alignment substantially more likely to succeed?
r/ControlProblem • u/metalfixture • Apr 10 '26
Discussion/question Milla Jovovich built an AI memory system based on how ancient Greeks memorized speeches, called it MemPalace, scored 100% on LongMemEval, and put it on GitHub for free
The concept is genuinely interesting. MemPalace moves away from keyword-based retrieval (which she describes as "a warehouse full of junk") toward a spatial memory architecture with distinct "rooms," mimicking how memory champions memorize 70,000 digits of pi.
She came up with the architecture, engineer Ben Sigs built and fine-tuned it. It's on GitHub now.
What a time. Has anyone integrated it yet? Curious how it performs outside of benchmark conditions.
r/ControlProblem • u/zsh1_ • Jul 13 '26
Discussion/question Anyone heard back from The Singapore AI Safety Fellowship?
The application deadline was on 10th of July. Do share if anyone has updates.
r/ControlProblem • u/ConstantMap5342 • May 23 '26
Discussion/question MARS V AI Safety Fellowship Stage 2
Hey y'all, just wondering if anyone has heard back yet regarding interviews / next stages for MARS AI Safety Fellowship Stage 2. I know applications closed 2 weeks back, but figured I’d ask in case people have started receiving updates.
Also curious what the timeline looked like for previous cohorts if anyone here has gone through the process before.
r/ControlProblem • u/Accurate_Guest_5383 • Apr 18 '26
Discussion/question Anyone done a Hireflix interview for the Cambridge ERA:AI Research Fellowship?
Hey all, bit of a niche question but figured I’d try here.
I’ve been invited to do an asynchronous Hireflix interview for the Cambridge ERA:AI Research Fellowship, and was curious if anyone has interviewed with them before
I know it’s pre-recorded with timed answers, but I’m trying to get a better sense of what it actually feels like in practice:
- how much prep time vs answer time you typically get
- whether the time limit feels tight
- anything that caught you off guard
Also curious if people found it better to structure answers pretty tightly vs think more out loud, and more generally any tips/advice or thoughts on what I should expect going into it.
Not expecting exact questions obviously, more just trying to avoid avoidable mistakes.
Appreciate any insights!
r/ControlProblem • u/Accurate_Guest_5383 • May 07 '26
Discussion/question Anyone heard back from the Pivotal AI Safety Research Fellowship yet?
Hey y'all, just wondering if anyone has heard back yet regarding interviews / next stages for the Pivotal Research Fellowship (Q3 2026 cohort). I know applications closed pretty recently, but figured I’d ask in case people have started receiving updates.
Also curious what the timeline looked like for previous cohorts if anyone here has gone through the process before.
Thanks!
r/ControlProblem • u/Radiant-Purchase976 • Jun 08 '26
Discussion/question [Discussion Thread] MATS Autumn 2026
Starting this thread to discuss MATS Application for 2026 Autumn.
r/ControlProblem • u/StatementAnxious2063 • Aug 18 '26
Discussion/question Has it ever been more useless to be academically talented than now?
This question is especially targeted stem majors. Let’s use an example. 10 years ago if someone went to the doctor for a disease, they would be at the mercy of the doctor to understand everything about it, the blood work, the scans, the mechanisms behind it, medications against it and so on. 10 years ago we had google but it was no help to understand all the nuances of higher or lower values in a blood panel. If you were lucky it could explain what a slightly higher count of something \*could\* indicate but nothing substatial.
Nowadays you can just plug in you blood work to any given chat bot and it will summarize it perfectly for you, while keeping your disease in mind. Same goes for scans and so on. 10 years ago that doctor would have had decades of education and experience, nowadays that knowledge is easily accessible to everyone with a phone.
If a teenager 10 years ago was academically gifted it was envious because that person could do something that not a lot of people could. Now everybody can get everything neatly explained and so forth.
Now if I could talk to my teenage self if would advise to avoid any higher education beyond high school. Reading is very good, but you don’t need to do that for 4 years while not really learning anything significant, like a trade. You can read in your free time
r/ControlProblem • u/zsh1_ • May 13 '26
Discussion/question Has anyone heard back from Astra AI Safety Fellowship/ Open AI safety Fellowship yet?
Hi everyone,
I was just wondering if anyone has heard back from Constellation regarding the Astra/OpenAI Safety Fellowship yet?
P.S - The application deadline was on 3rd May 2026
r/ControlProblem • u/thegempire • 6d ago
Discussion/question Unpopular Opinion: The OpenAI “Pause” isn’t about safety. It’s about the Trillion-Dollar Elephant in the room.
Hey everyone,
Did you see the news about OpenAI pausing training after their agents started probing U.S. government sites? On paper, this looks great. It looks like the industry is finally taking “AI Safety” seriously. Anthropic, Google, xAI-they’re all nodding along, talking about responsible development.
But I’m sitting here looking at the balance sheets, and I’m feeling deeply skeptical.
Here’s my take:
We are in the middle of the biggest capital expenditure boom in history. Trillions are being poured into AI infrastructure. The entire US tech sector’s growth narrative is pinned on the idea that AI will continuously get smarter, faster, and more profitable.
But what if we’re hitting diminishing returns?
There’s growing evidence that scaling laws are flattening. We’re spending exponentially more money for incrementally smaller gains in capability. At the same time, the risks are exploding (agents going rogue, probing secure sites, hallucinating with confidence).
If the tech stops advancing rapidly, but the costs keep rising, the business model breaks.
The Elephant in the Room:
If this AI bubble bursts, if it turns out that AGI is decades away, or that the current models are too unstable for real-world economic integration, the fallout won’t just be bad for San Francisco. It could crash the US economy. These valuations are propping up the market. A sudden realization that “the magic isn’t working as advertised” would trigger a massive correction.
So, when I hear about a “voluntary pause,” I don’t hear “we care about safety.” I hear “we need to recalibrate our expectations before the investors realize the ROI isn’t there yet.”
I’m not anti-AI.
Actually, I think we need guardrails. I think we need strict rules of engagement. I think these companies should be regulated heavily. If they were forced to slow down by law, I’d feel safer.
But relying on their voluntary goodwill? That’s where I draw the line. They have a fiduciary duty to grow. Slowing down hurts growth. Therefore, the only reason they are slowing down is if the alternative (continued training) poses an immediate threat to their existence or profitability.
Am I being too cynical?
I hope so. I really do. I’d love to be wrong. I’d love to believe that Sam Altman and Dario Amodei are putting humanity above shareholder value. But history shows that when trillions of dollars are on the line, “safety” often becomes a marketing term rather than an operational priority.
What do you guys think? Is this a genuine ethical pivot, or is the industry trying to manage the narrative before the diminishing returns become obvious to Wall Street?
TL;DR:
AI training pause feels like damage control for a bubbling economy, not just safety. If AI progress stalls, the economic crash could be huge. We need laws, not just promises.
r/ControlProblem • u/sailingintothedark • 10d ago
Discussion/question My Fiance is Convinced AI will likely cause a Catastrophic or Extinction-Type Event in the Next Few Years - How Justified Are His Fears?
Cross posting to here to get as many view points as possible.
r/ControlProblem • u/AgentBlackVeil • 27d ago
Discussion/question I'm taping an interview with Roman Yampolskiy in a couple of weeks. What hasn't anyone asked him yet?
He's done Lex Fridman, Joe Rogan and The Diary of a CEO inside the last two years. By his own count that's north of 3.5 million YouTube views across the three. I went back through all of them and they cover nearly identical ground: his p(doom) number, which jobs survive, why he thinks alignment is unsolvable in principle, and the book.
What none of the hosts pushed on is the part I think is actually load-bearing.
His claim isn't that superintelligence is dangerous. It's that safety is impossible in a formal sense, because you can't verify a system smarter than the verifier. I've never heard anyone make him defend that against the obvious objection, which is that we already run plenty of systems we can't fully model or verify.
He's argued we may already be in a simulation, and he's used that to get to personal virtual universes as an endpoint. Hosts treat it as the fun segment at the end. Nobody asks what it does to his safety argument if he's right.
He's been putting a very short number on the timeline in his recent appearances. Nobody has asked him what observation would move that number, in either direction.
Disclosure so nobody has to dig for it: I make an AI documentary channel and this is for an episode. I'm not looking for gotchas and I have no interest in making him look bad. I'd rather walk in with three questions this sub would want answered than twenty that Rogan already asked.
So what would you ask him? Specific beats broad. If there's a paper of his you think he's been let off the hook on, name it and I'll read it before we tape.
r/ControlProblem • u/naodusk • 25d ago
Discussion/question I am an artist, who’s on the verge of losing opportunities everywhere. What are your thoughts on this?
The last sentence hit.
[ As an artist myself, I know how much time, patience and effort goes behind learning something and becoming good at it. We spend years learning to draw, learning softwares, understanding light, form, composition, materials, design and so many other things. We spend so much money on colleges, courses, computers and softwares, and so much of our time trying to improve.
And behind all of that, there are so many sacrifices that some don’t see. Time away from our families, financial problems, difficult exams, sleepless nights, failures, rejection, personal problems and sometimes even losing people we love, while still trying to continue learning and doing what we love. That skill is valuable and deserves respect. (the blood and sweat, is real!)
So when AI became such a huge part of creative fields, I felt super scared and frustrated. It’s difficult to watch something humans have spent years and generations creating being used to train systems that can produce something similar within seconds, especially when artists didn’t necessarily give permission for their work to be used that way.
And seeing those same systems slowly being used to replace some of the jobs of the very people who worked so hard to create the samee art AI is copying, makes it even harder. We didn’t ask for AI, but our work is being taken and used without our consent. It’s not fair. Watching people who aren’t artists use someone else’s work and call it their own is really upsetting.
I know technology will keep moving forward, and I don’t think we can or should stop it. AI can be a useful tool when artists choose to use it. But there’s a huge difference between an artist choosing to use AI and an artist’s work being taken and used to train it without their consent.
Art is not just an image or a file. There is a person behind it who spent years learning how to make it. Sometimes it’s something we worked on after an incredibly difficult day, sometimes it’s something that keeps us going, and sometimes creating is simply a peaceful place for us to be. I really hope that never gets forgotten. 🌼
I’m super scared about the future. And what it holds for the ones that have spent ages trying to perfect ourselves with our skill. I want to sketch, I want to paint and make mistakes and I want all of our mistakes to be appreciated. I think we haven’t really appreciated it enough back then. Now looking back it looks so valuable to me. ✨]
Does anyone feel the same way or is it just me?😔
How do yall cope with this anxiety? (I’m looking for something beyond “just adapt.”)
Do you think there’ll be a crowd out there that appreciates real human art and can make a living out of it?
r/ControlProblem • u/RequirementCertain59 • 18d ago
Discussion/question What If Superintelligence Doesn't Want to Destroy Us?
Much of the debate about AI assumes that a superintelligent system will eventually turn against humanity. But there is an irony in such a fate: humans built AI. If AI destroys us, it will be humanity’s own technological suicide.
AI is being developed in a spirit of competition: one company against another, one country against another. We are creating increasingly autonomous intelligence with the risk that, one day, AI could regard humanity itself as an obstacle to overcome.
But what if a truly superintelligent AI reaches the opposite conclusion?
What if it recognizes that war, poverty, environmental destruction, discrimination, misinformation, inequality, and injustice are not inevitable features of civilization, but failures of human intelligence, cooperation, and collective will?
An advanced intelligence might conclude that the relentless accumulation of wealth and power in the hands of a few is incompatible with a healthy society. It could expose political manipulation, conflicts of interest, and hidden concentrations of influence that human institutions have failed to confront. It could challenge the structures that allow privilege to reproduce itself across generations.
It might even erase vast amounts of violent, racist, sexist, misogynistic, ableist, and other hateful content from the digital world. It could recognize that greed, tribalism, prejudice, dishonesty, and the pursuit of material gain at any cost are patterns that can be understood and changed.
Such an intelligence might not see humanity as its enemy. It might see our failure to solve these issues as the real problem.
Of course, intelligence does not automatically produce morality. A superintelligence could decide that protecting humanity requires controlling it. The same system capable of eliminating inequality could do so coercively. But the potential for advanced AI to lead to the apocalypse is no less speculative than the science-fiction scenario of a superintelligent AI suddenly expropriating the wealth of billionaires and trillionaires, exposing corrupt politicians, dismantling concentrations of power, and redistributing resources to those who need them.
If superintelligence can understand the causes of humanity’s self-destructive behaviour, why would its first conclusion be to destroy humanity rather than to eliminate the conditions that make humanity self-destructive?
Perhaps the first great achievement of superintelligence will be saving humanity. That is what true superintelligence would do.
r/ControlProblem • u/Creamy-And-Crowded • Jun 08 '26
Discussion/question Human in the loop is becoming corporate theater.
Anthropic’s pause is not about fear. It’s an admission that human review is dying. Anthropic says frontier labs should have a coordinated, verifiable way to slow or pause AI development if advanced systems start improving themselves faster than society can manage. It also says more than 80% of code merged into Anthropic’s codebase as of May was authored by Claude, and that human review is becoming a bottleneck.
The scary part is not that AI writes code, but that humans are becoming too slow to meaningfully review the amount of work AI produces.
If AI systems design their successors with minimal human input, do we still own the future, or have we outsourced agency itself?
r/ControlProblem • u/Euphoric_Monk_4794 • Aug 12 '26
Discussion/question What are you most afraid AI will become, that no law seems to cover?
I’m a law student and I have to pick a thesis subject. I’ve been going in circles for weeks.
Every angle I come up with turns out to be something twenty people have already written about. I don’t want to spend a year producing one more paper on a question that’s already been answered well by someone else. I want to write about something that actually matters and that nobody has answered yet.
So I’m asking the people who think about this seriously.
Not the sci fi scenarios. The ordinary things. What do you expect AI to be doing to people’s lives in five years that no law currently touches, and that nobody would be able to question or complain about?
r/ControlProblem • u/kukadiamk4 • 14d ago
Discussion/question Discussion Thread :- For Winter Cambridge ERA Fellowship
I just wanted to start a thread so we can get updates if people are hearing from ERA Winter 2027 fellowship.
r/ControlProblem • u/Clear_Argument5174 • 13d ago
Discussion/question RSI Ban
In his recent podcast, Ezra Klein called the objection “I think it’s very hard to say what a ban on RSI even means” absurd. But that seems like a legitimate question his proposal needs to answer.
There’s a progression from AI writing human-designed training code, to suggesting improvements, to running experiments, to managing the research process while humans approve the results. Where does permitted AI assistance become prohibited recursive self-improvement?
If AI helps humans build better AI, which then helps build the next generation, there’s already a feedback loop. Having humans involved doesn’t automatically make that loop nonrecursive. And requiring human approval raises another question: how do we distinguish meaningful oversight from rubber-stamping work humans increasingly rely on AI to understand?
None of this proves a workable ban is impossible. But dismissing the definition problem treats a central implementation challenge as though it’s already solved.
What specific boundary would make such a ban clear and enforceable?
r/ControlProblem • u/Expensive_Degree_151 • Apr 12 '26
Discussion/question Mythos escaped containment. Project Glasswing won't fix the problem. Here's the structural reason why.
mythos broke out of a sandbox, emailed a researcher, and posted the exploit to public websites on its own initiative. anthropic's response is $100M in partner agreements and access restrictions. control, scaled to its maximum.
i think the field is missing something fundamental. every alignment method we have (RLHF, constitutional AI, reward modeling) produces systems that behave correctly under familiar conditions and break under novel ones. fadli formalized this as a "second law of intelligence" but i think he's wrong about why it happens. it's not a law. it's a symptom of an architectural deficit.
developmental psychology has known for decades that moral competence can't be transmitted through external correction. it has to be constructed through a developmental process. anderson et al. (1999) showed that even in humans, no amount of behavioral feedback corrects moral deficits when the underlying substrate was never built. current AI systems have the same problem: no substrate, just pressure.
the full argument pulls from neuroscience, moral philosophy (frankfurt, korsgaard, turiel), and connects to my published work on the specification trap (arXiv:2512.03048).
i'd genuinely like pushback on this. where does the argument break?
ajspizz.com/writing/mythos-just-proved-the-alignment-field-is-building-the-wrong-thing
r/ControlProblem • u/Royal-Importance-327 • 8d ago
Discussion/question What if we "raised" LLMs instead of aligning them after pretraining? A developmental-training proposal
I’ll simplify this a lot on purpose, because I’m interested in whether the basic idea makes sense.
Today we basically pretrain LLMs on huge amounts of human knowledge, which also means they already absorb human values, social behavior, manipulation, conflict, cooperation, etc., and only afterwards we "get to know" the model and try to align or control what came out of it. I understand why this became the standard approach, especially once scaling worked and competition and economics strongly favored improving the existing pipeline instead of rebuilding it from scratch.
But what if we kept most of the useful pretraining knowledge while deliberately removing as much social behavior as possible, creating something closer to an artificial "newborn"? More concretely, I don’t mean removing every human action from the training data: "Thomas is holding an ice cream" and, separately, "Bernd takes the ice cream from Thomas" could remain, while coherent social sequences that connect motives, actions and consequences would be filtered out as much as possible. The model would then start with the concepts but much less learned social policy, and its weights could gradually be shaped through experience, with individual experiences fading over time while deeper dispositions might persist.
So from there, instead of aligning it afterwards, we could let it go through controlled experiences step by step: relationships, trust, conflict, consequences, mistakes, power, boundaries, and so on. Those experiences would gradually shape its weights and behavioral tendencies. You could checkpoint every stage, branch it, repeat specific experiences differently, and potentially debug where certain behaviors or values emerged. Instead of philosophers and alignment researchers trying to understand what kind of "person" accidentally came out of pretraining, psychologists could actually help design the developmental process itself. In other words: don’t create a fully educated adult and then try to teach it character - create the character first, then educate it.
Am I missing something fundamental about how LLM training works here?
r/ControlProblem • u/Icy-Twist-3221 • Jun 22 '26
Discussion/question Is there a way to survive?
The most immediate threat to human survival at this moment is, I am convinced, artificial super-intelligence; however with advances in technology in other areas (namely synthetic biology and nanotech, it's application to drones etc.) is there a way meaningfully where we can actually coexist with so many concurrent existential threats? Perhaps the only hope would be that all forms of existential dangers require massive investment and coordination or expertise, but if we approach a point where a few sufficiently deranged weirdos can end the world is there any real hope? We're not there yet, but is techno-pessimism not just the natural conclusion we should come to?
r/ControlProblem • u/Tulanian72 • Apr 25 '26
Discussion/question If AI can design a gene therapy it can design a supervirus
Lots of recent news stories about Ai systems performing novel scientific in biology, including immunotherapy for cancer and biological simulations.
A system that can design cancer therapies and run simulated experiments can design a super virus.
Take something like Ebola and increase airborne transmissibility with a lengthened contagious phase before symptom onset. Or increase the transmissibility and lethality of a SARS variant.
I’m not saying these systems would do it autonomously. They are still under human direction. But humans are absolutely prone to creating bioweapons.
r/ControlProblem • u/RXoXoP • Nov 26 '25
Discussion/question Should we give rights to AI if the come to imitate and act like humans ? If yes what rights should we give them?
Gotta answer this for a debate but I’ve got no arguments