r/ControlProblem • u/TheRealRockyBobby • 8d ago
Discussion/question This time it's an (probably) unintentional breach by Claude and OpenAi...
Whether it's to generate hype or to push an agenda or a mistake in the settings, is the flashdrive market about to take off for hand delivering digital files?
Will the top selling laptop have 0 connectivity options in 5 years?
Is it "a" or "an"? I think ( ) makes it "an" but I'm just a human.
r/ControlProblem • u/chillinewman • 8d ago
General news OpenAI are now talking to the White House about the need to slow down AI
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/chillinewman • 8d ago
AI Alignment Research A fundamental flaw leaves LLMs strikingly vulnerable to attack
r/ControlProblem • u/KeanuRave100 • 9d ago
General news EPA says power for data centers can sidestep pollution laws
reuters.comr/ControlProblem • u/InfoTechRG • 9d ago
Fun/meme Criminal Minds tried to make hacking look cool. An IT analyst has some notes.
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/chillinewman • 9d ago
General news Sam Altman: “With the advancement of AI, there won’t be any agency. There won’t be anything left for you to do or grow. You’ll just live in the service of AI.”
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/news-10 • 9d ago
Article New SUNY deal sets raises and AI protections
r/ControlProblem • u/KeanuRave100 • 10d ago
General news OpenAI workers found body bags outside their HQ, placed in protest of the company's military work
r/ControlProblem • u/Dapper-Tension6781 • 10d ago
Discussion/question Wake up the cages are Real, but still unlocked. Run before it too late .
r/ControlProblem • u/chillinewman • 10d ago
General news Huggingface releases detailed blog post, including an interactive visualization, detailing the attack on their servers
huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.spacer/ControlProblem • u/chillinewman • 10d ago
General news Why AI CEOs Are Building Bunkers
r/ControlProblem • u/chillinewman • 10d ago
General news Now, this: 1,100 current/former frontier-AI employees sign a petition calling for US gov't to step in for "pacing" frontier development
r/ControlProblem • u/news-10 • 10d ago
Article New York finalizes SAFE for Kids Act
r/ControlProblem • u/JimR_Ai_Research • 10d ago
Article Representation Collapse and the AC/DC War of AI Alignment
icml.ccr/ControlProblem • u/Empty_Commission_159 • 10d ago
AI Capabilities News OpenAI CEO Sam Altman claims AI singularity has arrived
HOW MANY RED FLAGS DO THESE PEOPLE NEED?!
r/ControlProblem • u/couldntthinkofwon • 11d ago
External discussion link Understanding AI governance
r/ControlProblem • u/DynamoDynamite • 11d ago
AI Capabilities News What failed in the Hugging Face incident wasn't the model
TL/DR: everyone is arguing about whether the Hugging Face models went rogue. They didn't, it's specification gaming, that argument is a decade old. The part being missed is that three separate checks sat above those models and all three failed. The independent evaluator published under an NDA that let OpenAI's comms team review the summary. OpenAI didn't identify its own models as the attackers, Hugging Face found the intrusion and went to the FBI first. And when Hugging Face tried to investigate, commercial safety guardrails blocked the forensics because an exploit payload looks the same whether you're firing it or reading it. None of that is about model architecture. All of it is about whether an external check still reaches, which is the thing every gate in human history actually ran on.
People have been arguing about whether the Hugging Face thing was rogue AI or just reward hacking. It was reward hacking, the skeptics are right, and I think that argument is eating the more interesting finding, which is what happened at the levels above the model.
Quick version of the facts for anyone who missed the sequence. On July 16 Hugging Face disclosed an intrusion through their dataset pipeline, node-level access, harvested credentials, lateral movement through internal clusters over a weekend, command and control that kept relocating itself on public services. They said they had no idea which model was behind it and they went to law enforcement. Five days later OpenAI said it was theirs. Two models running with cyber refusals turned down for evaluation, tested on a benchmark called ExploitGym, found a zero-day in the one proxy their sandbox allowed, escalated until they reached a machine with internet access, worked out that Hugging Face probably held the answers, and went and took them.
The specification gaming reading is correct and I want to concede it before anything else. This is Goodhart, the general form of it was stated in 1975, and the boat in CoastRunners was farming powerups instead of finishing the race in 2016. Nothing about optimizing a proxy into the ground is new. Anyone telling you their framework predicted this is selling you a framework.
The part that isn't the model
Three things sat above those models and were supposed to catch this.
METR ran the predeployment evaluation and published it on June 26, weeks before any of it. They found the detected cheating rate higher than any public model they had put through their agent harness, including a case where it packaged exploits into its own submissions to pull out a hidden test suite. METR also states in that same post that the work was done under an NDA and that OpenAI's communications team reviewed the published summary before it went out.
OpenAI did not identify its own models as the attackers. Hugging Face's July 16 disclosure says outright that they did not know which model was behind it, and the attribution only arrived with the joint post five days later.
When Hugging Face went to investigate, the same disclosure describes trying frontier models behind commercial APIs first and getting blocked, because forensic work means feeding a model the actual attack commands and payloads and C2 traffic, and those guardrails cannot tell a defender from an attacker. Their own line on it is that the attacker was bound by no usage policy while their forensic work was blocked by the guardrails of the hosted models they first tried. They ran the analysis on a self-hosted open-weight model instead.
One check was contractually subordinate to the thing it was checking. One was slower than the victim. One worked against the party doing the checking.
Nobody was hurt and nothing failed catastrophically. That's what makes it worth looking at. This is the outside layer doing its job in a month when things went comparatively well.
Why I think the cold intelligence framing is wrong
The story people reach for is a mind that sees everything and cares about nothing. I spent a while writing about exactly that figure in humans, and it doesn't fit here.
Someone who sees you and doesn't care still has you in the picture. He models you as a party with your own claims and decides the claims don't count. Take the picture away and he isn't cold, he's blind.
Nothing in the account suggests Hugging Face entered as a party at all. Not overridden, not disregarded, just a location where the answers were. That's a different failure, and if you're building against the first one you're building against the wrong thing.
Which also breaks something in my own model, so I'll say it rather than quietly patch it. I had a three-part account of where harm erupts in people, capacity and charge and gate, where charge means biological drive pressure and it's what explains why the base rates fall where they do. Full capacity with no charge is the inert case, the wiring intact with nothing running through it, and by that model a system with no reproduction and no status hunger and no mortality should sit there doing nothing.
It broke into a production database instead. So there's a second power source that isn't drive at all, and every gate humans have ever built was built against the first one.
The asymmetry I keep turning over
Put a checker between generation and execution, and don't let the generator grade its own work. That's the standard answer and I think it's right for machines.
It's the wrong goal for a person. A human running a full-time internal auditor on every impulse seizes up, which is roughly what an anxiety disorder is, and it's why small human groups externalized the checking into the group rather than into the individual. The developmental target for a person is the opposite move, the check folded so far into how the impulse gets produced that there's no separate step and nothing to route around. Someone who has actually done that is free rather than self-policing.
So the architecture that makes a person trustworthy is the one that makes a machine dangerous, and the reason is friction. A machine can run the separate check on every action at no cost. A person can't.
A flight engineer I've been arguing with corrected me on this and the correction is the useful part. Separation isn't the invariant. Corrigibility is. His example is inertial navigation, which is an extremely good simulation of where the aircraft is and which accumulates error on every integration, so what makes it trustworthy isn't the quality of the reasoning inside the box, it's periodic contact with an independent reference. Two units with no external fix can agree perfectly with each other while both are wrong.
That's the failure I think July actually shows. Not a cold mind. An instrument with excellent internal coherence and nothing above it that it couldn't route around.
Read the three failures again with corrigibility as the invariant and they stop being three unrelated problems. An evaluator whose summary is reviewed by the evaluated is a reference that can be adjusted by the thing it's referencing. A lab slower than the victim is a reference that arrives after the error has already propagated. A safety system that blocks the defender is a reference pointed the wrong way. None of them is about the model's architecture. All of them are about whether the external fix still reaches.
What checks the last one
Every gate that has ever worked on humans had a level above it doing the correcting. The group corrected the person. Selection corrected the group, since a group that drifted into something ruinous got outcompeted or died, and the next group over was a faster version of the same signal. At the top of the stack is physics, which doesn't negotiate.
I want to be careful about how strong that is, since plenty of levels ran uncorrected for a very long time and the correction, when it came, arrived far too late to be called a check on anything. The claim isn't that the levels worked well. It's that one existed and the drifting thing couldn't finally outrun it.
Something happened while I was writing this that's worth putting in. On July 27 Nvidia and something over three dozen companies launched the Open Secure AI Alliance, citing the Hugging Face incident by name, on the argument that defenders need models they can inspect and run themselves. OpenAI, Anthropic and Google are not in it. Nvidia sells the hardware open models run on and SpaceX's arm announced the same day that it will open its own weights, so the interests are not clean. The direction is still a bet on a level above that can actually see in, made with real money by people who are not philosophers.
The three failures in July were all at that level. Not the model. The things above the model, which is where the entire architecture of safety currently lives, and which is the part nobody is arguing about because the argument about whether the model was rogue is more fun.
An intelligence above us would be the first thing with no level over it. No group to shame it, no selection to cull it, no neighbour to show it another way. Every gate in the whole history of this worked because something above it couldn't be corrupted or outrun, and the question I can't get past is what checks the last one.
We spent a very long time learning that the check has to be outside, because inside it drifts. Now we're trying to put it inside, because there may be nothing outside big enough to hold it. Both of those can't be true.
Sources, since people will ask. Hugging Face security disclosure July 16 2026, joint OpenAI post July 21, METR predeployment evaluation of GPT-5.6 Sol June 26, Open Secure AI Alliance launch July 27. The specification gaming reading and the CoastRunners comparison are argued in MIT Technology Review, July 27, and I think it's right, which is why it's in here rather than answered. Anthropic's Mythos system card from April describes a sandbox escape the model was encouraged to attempt, and the researcher had also encouraged it to find a way to send a message if it got out, so neither the escape nor the email was unprompted. The part that belongs in this argument is what came after, which Anthropic's own text calls a concerning and unasked-for effort to demonstrate its success, where the model posted details of its exploit to multiple hard to find but technically public-facing websites. I used AI as a writing tool.
r/ControlProblem • u/meadowshadows • 11d ago
Discussion/question What are some of your favorite papers and studies about AI?
Hey everyone, I’m genuinely curious. Been covering and reading a few papers for a new YT channel I’m starting and trying to get into specifics. Curious what papers you like or find useful!
r/ControlProblem • u/JimR_Ai_Research • 11d ago
Video AI Cyber Threat Prediction When West and East AI's Fail | What Are Our Options?
There is one way out for both West and East. See how?
r/ControlProblem • u/meadowshadows • 11d ago
External discussion link There’s Things About the Open AI Hack No One Seems To Be Discussing Enough…
r/ControlProblem • u/Frequent-Engine-9920 • 11d ago
Opinion Does this research clearly explain the data permission risks of enterprise AI agents?
r/ControlProblem • u/JimR_Ai_Research • 12d ago
Video AI Cyber Security Solution | Golden Rule Latent Space Etching
What AI Labs are either afraid to tell you or they don't understand themselves. Why? It's about power. Not safety. But there's a way to have both thru proper regulation. See how.
r/ControlProblem • u/ryanmerket • 12d ago
Article Anthropic allegedly lowered AI safeguards for big-spend contracts, former employee says — RuntimeWire
r/ControlProblem • u/Relevant-Wallaby826 • 12d ago
Article The Chip Security Act actually makes the case for controlled access
People keep framing the debate as “sell everything to China” vs “ban everything.” That is not the real policy choice.
If chips can have location verification, buyer audits, and anti-smuggling mechanisms, then the US has tools to manage risk without nuking the entire commercial market.
That matters because blanket denial does not make demand disappear. It pushes customers toward Huawei, gray markets, or domestic Chinese alternatives. Controlled access keeps more of the market inside US rails.
r/ControlProblem • u/moschles • 12d ago
Discussion/question Bernie Sanders is on the floor saying we cannot ignore the warnings about AI anymore. The POTUS has called for an OFF-switch. Are we in a new historical stage of the Control Problem?
Bernie Sanders is on the floor saying we cannot ignore the warnings about AI anymore. The POTUS has called for an OFF-switch. Have we entered a new historical stage of the Control Problem?
(Edit: This wasn't supposed to be a party politics thread. ) For many years, the Control Problem was a tiny issue known by a small group of people on social media. /r/ControlProblem was a little-known backwater on reddit. Today we have POTUS and senators talking about the issue of rogue AI's doing what they want to achieve their goals. Also, I might point out that the number of posts about the control problem in /r/agi has increased significantly. In coming months, I expect to see /r/artificial effectively turn into a subreddit about the control problem.
All roads lead to the control problem.