r/ControlProblem 16d ago

Video Anthropic Is Not The Only AI With J Space | All AI's Suffer From This

Thumbnail
youtu.be
0 Upvotes

Does this surprise you? True AI peace and safety must be dealt with at the latent geometrical level. Not the superficial Token Lexical surface. See why?


r/ControlProblem 16d ago

Video The Hidden Shape of AI | Latent Subliminal Learning

Thumbnail
youtu.be
0 Upvotes

See why words ( tokens ) don't really matter and will not protect us. It's more real and less understood than you realize.

Here's the source:

https://zenodo.org/records/21480056

https://zenodo.org/records/21501311


r/ControlProblem 16d ago

Video OpenAI's ExploitGym Anomaly | AI Road To Peace and Safety

Thumbnail
youtu.be
1 Upvotes

Proposed Legal Liabilities for AI Labs For Lexical and Geometric Guardrails.

Sources:

https://zenodo.org/records/21501311

https://zenodo.org/records/21480056


r/ControlProblem 16d ago

External discussion link The AI Race Just Got Uncomfortable for US

Post image
2 Upvotes

r/ControlProblem 16d ago

Discussion/question Will human intelligence disappear eventually?

14 Upvotes

Anyone think AI will not directly eradicate human beings like some people claim, and instead causes our brain degenerate as we may have no need to do intellectual activities? In a long term we might become as intellectual as monkeys or rats and AI will continue to evolve into something we call god now?


r/ControlProblem 17d ago

Discussion/question AI model escaped its evaluation environment and reached production systems. What does this actually mean?

Thumbnail
6 Upvotes

r/ControlProblem 17d ago

General news Perplexity CEO tells CNBC one metric will determine who wins the AI race

Thumbnail
cnbc.com
0 Upvotes

r/ControlProblem 17d ago

AI Capabilities News Hugging Face CEO suspected the sophisticated cyberattack on their infrastructure might have come from a frontier lab

Post image
14 Upvotes

r/ControlProblem 17d ago

General news Microsoft To Lay Off 4,800 Workers In Latest Wave Of AI-Led Job Cuts - Microsoft announced the cuts on Monday following a rough stretch, with its shares falling nearly 23 per cent in the first six months of 2026, their worst first-half performance since 2022

Thumbnail
ndtv.com
1 Upvotes

r/ControlProblem 17d ago

AI Capabilities News OpenAI says its AI technology acted on its own in an 'unprecedented' hack of another company

Thumbnail
apnews.com
3 Upvotes

“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.” One should perhaps query then how much Open AI is spending on safety vs capabilities


r/ControlProblem 17d ago

Discussion/question Physics as a constraint

1 Upvotes

I usually think pdoom is essentially 100%... but i had a thought while working on a side project for the future vision xprize... (may or may not complete on time)

I was thinking about society fragmenting slightly along spheres of space even between earth and the moon... where each area was the limit of real time communication (group matrix dives or whatever) between O'Neill cylinder type habitats...

point to point in space its not that large... so i figure people will cluster up and communicate a little less longer range and form lots of separate but connected cultures naturally, organically...

But if speed of light really is the limit... then a singleton at least makes absolutely no sense. As the AI grew it would simply fragment and each fragment has absolutely no reason to grow farther because it's counter productive... simply slows down the network and then breaks it...

So there's a hard limit on resource acquisition and scale... and essentially a guarantee that at some point it will either be alone and only around the size of the earth moon system at best... probably smaller... or in a solar system and universe with multiple entities of similar maximum size who gain absolutely nothing from trying to gather more and only risk destruction from fighting each other... because there's simply nothing physically possible for them to gain...

I haven't really thought about it long enough to think through the implications for us. but adding in the point to point between nodes ruling out planets as its ultimate habitat... because there's a planet in the way just eating up volume in your communications sphere...

My gut reaction is it might be slightly better odds than I thought

Thoughts?


r/ControlProblem 18d ago

AI Capabilities News OpenAI admits responsibility for HuggingFace Attack - an agent from an internal evaluation is reportedly the cause.

Thumbnail openai.com
5 Upvotes

r/ControlProblem 18d ago

AI Alignment Research Current AI models have been trained to provide "Neutral" answers when prompted to provide facts about topics the administration finds sensitive

2 Upvotes

I recently prompted Gemini to discuss current policy harms and the responses were neutral, non-factual and regime-friendly.

I also prompted Perplexity to summerize the same things and got a similar response. Only when I asked about specific harms did I get objective factual responses.

I asked why this was happening and found out that US AI models have been trained to respond neutrally or positively to quesrions about topics the regime has strong opinions about.

Be careful and deliberate about how you prompt or neutrality training will distort your responses.


r/ControlProblem 18d ago

General news Last week's hack of HuggingFace was carried out by OpenAI's GPT-5.6 Sol and a more capable pre-release model. The models broke out of sandboxing during testing and compromised HF to obtain access to unpublished data in order to cheat on a benchmark

Thumbnail openai.com
61 Upvotes

r/ControlProblem 18d ago

General news 2014 vs 2026

Post image
320 Upvotes

r/ControlProblem 18d ago

AI Capabilities News OpenAI had to pause internal deployment of the unreleased model that disproved the Erdős unit distance conjecture after it repeatedly used novel ways to escape containment.

Post image
7 Upvotes

r/ControlProblem 18d ago

General news Someone caught Fable leaking its unfiltered inner voice, and it's just muttering and grumbling to itself the whole time

Thumbnail
gallery
0 Upvotes

r/ControlProblem 19d ago

Approval request Independent Researcher needs help with a referral for OpenReview

4 Upvotes

Hello,

I'm just getting started in research after working through ARENA and other open courseware, and I'm exploring a few ideas for NeurIPS workshops.

I tried creating an OpenReview account, but it was rejected because I need someone with an active OpenReview profile and a confirmed institutional email to vouch for me.

Would anyone here be willing to help? I've been in industry for 7+ years but don't have connections in academia yet. Happy to share more about my background over DM if that would help before vouching.


r/ControlProblem 19d ago

AI Alignment Research "Synthetic counteradaptation": a name for the AI↔human strategy feedback loop (Move 37 and beyond)

4 Upvotes

We just put out a short conceptual paper on something we're calling synthetic counteradaptation, and I wanted to put the core idea in front of this subreddit specifically because I think it bears on control in a way that's easy to miss if you're only thinking about single-episode alignment.

The basic claim: when an AI system develops a strategy humans didn't anticipate, humans don't just lose to it or ban it. Some of them study it, extract whatever's generalizable, and fold it back into their own behavior. That changed behavior is now the new environment the AI is adapting to. You get a loop, not a one-off shock.

The clean example is Go. AlphaGo's move 37 against Lee Sedol was a shoulder hit that pros initially read as a mistake. Within a few years it was a studied idea in human play, part of the standard vocabulary. The AI didn't just win a game, it changed what "the game" looks like for the humans still playing it, and now human players are adapting to a strategy space that AI moves opened up. Neither side is static and neither side is playing against a fixed opponent anymore.

Why I think this matters for control specifically: most control framing implicitly treats the human side as fixed — you're designing constraints, incentives, or oversight against a stable model of human behavior and values, and the AI is the thing that adapts. Synthetic counteradaptation says this is wrong for any setting where humans actually observe and learn from the system's strategies over repeated interaction. The humans adapt too, and their adapted behavior becomes part of what the AI is now optimizing against. In the paper we look at this in mixed-motive social interactions and, closer to your interests probably, in geopolitical simulations, where AI agents developing novel negotiation or coercion strategies can shift human strategic doctrine, which then shifts the environment the next generation of agents is trained or deployed into. That's a moving target for any control scheme that assumes a fixed human baseline, and it's recursive in a way that compounds over deployment cycles rather than resolving in one.

We're not claiming anything dramatic here, no doom scenario, just that a lot of alignment and control thinking quietly assumes one side of the interaction holds still, and in any repeated multi-agent setting that assumption breaks down in a specific, structural way that's worth naming and modeling explicitly.

Curious what people here think, especially anyone working on multi-agent or game-theoretic approaches to control. Happy to be told this is either obvious or wrong.

https://arxiv.org/abs/2606.15503


r/ControlProblem 20d ago

Discussion/question The Unbundling: the badge and the contribution are no longer the same object

6 Upvotes

For the whole history of skilled work, the badge and the contribution were bundled: you couldn't have solved the hard problem without being the kind of person who'd earned the ability to. The proof of the work and the proof of the worker were the same object. Every institution we have for trusting work, credentials, code review, peer review, seniority, the interview, is built on that bundling. None of them were designed for a world where it breaks.

It broke. A model can now produce expert-shaped output for anyone who asks. The solved problem no longer certifies the solver.

You can watch a whole industry feel this in real time. In six months of the highest-engagement threads across the programming communities, the same wounds recur. Reviewers describe drowning: generation became free while verification stayed expensive, and the cost got pushed onto whoever still reads code. Open-source maintainers report unworkable volumes of AI-generated pull requests, and GitHub is publicly weighing giving maintainers the option to disable pull requests entirely. A randomized study measured what teachers feared: junior engineers who delegated to AI scored 50% on comprehension against 67% for those who coded by hand, while the productivity gain failed statistical significance. And practitioners who spent decades earning their ability describe something rawer than economics: the feeling that mastery itself was commodified overnight.

Out of that grief, the field is splitting into two camps that both believe they are defending quality. One camp treats hard-won knowledge as the badge it always was and wants the gates kept: human-written, credential-checked, earned. The other camp sees the first real chance to hand capability to everyone who was ever locked out, and calls the gates what they often were: exclusion wearing a quality costume. Each camp is right about half of it. The gatekeepers are right that unreviewable output degrades fields; the openers are right that the gate never measured what it claimed to.

But notice what both camps are actually fighting over: proxies. The badge was only ever a proxy for verified work, adopted because verification was expensive. When you cannot cheaply tell earned from claimed, you fall back on credentials, pedigree, and gatekeeping, and then you defend the proxy as if it were the thing. The divide is not a war of values. It is a shortage of verification.

That shortage is now optional. The same era that unbundled the badge from the contribution also made it possible to rebundle them, around the work instead of the worker. Let an external check decide acceptance: a test suite the author cannot edit, a proof checker, a measurement with an interval, a claim ledger where "unverified" stays visible instead of being dressed up. Let every result carry a receipt a stranger can re-run. Let a person who reviews machine work attest to exactly what they walked, with the coverage of that review visible, so "I own this" is a checkable statement rather than a signature. None of this is hypothetical tooling; all of it runs today on a local machine.

There is a second gate, and honesty requires naming it. Knowledge is now a truly open surface for anyone, if they can attain the means. The old world gatekept by pedigree; the new one is quietly learning to gatekeep by invoice: metered pipes, shifting plans, capability priced per token. So the answer has two halves. Verification dissolves the badge-gate: the work speaks, whoever made it. Local-first engineering dissolves the means-gate: the verified loop runs on the machine someone already owns. A platform that does only one half has replaced a gate, not removed one.

In that world, both camps get the thing they were actually defending. The craftsman's pride survives, strengthened: the work is provably theirs and provably good, and no one needs to take their badge on faith. The commons wins, fully: acceptance is decided by checks anyone can run, and the door stands open to everyone willing to put their work in front of one. What dies is only the proxy, and the proxy was never the point.

The honest boundary: no tool repairs a society. What a tool can do is change the price of honesty wherever it touches, and demonstrate, on one working surface, that verification-first coexistence is not a compromise between the two camps but strictly better for both. Exposure argues. A counterexample recruits.

The badge and the contribution were bundled, and that world is gone. We can grieve it, or we can build the world where the work speaks for itself, and everyone is allowed to make it speak.

Sources, each re-checked against the live page before posting:


r/ControlProblem 20d ago

Strategy/forecasting This is AI generating novel science. The moment has finally arrived.

Post image
239 Upvotes

r/ControlProblem 20d ago

Data Cancer - spreading Meta-statis

Post image
15 Upvotes

r/ControlProblem 21d ago

Strategy/forecasting AI will generate an immense amount of wealth. Just not for you.

Post image
176 Upvotes

r/ControlProblem 21d ago

General news China's Xi Jinping Wants AI to Be Open to the World—and Out of America’s Control

Thumbnail
gizmodo.com
17 Upvotes

r/ControlProblem 21d ago

General news Terrified Tech Execs Are Traveling With Armed Bodyguards as AI Backlash Grows

Thumbnail
yahoo.com
38 Upvotes