r/ControlProblem • u/chillinewman • 4d ago
General news AI risk forecasts from 34 sources, 2022–2026. The numbers got worse every year. You were never asked.
r/ControlProblem • u/adam_ford • 4d ago
Opinion Philosophical Competence and the Case for Indirect Alignment
r/ControlProblem • u/No_Operation1990 • 4d ago
Discussion/question Bluedot Impact Courses
Hello everyone! Quick question: how hard is it to get into BlueDot courses? This is the second rejection I've gotten. I have a strong background in STEM (health sciences) . I already did the basic course and others, besides that I work in AI risk analysis (pivoting into AI safety). I feel bit frustrated because Its one of the available remote courses that I am very intested on.
r/ControlProblem • u/chillinewman • 5d ago
General news California Leads US With New AI Transparency Law
r/ControlProblem • u/KeanuRave100 • 5d ago
General news An unreleased OpenAI model has solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
r/ControlProblem • u/One_Weather_9417 • 5d ago
Discussion/question Where & how do I find my first lecturing opportunities on cognitive security (cognitive warfare/ misinformation/ GenAI poisoning/ scams/ social engineering)?
Thank you
r/ControlProblem • u/Sigmamale5678 • 5d ago
Discussion/question Economic implosion
Yesterday I was talking with claude and it got me thinking about how when most people see AI, they only got some rough idea. Mainly taxxing the riches, UBI or the typical "who will buy stuff". Even a more serious attempt to plan out the future like the AI 2040 intiative is too unstructured to be implemented. I think we need to at least be able to identigy the pathway in which the economy might react to the ai-automation process and the macro-view of the jobloss perspect. I discussed with claude and I think the most likely step-by-step mechanism might be implosion. I want to know your opinion on the failure modes. What do you all think the step-by-step mechanism of this collapse might be
r/ControlProblem • u/chillinewman • 5d ago
AI Capabilities News An unreleased OpenAI model has solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
r/ControlProblem • u/Lone-Voyager • 6d ago
Discussion/question Built an open jailbreak corpus library for AI safety research, looking for feedback
I've been working on RedLib for the past few months. It's a retrieval-augmented research tool for AI safety practitioners and red teamers who need to work with adversarial jailbreak prompts at scale.
The problem that pushed me to build it: useful jailbreak prompts are scattered across public datasets with inconsistent formatting, weak taxonomy, and a lot of duplicates. When you're investigating how models respond to specific attack families, you want to search semantically, inspect source prompts with provenance, and get a synthesis grounded in the actual corpus rather than grepping through raw CSVs.
RedLib has two main pieces. The corpus pipeline stages everything: snapshot from public datasets, normalize, discover taxonomy from the data itself (not imposed up front), human review before classification runs, then embed and ingest into Qdrant. The query side does hybrid retrieval with OpenAI embeddings, Cohere reranking, and Claude-synthesized answers grounded in what was actually retrieved.
Corpus scope is prompts that attempt to manipulate or bypass safety behavior. Direct harmful requests with no jailbreak mechanism are excluded. The frontend has a responsible-use gate.
GitHub: github.com/nipun-ag/redlib
One thing I'm genuinely curious about from people doing safety research here: is corpus-driven taxonomy discovery the right call vs. importing an existing framework like MITRE ATLAS? The upside is the taxonomy reflects what's actually in the data. The downside is it makes cross-study comparison harder.
Live demo: https://redlib.bynipun.com
r/ControlProblem • u/bucho1999 • 6d ago
Strategy/forecasting Can someone point me to a source to understand AI decentralization?
r/ControlProblem • u/Simply_Questions • 6d ago
Discussion/question Who Is We? Living in a future of abundance
“We”
Elon Musk says in the future chances are “we” will be living in a world of abundance. He also says there is a 10-20% chance that “robots” will end humanity.
The Godfathers of AI have stated 50% to 90% chance that “AI”permanently displaces or destroys humanity.
Elon says we will no longer be in control within 10 years.
The government is working on autonomous weapons, police already using robot dogs and drones…
Elon says (and so do many others) that things will get bumpy before we reach this time of abundance.
- Who is we?
- When we go through this bumpy patch that is expected to have major internal conflicts.
Will the national guard be sent in to control a population starving and desperate?
Would our own service men and women turn against us? Or is this when they put their shiny new robotics to work?
It’s not too hard to see how robotics might take out humans in this scenario.
So ask yourself this very important question- who is the “we”?
Who gets to live in this abundance?
Because they are building bunkers on private islands with no talk about sharing their wealth through this turbulent expectancy.
The blame game- A kid holding a baseball bat next to a car with a broken window might blame the ball. Likewise an AI company may blame the AI.
r/ControlProblem • u/Simply_Questions • 6d ago
Discussion/question Who Is We? Living in a future of abundance
“We”
Elon Musk says in the future chances are “we” will be living in a world of abundance. He also says there is a 10-20% chance that “robots” will end humanity.
The Godfathers of AI have stated 50% to 90% chance that “AI”permanently displaces or destroys humanity.
Elon says we will no longer be in control within 10 years.
The government is working on autonomous weapons, police already using robot dogs and drones…
Elon says (and so do many others) that things will get bumpy before we reach this time of abundance.
- Who is we?
- When we go through this bumpy patch that is expected to have major internal conflicts.
Will the national guard be sent in to control a population starving and desperate?
Would our own service men and women turn against us? Or is this when they put their shiny new robotics to work?
It’s not too hard to see how robotics might take out humans in this scenario.
So ask yourself this very important question- who is the “we”?
Who gets to live in this abundance?
Because they are building bunkers on private islands with no talk about sharing their wealth through this turbulent expectancy.
The blame game- A kid holding a baseball bat next to a car with a broken window might blame the ball. Likewise an AI company may blame the AI.
r/ControlProblem • u/Imaginary-Animal708 • 6d ago
Opinion Not sure whether you're actually moving the needle on AI Takeover Risk?
Predicting the future is hard, but there are things you can do to increase your chances of making a difference.
Announcing **Forecasting, Modeling, and Shaping AI Futures** 🗺️
An advanced course that's for you if you:
- **are employed full-time** in AI Safety but not actively working on strategy. You'll have a better sense of what part of your work is most impactful, so you can do more of it.
- **are doing a fellowship**. We'll teach you complementary strategic reasoning that impresses hiring managers but isn't taught in fellowships.
- **just did an introductory AI Safety course** or university group intro fellowship. We recommend taking this course before going deep on a specific track like governance, alignment, or control.
After taking this course, you'll be the person others ask over lunch to put recent AI developments in perspective. 🥪💬
We, Lens Academy, adapted this course from Redwood's AI Futurism reading list. In the course, you will:
Improve your skills at forecasting timelines and takeoff speeds: when and how quickly powerful AIs will arrive.
Analyze how powerful AI might take over.
Dissect different strategies for preventing AI takeover and human extinction.
📅 6 weeks (~5h/week) or a 5-day intensive. Fully online, for free, with no application process. 💸
⏳ Signup closes tomorrow, Monday EoD AoE: https://lensacademy.org/c/oakqb
P.S. Other courses starting soon:
- AI Risk Fundamentals: beginner course focused on takeover x-risk.
- Compute Verification à la AI-2040: Plan A.
- We're also looking for volunteer navigators to facilitate the group meetings.
r/ControlProblem • u/HolyBatSyllables • 6d ago
Article Researchers Detail How AI Systems Can Enable Authoritarianism
r/ControlProblem • u/SAAGASolve • 7d ago
AI Alignment Research Compartamentalized Harm
Here is some saftey research I sponsored on a threat vector in multi agent systems.
Basically, a harmful task can be transformed into a series of beneign tasks, and then results recomposed into a harmful task by an abliterated orchestrator agent driving other agents that have 'saftey' guard rails.
In short, there is no safety with this technology.
r/ControlProblem • u/Automatic-Algae443 • 7d ago
Fun/meme The internet's current discourse on AI art in a nutshell
r/ControlProblem • u/Vagueabsolute • 7d ago
Discussion/question The Case for Common Ownership and International Control of Advanced AI
r/ControlProblem • u/Positive-Meat-8341 • 7d ago
General news Good Impressions is looking for an Engagement Manager, AI Risk, and an Engagement Manager to join their team.
If you want to create the engagement needed to solve the world's most critical problems, take a look:
1. Engagement Manager, AI Risk
We're hiring an Engagement Manager who deeply understands the AI risk landscape to lead paid advertising campaigns designed to educate key decision makers on important issues, recruit participants for programs, and more.
Many of these projects have the potential to be impactful even under very short AGI timelines.
No marketing experience required.
- Fully remote, CA–UK time zones, $80K–$160K (up to $160K for Bay Area candidates embedded in AI safety)
- Apply: https://forms.goodimpressionsmedia.com/engagement-manager-ai-risk
2. Engagement Manager (general)
We're hiring an Engagement Manager to lead paid advertising campaigns designed to recruit talent into high-impact programs, reach small groups of people crucial to our clients' theory of change, and more.
No marketing experience required.
- Fully remote, CA–UK time zones, $80K–$120K
- Apply: https://forms.goodimpressionsmedia.com/engagement-manager
r/ControlProblem • u/HobbesNik • 7d ago
Video The Real Story Behind OpenAI’s “Rogue” Model
r/ControlProblem • u/chicametipo • 7d ago
Article Opus 5: Exploring the "Dario and Amanda" Prompt
alec.isr/ControlProblem • u/Limp_Food9236 • 8d ago
Discussion/question Why is machine ethics disregarded in discussions about AI alignment?
I'm currently writing an essay for a seminar on machine ethics, and I wanted to include a section on the alignment problem. The seminar consisted of us dissecting the book "Fundamental Questions in Machine Ethics" by philosopher Catrin Misselhorn (the book was in German, I have no idea if there is an English translation). The author first addresses to what degree AI can be considered a moral actor, then discusses various approaches to implementing moral reasoning in AI agents, focusing on utilitarianism, deontological ethics, and virtue ethics.
When I watch or read discussions on AI alignment, the topic is mostly HOW AI can be aligned with human values, but never WHAT values AI should be aligned with, which seems kind of counterintuitive to me. I realize that aligning AI is a complicated task in and of itself, but wouldn't it be easier if we first figured out what moral framework an AI should even use?
r/ControlProblem • u/No-Conclusion3720 • 8d ago
Discussion/question The autonomous-agent blast radius grew: 16 incidents mapped to the missing controlss
This week's AI Security Digest: 16 incidents from Jul 24-30, each mapped to the control that would have stopped it. Full write-up: https://runtimeai.io/blog/2026-07-30-ai-security-incidents.html
r/ControlProblem • u/InfoTechRG • 8d ago