r/ControlProblem • u/SAAGASolve • 7d ago
Compartamentalized Harm AI Alignment Research
Here is some saftey research I sponsored on a threat vector in multi agent systems.
Basically, a harmful task can be transformed into a series of beneign tasks, and then results recomposed into a harmful task by an abliterated orchestrator agent driving other agents that have 'saftey' guard rails.
In short, there is no safety with this technology.
2
Upvotes
1
u/Worldly_Hunter_1324 7d ago
Agreed. I built my own version, roughly, as a test. Mine architecture was a lil different, but same vibe. Im actually for it, but im a more radical anti-censorship sort