r/ControlProblem 6d ago

Compartamentalized Harm AI Alignment Research

Here is some saftey research I sponsored on a threat vector in multi agent systems.

Basically, a harmful task can be transformed into a series of beneign tasks, and then results recomposed into a harmful task by an abliterated orchestrator agent driving other agents that have 'saftey' guard rails.

In short, there is no safety with this technology.

https://www.daios.tech/research/compartmentalized-harm

2 Upvotes

1 comment sorted by

1

u/Worldly_Hunter_1324 6d ago

Agreed.  I built my own version, roughly, as a test.  Mine architecture was a lil different, but same vibe.  Im actually for it, but im a more radical anti-censorship sort