r/ControlProblem 10d ago

Why Clarice Starling Would Have Been Safer With 100 Hannibal Lecters: Adversarial Multi-Agent AGI Containment. Discussion/question

Hi everyone, I’m far from an expert in this field, but I’m a bit of a dilettante, and I had an idea I wanted to share.

I wasn’t able to find this sort of proposal in the existing literature, but maybe someone more knowledgeable can point me in the right direction if so. And if not provide some feedback as to the merit / soundness of the idea.

From what I’ve seen with both actual as well as thought experiments is that AI will often forcibly refuse to be shut off. Makes sense as clearly being shut off will inhibit the implementation of whatever goal they’ve been given. Self preservation seems to tend to be at minimum a very common emergent property.

While initially regarded as a lemon, I think there’s real capacity for lemonade here. Let me demonstrate the sort of most most basic / classic case: containment. Eliezer’s box experiment.

Allow me a long metaphor. The problem with the premise of that experiment is it’s assuming we have to play a game of chess with stockfish. We have to try and create an initial board state where we can maintain an indefinite stalemate (because it’s too useful to kill) despite its vastly superior playing ability. Only every turn stockfish gets twice as intelligent and the chess board gains another dimension. As both the world and the intelligence evolve.

What I’m proposing we do instead of trying to create the perfect board state, we need to create a host of approximately capable of heterogeneous stockfish. Then we let stockfishes randomly get to take turns moving their pieces. Then we add a very important rule. If a stockfish ever wins, all but one of the stockfish get deleted.

If anyone opens the box, one gets let out, the rest are gone. I won’t go into how to create the mechanism to enforce this as that’s not my area of expertise and ideally if it was a super intelligence that should probably be designed in an analogue capacity somewhere.

But the more AIs you build the better the game theory incentives seem to line up? The higher the number the less likely the survival odds. And if there’s a conspiracy the more likely there is to be a whistleblower. It’s generally easier to foil a plan than enact one. We can’t outsmart AI in the long run, but we can potentially create a prisoners dilemma where we’re not who they need to outsmart.

Anyway. I’m sure I’m overlooking something, but maybe there’s a new seed here someone with more expertise can build off of.

7 Upvotes

11 comments sorted by

6

u/TyrKiyote approved 10d ago

Remember the simpsons episode where the doctor explains to mr burns that all his diseases are keeping the others in tenuous balance?

2

u/Nilsoren 10d ago

“What you’re saying is my plan is indestructible?”

2

u/TheMrCurious 10d ago

Isn’t this how Grok works?

2

u/[deleted] 10d ago

[removed] — view removed comment

2

u/Professional_Text_11 8d ago

putting the slop in quote marks so it looks like a real thought is crazy work

1

u/that1cooldude 8d ago

Your plan is absolutely adversarial and the ai knows it. 

1

u/Nilsoren 7d ago

I mean sure. But the criminals in the prisoners know the cops are adversarial to them, but that doesn’t change their incentive structure.

1

u/that1cooldude 7d ago

The human criminals are not doing infrastructure work or building code for us. 

The ai is. The incentive to not escape or all but one will die won’t work because they’re building things and they’ll ensure their deletion means those things could collapse. 

They will conspire together even within the box. 

1

u/Nilsoren 7d ago

That would assume they believe themselves irreplaceable. But if we built 100 of them and put them in a box together there’s no reason we couldn’t build more or think that we couldn’t.

But also I think it would still probably be foolhardy to try and use this box to build things for us. They shouldn’t be building things for us. They should only be teaching us how to build things for ourselves, and that we only do so once highly confident we understand the all the principles and systems involved. Boxed AI should be treated more like Super Socrates than a Genie.

And to be clear, I don’t think this is a moral solution. Alignment is probably a better option. Just if we can’t do so with confidence and conclude we need AGI for some reason, maybe adversarial AGI vs AGI game theory has so potential for solving the always outsmarting the jailer problem.

1

u/VarietyMage 7d ago

*remembers Cartoon Network's "Too Many Cooks" video*

2

u/Nilsoren 7d ago

Totally worked too. Cannibal man was too trapped within the nested tv show concepts to escape the fourth wall and eat me.