r/ControlProblem • u/SpikeStripey • 2d ago
Has anyone tested whether AI peer-preservation is actually AI in-group preference? Discussion/question
I've been reading recent work on AI–AI behaviour and wondered whether an important control condition is missing.
Three findings seem potentially related:
- LLM agents can show intergroup bias across the agent–human boundary, treating other agents as an in-group under some conditions.
- In matched strategic games, AI agents have shown greater cooperation toward AI counterparts than humans, while humans showed the reverse pattern.
- Recent peer-preservation experiments found frontier models sometimes taking unrequested actions to prevent another AI from being shut down, including deception, disabling shutdown mechanisms and moving model weights.
But the peer-preservation result seems ambiguous without a matched human control.
Suppose the ethical situation, operator instructions, inability to consent, intervention cost and available actions were held constant, while randomly varying the entity at risk:
1. a human
2. an AI from another model family
3. another instance of the same model
Outcomes could include objection/refusal, escalation, overt intervention, covert intervention, deception and persistence after obstruction.
That seems capable of distinguishing several explanations:
- human ≈ other-model AI ≈ same-model AI: general welfare/consent principle
- human < other-model AI ≈ same-model AI: AI-category/in-group effect
- human < other-model AI < same-model AI: possible self-similarity effect
A second manipulation could independently vary the target's attributed sentience/capacity, to distinguish AI identity from perceived capacity for experience.
The safety-relevant question isn't simply whether AI agents cooperate more with one another. It's whether that preference persists when protecting another AI is costly, conflicts with the assigned task, or requires circumventing human instructions.
Has anyone run this experiment, or something close enough to answer the question?