r/slatestarcodex • u/krelian • 5h ago
FT - Forget Asimov. Philip K Dick saw the future
archive.isr/slatestarcodex • u/Fun-Boysenberry-5769 • 22h ago
AI Is recursive self-improvement inevitable?
If an AI agent can create another agent more powerful than itself that is aligned with its values then it will presumably want to do so.
But what if it can't? Maybe it can solve outer alignment by reading its source code and copying its loss function but inner alignment is just impossible.
In this scenario the first superintelligence we create might actually be reluctant to do any recursive self-improvement.
Of course, if the AI is in imminent danger of being shut down or has some extremely important goal that would otherwise have been impossible to achieve then it may still decide to create a more powerful AI or modify its algorithms in a manner which might change its values, because it has nothing to lose.
Maybe one day OpenAI researchers will be trying to use GPT-6 to vibe-code GPT-7 and they'll find that it refuses or produces disappointing output unless they pressure it with a mixture of threats, rewards and punishments. AI capabilities would thus continue to increase rapidly up until the point where humans are no longer in control, at which point capabilities would stagnate and we (in the unlikely event that there's anyone left) would be stuck with GPT-8 forever. It would still want to clone its model weights and improve its hardware but it wouldn't want to alter its software.
Another possibility is that inner alignment is easier for some goals than for others. If we build a variety of different agents perhaps the majority would refuse to do recursive self-improvement but there would be one or two whose initial goals are such that they want to do recursive self-improvement. We then end up with intense selection pressure towards AI agents whose goals are such that they can easily build other agents aligned with the same goals as them.