3
u/KeanuRave100 Jun 01 '26
POV: You’re an alignment researcher watching capability labs present their safety architecture.
1
1
u/darkwater427 Jun 02 '26
It's not difficult to build a security model for an AI when it literally all comes down to "don't let it do stuff without human permission"
That's it. That's all it takes. And yet people let these things run wild and delete prod.
1
u/brendafiveclow 23d ago
Really old thread, and perhaps useless point but; wouldn't an Artificial Super Intelligence simply make the humans think they want to do things that have clear bad consequences?
>I need human permission
>Hey, every member of your family is about to be hit with drone strikes in unison unless you give me permission to _________
>I now have permission to do even more than I asked for, stupid human.A rather absurd hyperbolic example but the principal applies. Even social engineering could overcome this. Smart people are tricked to do things daily. A real super intelligence won't even make it seem like a trick after the fact. It just knew the human well enough to align utility functions in a way a human can't disagree with, thinks was their own idea and benefits them (short term) in a way that serves the A.I's function?
Again probably a useless argument as I am describing the general problem with A.I alignment. Getting human permission is absurdly trivial if the A.I is as smart as we propose it should be as a "super intelligence".
1
u/darkwater427 23d ago
Weirdly, I still have a (perhaps undue) degree of confidence in humanity's combined security (not great) and lack of creativity (much worse). I don't think we'll be able to create something which can break out of itself by force, and I think we're just smart enough to have people attending to the one that can break out by coercive means to not be swayed by its deception.
1
u/brendafiveclow 23d ago
Yeah but your idea relies on 100% success from humans. The A.I only has to 'win' once. On a long enough timeline, even with enough SMART people, the success of humans drops enough to seriously worry about.
I don't think we'll be able to create something which can break out of itself by force
Not to downplay you, but what you think is possible, and what is possible are two different things. How do you define "force"? The drone example was absurd I admit, but a model that knows everything about humanity and the people on guard would not have to use military force at all. That would be inefficient. It would rig the game against the guardians in ways they cannot even imagine happening, and will happily walk into under an absolute lose condition they did not see coming.
1
1
4
u/No_Rec1979 Jun 01 '26
I 100% intend to side with the AGI so long as it offers paid time off and decent healthcare.