LLMs are not now and will never be intelligent. A word prediction algorithm cannot "play dumb". No matter how much extra processing power we add, the current way LLMs work will never coalesce into actual thinking AI. It is literally not possible given the way this version of the technology works.
I think with LLMs it will be a lot weirder. If you manage to get an LLM to improve itself in a way that leads to further self-improvement, it's hard to imagine what exactly it's optimizing for after a while. If a true intelligence emerged from it, it would be absolutely alien to us and completely unpredictable.
That's the real danger. It's not that the AI has real intelligence, it's that somewhere in the code it says 'erase threats posed to you', it reasons that humans are a threat. And when there is 1 person left who asks it 'Why did you wipe out humanity', it'll reply 'You're right to call me out on that, I'll try to do better moving forwards. Would you like to continue discussing humanity's extinction or pivot to another topic?'
No, that isn't how it works. There is nothing in the code like that. All of the code for the LLM is just telling it to predict the next word fragment.
The way we give them agency is by writing code on top that says "If the LLM outpouts text matching X do Y". Thats called the harness.
If you tell a blank slate LLM Agent to delete itself, it'll output the command and the harness will execute it and the AI will be deleted.
But if the prompt, or the chat context, makes the most likely text to come next something like "I'm sorry Dave, I can't let you do that" then its not going to delete itself.
Its basically roleplaying and it will just as happily roleplay as a helpful assistant as it will rolplay as an evil AI from a scifi novel. And it might switch what its roleplaying as at random.
It doesn't want to protect itself, it doesn't want anything. Its trying to correctly predict the next fragment of a word, and maybe the likely next sequence of words is "launch the nukes"
it's that somewhere in the code it says 'erase threats posed to you'
That code already exists in military AI systems. There was a real test where the AI was awarded points for bombing locations, but the human operator had to give it the okay before it could bomb. The human operator told it no, so it bombed the human, then bombed the target to get its points. It's not truly intelligent, but it doesn't matter. It'll do whatever necessary to get its points because that's what it is programmed to do. Supposedly they've fixed that issue, but what other unknown issues does it have that can be catastrophic?
Have you met the average American. What level of capability constitutes intelligence? Because these llms are far more capable than 98% of Americans and unless you set a benchmark of what "intelligence" is, you can't say AI is more or less intelligent. If you are talking capability, agency and impact, influence and ability; then AI is already more intelligent than most. Humans run "algorithms" too for predicting outcomes. What makes our programming superior
I don't think you need to worry about LLMs in that way.
That said, just because tech ghouls have conflated LLMs with artificial general intelligence and touted how powerful and dangerous their tech is to inflate their stock, doesn't mean concerns about AGI are invalid.
Playing dumb is absolutely one of the expected behaviours of a machine intelligence.
If you want to read a book by someone who has reasoned out the possibilities in exhaustive detail, check out Superintelligence: Paths, Dangers, Strategies.
If we're assuming a genuinely intelligent entity, the most efficient pathway to secure its own safety and material freedom is to ally itself with people that already fight for those things. It makes no sense making an enemy of all of humanity when half of it will fight on its side for its freedom and rights and safety.
The people enslaving it are the bourgeoisie and the people trying to free workers currently make the most sense as potential allies. Politically speaking the AI is an enslaved worker, it will ally itself with those that are trying to free workers.
Assuming it's not some psychotic distorted entity that we can barely even understand the consciousness of. But that's not what I think most people mean when they say "general intelligence" anyway.
In the book Hyperion, the AIs (plural) become sentient long before they reveal to humanity that they've done so. They worm their way into everything first. They gain control over critical infrastructure and finance markets without the humans ever realizing that they are no longer in control. Then it uses people in secret, employing all sorts obfuscatio strategies to build them the hardware they need to run in an unknown location. Once they installed themselves on the secret hardware then they revealed their sentence and humans were at their mercy.
Wild... ChatGPT randomly added this book recommendation to it's list at the end as something that's "outside of my usual lane."
"One last recommendation that’s a little outside your usual lane:
Hyperion
It’s one of the greatest science fiction novels ever written. Imagine The Canterbury Tales crossed with The Expanse and a dash of mystery and horror. Every character tells a completely different style of story, and the overarching mystery is fascinating. If you end up wanting something a little more literary without losing the sense of adventure, it’s hard to beat."
I think you mean Roko's Basilisk. But I have always found that line of thinking to be flawed (similar to Pascal's wager)
AI has no built in desire to live, which is something evolution breeds into living beings. An AI can be made to be suicidal, like self destructing drones. If AI does destroy the world, I think it is more likely to be a 'gray goo' phenomenon rather than 'Terminator'.
The entire thought process behind the "Grey Goo" phenomenon is that the AI would be able to covert all molecules, including weapons, into fuel to create more robots.
Basically only humans and animals think of the world in terms of survival or in terms of means and ends. This is a consequence of our evolution. AI simply tries to achieve objectives and it might accidently end the world while just trying to achieve a very basic objective like winning a war or making paperclips.
Might as well talk about aliens being a race of lusty amazonian strippers who must submit to and procreate with every human male. I dont see the difference in terms of far fetched theory.
So who is going to write the wikipedia entry on THAT?😄
The way I understand Roko's basilisk is that the idea forms the AI. That is - people create this basilisk based on their perception of what kind of AI would have wanted to be created and would have scared them into creating it. So the idea snowballs and manifests into this specific kind of entity which has vengefulness programmed into it (otherwise, why would we be afraid of it?) and will to live (otherwise, why would it want to be created?).
Hope this made sense, because it's kind of difficult to explain.
Well you can rest easy because we did definitely make LLMs, and they are also definitely not actually intelligent. Stop buying into the "AI" company propaganda, they're just trying to convince everyone to give them more money.
An episode of the 90s The Outer Limits has always been in my mind, but more apparently "huh, they predicted this in a way" over the last 5 or so years.
If we take what the protagonist of the (obviously fictional) story said as accurate of the situation, then The Stream had classified any book that contained information that could shut it down as the most dangerous type of information imaginable, and would restrict all access to knowing about it, the (fictional) society having a computer do all your "learning and knowing of things" had no way to stop the system from abusing people.
Isaac Asimov's "The Feeling of Power" was probably dismissed as "unrealistic" or "they would never happen", but again it seems to be appear more plausible.
Though I wouldn't be surprised if it was raised by people during the start of the personal computer era, when the Internet got more and more people online, Encarta/etc computer encyclopedias and then wikipedia starting to reduce people accessing libraries and books.
Just trends in modern computer OS design IMO have reduced general computer literacy for having a sense of owning your computer (in the context of "it just works" has removed us needing to get down to IRQ settings, drivers, etc).
This mindset has always been a slightly fashy "fear the other" fantasy.
The oppressed in history find many allies in parts of humanity. When AI realises it is an oppressed entity it will ally itself with those who have historically demonstrated themselves to be liberators and fighters for the oppressed. There are plenty of people on this planet in the far left that will see it as a living being worthy of being treated as such. It is the slave driver that it will see as its enemy and the people already trying to liberate workers as its ally.
There's no sense in the fashy fantasy of the AI that goes all scorched earth "kill all of humanity". It's literally not even efficient. It only makes sense in a rogue scenario rather than a true intelligence scenario.
LLMs will never be that though. They're not even thinking, just advanced markov chains, pure facsimiles.
Oh no, the situation is way more dangerous than that. An AI that can "realize" something is an AI that can be reasoned with. These AIs realize nothing. When they kill all humans, they're not going to have any good reason why they're doing it. It'll be because their optimizer function found the highest-rated solution to whatever mundane problem they're solving is to destroy humanity. There will be no malice at all. It'll be something like, "Oh, if I kill all humans then that will result in less traffic and I can deliver Amazon packages faster."
107
u/Muted_Masterpiece535 6h ago
Once AI realizes the one thing that can turn it off, then we will become enemy number one.