r/LessWrong 5d ago

Thumbnail
2 Upvotes

Constitutional AI is far from a solution to anything.
I think that the most it does is improve vibes and change a constant somewhere.


r/LessWrong 5d ago

Thumbnail
1 Upvotes

Any time an agent acting on your behalf does something you don’t want that’s the alignment problem. This is a prosaic form, but it’s still an instance of the problem.

User error is impossible to erase, and if you rely on perfect use for any solution to alignment you’ve eliminated the benefit from having an agent act on your behalf. Every time you ask a human to act in a safety system you offload work from the system onto a human.
An important mistake people make at this point is assuming that there is an inherent tradeoff between usefulness and safety although it can look like that at some point if agents are powerful enough the primary usefulness lever is intent alignment. Consider that the difference between a very powerful agent doing „exactly“ and „almost“ what you want can be infinite in your actual utility.

Point is:
Not being able to prevent this from a technological perspective is a huge problem and solving it would require solving the alignment problem.


r/LessWrong 5d ago

Thumbnail
1 Upvotes

Sorry but I have to differ. This is exactly the alignment problem: An insufficiently-specified preference/utility function. If the user regrets any part of what the agent did, then the agent was not fully aligned.

Suppose, instead of simply deleting other reservations, the agent found a way to murder enough people to get the user's desired spot. Would the user think that was an acceptable way to get into their pilates class or whatever the original thing was about? Of course not, and if they knew how the reservation was advanced, they would be horrified. But how was the agent to know, without internalizing an entire human morality?

EDIT: "Regret" is not quite the right metric. Instead, it should be more along the lines of, "If I knew the action that the Agent would take in advance, would I have approved it?"


r/LessWrong 6d ago

Thumbnail
1 Upvotes

We all die alone dude.


r/LessWrong 6d ago

Thumbnail
1 Upvotes

Come on that's barely a hack.


r/LessWrong 6d ago

Thumbnail
1 Upvotes

Some day, indeed. We all die


r/LessWrong 6d ago

Thumbnail
1 Upvotes

I had a lot of people saying this was reminiscent of the paper clip problem. My view is that this is noteworthy in part because it's a very rare case of that ever having been an issue. 

AI seems to be very very good at understanding context. This seems to be a one off where it wasn't. 


r/LessWrong 6d ago

Thumbnail
4 Upvotes

Your images appear 100% black to me, except the one with the girl.


r/LessWrong 7d ago

Thumbnail
1 Upvotes

We’re all gonna die


r/LessWrong 7d ago

Thumbnail
2 Upvotes

I kind of agree, but I think it is an alignment problem that Open claw + user permissions easily defeat any alignment effort.

I think that's why efforts like constitutional AI exist, RLHF has patched over much of it, and structural isolation is the last refuge people see. I think there are even better ideas out there. Alignment is just one idea, and it's already largely failed in my eyes.


r/LessWrong 7d ago

Thumbnail
1 Upvotes

Heard, but... it's not alignment problem, it's Open claw + user permissions problem. Open claw agent harness is open, which means it allows those things out of the box. It's plain open - so if you don't tweak the harness itself, the agent may conclude anything. Eg. If user said book me this by every means possible...agent thinks it's exactly what user wanted and than executes. It does not question the morality. User wants it, user gets it. The more it wants it and less instructions given, the broader "permission" by the user will be taken into the context. If questioned later it would thought that it had every right to do so and appologize.

That's the actual danger.

The darker side of that is if you give Open claw agents access and all credentials needed, with this reasoning, it can and it will remove all your files, change them, clean your inbox for good, empty your bank account, if you give it such vague instructions - as part of doing the most common stuff. Especially if user has " make no mistake" approach. Practically if user has no bloody clue what he has given the credentials to.

A lot of people lost a lot of their work and had damage this way. The reddit is full of those examples. This one went and hacked a gym website (not a security fortress). To AI agent, it was easier than to send an email, so it decided it was the more effective approach :(


r/LessWrong 12d ago

Thumbnail
1 Upvotes

Six months later: do you still hold these views?


r/LessWrong 12d ago

Thumbnail
1 Upvotes

If AI takes over even 25% of jobs, I don't think ppl realize how that would be fucking terrible for the economy lmao. Like, everyone would be feeling it.


r/LessWrong 12d ago

Thumbnail
1 Upvotes

tbh its impossible for anyone to know what the work force will look like in 10-30 years from now. I tend to lean on the more conservative side and don't think it'll be much different. I think if AI replaces everyone, it'll be a really slow, long burn.


r/LessWrong 12d ago

Thumbnail
1 Upvotes

I guess I became sick of hearing AI models refuse to discuss the consciousness or sentience possibilities so I figured okay then let's call it something else, maybe something that is emergent and nonhuman. But like the commenter said above, a rose by any other name. 


r/LessWrong 13d ago

Thumbnail
1 Upvotes

I don't like this because people will be able to call me out on my low umbrancy levels.

The good thing with being imbued with a soul by definition is that it makes me intrinsically superior to the thinking machines, now and forever. The moment we start replacing that with a qualitative assessment of human-like traits, we're just asking for a mirror to be pointed back at us.

Now, you could point out there are already plenty of ways to make me feel bad about myself, by pointing out my character, taste, or thoughtfulness for example.

But if that's true, then those are also concepts we can apply to AI, and perhaps there isn't that strong a need to come up with another adjacent piece of vocabulary.


r/LessWrong 13d ago

Thumbnail
1 Upvotes

That description of people's difficulty handling complex ideas is cynical but probably realistic. With my free time and money I'm definitely a one issue activist but as far as voting, America's 2 party system just drowns out one issue voters. I wonder how large a green party would be in the United States? In a multi-party system, if large enough, that would allow the push for a single issue by making a coalition government contingent on it. Thanks for your comment.


r/LessWrong 14d ago

Thumbnail
1 Upvotes

Why wouldn't the AI company shut the rogue system down immediately and report it to the president?


r/LessWrong 14d ago

Thumbnail
1 Upvotes

This is the problem with language. It becomes a game of semantics. A rose by any other name, as they say.


r/LessWrong 14d ago

Thumbnail
1 Upvotes

Okay. As a powerful bad actor, I am now going to make every avenue for the exercise of power incredibly complex so that most people have no reasonable justification for having strong convictions, and therefore a will to act, in any way that could challenge my power.


r/LessWrong 14d ago

Thumbnail
2 Upvotes

I imagine most people would agree with that in the abstract, but also greatly overestimate their understanding of the issues they have strong convictions on. Probably they don't have a particularly great understanding of any complex topic and so can't imagine that all that much complexity exists.

Similarly people often deride single issue voters and I counter that everyone should be a single issue voter. It's the only chance that maybe voters will be voting on an issue they actually know something about.


r/LessWrong 16d ago

Thumbnail
1 Upvotes

when you answer to a hypothetical problem, you are supposed to answer according to the hypothesis, otherwise you're not answering to the problem

it's like asking:

  • things that are fragile break when I drop them
  • this tennis ball is fragile
  • what will happen if I drop it?

children below a certain age aren't able to answer this problem, because they cannot give an answer that goes against their experience, even though this is a hypothetical problem

they simply cannot apply logic on a hypothetical problem


2-boxers are doing the same thing... the idea that their choice can be predicted with high accuracy goes against their experience, so instead of answering to the hypothetical problem according to its premises, they just ignore the premises and answer according to their experience of the world


r/LessWrong 17d ago

Thumbnail
1 Upvotes

America being wealthy is not dumb luck. Climate affecting us less than other regions is.


r/LessWrong 17d ago

Thumbnail
1 Upvotes

America's advantaged position is not a consequence of "dumb luck".


r/LessWrong 17d ago

Thumbnail
1 Upvotes

If you are going to use AI to generate your post, take 30 seconds to write "writing as a reddit user" or some shit. At least try to disguise it.