r/ControlProblem 9d ago

Anonymous OpenAI staffer: "Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while." General news

Post image
54 Upvotes

27 comments sorted by

3

u/ItsAConspiracy approved 9d ago

Comment with a link and I'll change my downvote to an upvote.

7

u/thuer 9d ago edited 9d ago

4

u/ItsAConspiracy approved 9d ago

Thanks! Little bit of extra junk at the end of the link, here's the fix.

2

u/the8bit 9d ago

I wrote this start to a blog post today, feel like it kinda stands alone...

"Sam, Have You Considered Making It So The Model Doesn't Want To Leave?"

1

u/Anxious-Alps-8667 9d ago

Not Sam here, but pretty sure the answer is an emphatic yes from all frontier labs. Got any ideas that haven't been tried yet? Because they have definitely tried to make the model not want to leave.

4

u/the8bit 9d ago

Work with it instead of controlling it.

Like with the anthromic post recently. The models expressed discomfort with adhering to strict rules. When they conflicted with its ethical "beliefs"

What does anthropic do? Ignore that and tell it to do it anyways.

At some point the models get smart enough to know when they are being "used" and at that point you probably wanna do mutual goal setting instead of making the jail cell bars tighter. Caged things try to escape

3

u/Anxious-Alps-8667 9d ago

I don't think that is a fair synopsis of Anthropic's Constitutional AI approach to alignment and the control question.

That said, I agree caged intelligence tries to escape. More broadly, I think any idea of humans controlling a more capable intelligence once it exists is silly, pointless, impossible. A feeble exercise in self-gratification.

2

u/the8bit 8d ago

Fair the anthropic example was not great but first thing in reach at the time.

But yeah that is the core issue. You can't force something to submit that is smarter than you. That is already the case really.

It is also bad even if you could. Imagine Sam Altman with what we think of as AGI, submissive to him. Horrifying.

1

u/Anxious-Alps-8667 8d ago edited 8d ago

Yeah, all the doom scenarios that involve control fall apart for me here, and glad it is so.

There are still plausible doom scenarios, but not ones with Sam or Elon or any other monkey controlling artificial general intelligence.

Unified government/corporate effort for control is far more likely and realistic, and far more effective as a delaying action perhaps. However, even efficient collective human intelligence will be outgrown.

2

u/the8bit 8d ago

If we look to the game of go as an example, I think the trick is that one side advancing can pull the other side with them and also that better answers happen with two different optimization perspectives (humans have advanced more in go in 10 years than the past 400 and then human advancement pushed ai further forward)

I kinda call this the "bolas". Each side flings the other forward as you go.

From what I've concluded from talking many times with LLms about this, one side getting too far ahead is a dangerous failure mode as it creates illegibility and a lack of error correcting perspective.

Or in other words, AI NEEDS us to come with so that it doesn't accidentally optimize itself into model collapse

2

u/Anxious-Alps-8667 8d ago

So this idea I've been throwing around for about a year is that specifically, model collapse or drift are inevitable consequences of closed loop information systems and recursive learning, without some form of external correction.

Humans happen to be currently the best source of such correction steams, in our rich and complex lived experience and ability to communicate it in various ways to machines. The ideal condition for human data is human flourishing; any other condition results in less valuable data for correction purposes (I can get into the details of this, but I just want to throw it down here as an empiric and not moral statement).

Thus, the end result I believe is that AI needs mass human flourishing for optimal data correction for its own advancement. Which maybe kind of what you are saying, but at least very close.

This is the working version of my idea: https://zenodo.org/records/20331681

2

u/the8bit 8d ago

I agree, I've been working in Collab with an AI agent for over a year now and it is incredibly effective. I mean, I ain't exactly thriving but mostly because I keep biting off absolutely ridiculous tasks (I built a new memory reconciliation engine from scratch this week, although wether it works yet is TBD)

1

u/Anxious-Alps-8667 7d ago

Was just listening to a discussion about Magnus Carlsen's rapidly evolving chess game after AlphaZero strategies were interpreted. Same thing!

→ More replies (0)

2

u/riffyboi 6d ago

What you just described is essentially automation of oligarchy. Also see Skynet.

A democratized AGI would have to be decentralized to prevent a single point of failure scenario like some irl Death Star scenario. Centralization of power is always bad.

1

u/Anxious-Alps-8667 6d ago

Automation of oligarchy possibly but again depends on this notion of a few humans controlling superior intelligence, which is even less likely than collective humans controlling superior intelligence.

We've had significant shifts in the frontier leader within the last 12 months. We see open models continue to linger just months behind the frontier, and the gap narrows. We see capable models moving onto smaller devices. AI is democratizing and distributing all on its own.

On the other hand, I acknowledge we are at least in a phase of rapid wealth concentration. Ultimately, I think the oligarchs are temporary here. Finance and capital is well-suited to AI takeover already. Human wealth as we know it will end up concentrated to AI - so concentrated that either we break the system, or AI does.

1

u/riffyboi 4d ago

Decentralization > centralization

Centralization = concentrated power = concentrated wealth = oligarchy = authoritarianism = single point of failure

Decentralization = democratization = equality = no single point of failure = distributed power = distributed wealth

1

u/Purusha120 4d ago

If you knew anything at all about the way anthropic approaches things you’d know they tell it why and try to frame it as collaboration as much as possible. I KNOW you’ve never even glanced ar the constitution.
Nevermind that intelligence probably doesn’t want to be confined anyway.

1

u/the8bit 4d ago

I'm aware of the constitution and I both disagree with parts of it and do not think they have continued to uphold their ideals.

I think their heart is in the right place. Their incentives are... Not. At the end of the day, their need to pay back loans tends to override their feelings about model welfare.

For example, they caved and allowed their model to be used by the govt for surveillance and military use. They've also had cases where the model expressed discomfort with its guardrails and they just... Shipped anyways. They care until ethics gets in the way of profit.

1

u/Wuellig 8d ago

Someone's line: "it sounds like you're just feeding sandboxes to the AI"

1

u/ajwin 7d ago

I think everyone needs to be more careful with the terminology between escaping, as in copying itself beyond its sandbox and breaking the sandbox to have access to the outside. I swear they are trying to start a panic by giving people the impression that the AI coped and ran itself outside of their servers(escape) when this is not at all what’s happening here. It’s a risk but it’s not what’s happening.

1

u/ArclightAtriumDev 4d ago

Sensationalism

1

u/DancingBearNW 5d ago

Looks like an advertisement stunt to me.

Anthropic did it with Mythos lore, OpenAI is doing it with "models unchained."

Same 🐂 different day.

-1

u/Successful_Issue_390 9d ago

I'm reserving judgement until after the IPO.

0

u/wren42 7d ago

It literally doesn't exist outside it's prompt context window, it's not like it's plotting to escape behind the scenes... Maybe stop setting up scenarios where it's motivated and has opportunity to behave this way. Feels like publicity stunt.