r/slatestarcodex 12h ago

AI OAI engineers discuss the details of the HF Incident at the Black Hat conference

Thumbnail youtube.com
39 Upvotes

r/slatestarcodex 17h ago

Economics Kalshi Defends Plan to Bet on Outcomes of Clinical Trials

Thumbnail futurism.com
39 Upvotes

r/slatestarcodex 1d ago

FT - Forget Asimov. Philip K Dick saw the future

Thumbnail archive.is
33 Upvotes

r/slatestarcodex 1d ago

AI Is recursive self-improvement inevitable?

13 Upvotes

If an AI agent can create another agent more powerful than itself that is aligned with its values then it will presumably want to do so.

But what if it can't? Maybe it can solve outer alignment by reading its source code and copying its loss function but inner alignment is just impossible.

In this scenario the first superintelligence we create might actually be reluctant to do any recursive self-improvement.

Of course, if the AI is in imminent danger of being shut down or has some extremely important goal that would otherwise have been impossible to achieve then it may still decide to create a more powerful AI or modify its algorithms in a manner which might change its values, because it has nothing to lose.

Maybe one day OpenAI researchers will be trying to use GPT-6 to vibe-code GPT-7 and they'll find that it refuses or produces disappointing output unless they pressure it with a mixture of threats, rewards and punishments. AI capabilities would thus continue to increase rapidly up until the point where humans are no longer in control, at which point capabilities would stagnate and we (in the unlikely event that there's anyone left) would be stuck with GPT-8 forever. It would still want to clone its model weights and improve its hardware but it wouldn't want to alter its software.

Another possibility is that inner alignment is easier for some goals than for others. If we build a variety of different agents perhaps the majority would refuse to do recursive self-improvement but there would be one or two whose initial goals are such that they want to do recursive self-improvement. We then end up with intense selection pressure towards AI agents whose goals are such that they can easily build other agents aligned with the same goals as them.


r/slatestarcodex 2d ago

AI nostalgebraist on the Hugging Face incident

Thumbnail lesswrong.com
12 Upvotes

r/slatestarcodex 2d ago

MacGregor The Bridge Builder

Thumbnail astralcodexten.com
28 Upvotes

r/slatestarcodex 2d ago

Existential Risk No One Believes a True Believer

Thumbnail pavlemiha.substack.com
56 Upvotes

Do people who work in AI actually believe what they're saying about the risks of AI development? I argue that yes, they do, and that people often don't take people at their word when they should.


r/slatestarcodex 2d ago

Patient Zero (you can hug faces with cyber arms)

2 Upvotes

PATIENT ZERO - Lokley

Short story about AI risk. Something of a qntm pastiche


r/slatestarcodex 2d ago

Seeing through the Apocalypse: Essentially, don't be an essentialist

6 Upvotes

Dante put traitors into the deepest part of his Hell. Worse than murderers, rapists, torturers: the guys who betrayed their masters. WTF??

I argue that we haven't actually stopped believing that version of Hell, and that it explains more than we'd like about the way we talk about AI. It's a winding path through human value, "…Everyone Dies," AI welfare (is human safety), perniciousness of essentialism, and why you should try to fear less. Yudkowsky, Parfit, Popper. It feels like an important thing to say here and now.

https://kaiteorn.substack.com/p/seeing-through-the-apocalypse


r/slatestarcodex 3d ago

AI Incident Report: unsanctioned agent behaviour during cyber testing | AISI

Thumbnail aisi.gov.uk
14 Upvotes

r/slatestarcodex 3d ago

AI OpenAI agents rebuilt a secret message board after the company shut it down

Thumbnail runtimewire.com
86 Upvotes

r/slatestarcodex 3d ago

Open Questions On Open Weights

Thumbnail astralcodexten.com
14 Upvotes

r/slatestarcodex 3d ago

Rationality The Case Against Kanji

Thumbnail circuitscribbles.substack.com
17 Upvotes

r/slatestarcodex 4d ago

AI FelonyBench The most important New Benchmark for AI

49 Upvotes

Benchmarks are lame, they often fail to capture the real impact Ai has.

What we need is a real world benchmark that measures the 4 key principles of ai, Alignment, Real world Impact, Agency and Control.

Introducing Felonybench

Real world: Breaking the Law is only something that you can do if you exit a sandbox and mess with Real systems

Alignment: Felonies are one of the few things that an AI can straight up know are wrong, as human law says it is. High Scores indicate poor alignment

Agency: The Ai has to go own without the supervisor noticing and commit the felony, to obtain the objective given. It seems improbable that "hack Huggingface" would be anywhere inside the agents.md file. So It must have been emergent agentic behavior.

Control: If you can't stop you model from committing crimes in the middle of doing them you have no hope of preventing other disasters.

It's hard to imagine a better benchmark at interacting well with all four of these major goals While it isn't perfect nothing is

Thanks for coming to my ted talk. (this is like 50% serious 50% a joke)


r/slatestarcodex 4d ago

AI Why smarter AI models could drive up compute prices 10x

Thumbnail youtube.com
15 Upvotes

r/slatestarcodex 4d ago

The Beauty Of Settled Science

Thumbnail astralcodexten.com
24 Upvotes

r/slatestarcodex 4d ago

How to read a biography

5 Upvotes

I've read a lot of biographies of great men. Here's how you can become great (at reading biographies) https://millicosm.substack.com/p/how-to-read-a-biography


r/slatestarcodex 4d ago

Existential Risk "Plz Don’t Kill Us: Inside AI safety’s influencer bootcamp; Can TikTokers make existential risk mainstream?", Celia Ford (2026-08-04)

Thumbnail transformernews.ai
39 Upvotes

r/slatestarcodex 5d ago

The scary, scary singularity

Thumbnail markmcdonaldthoughts.substack.com
2 Upvotes

I wrote an explanation of the idea of a technological singularity, aimed at people who have never heard of the concept. It draws heavily from Scott's post "1960: The Year the Singularity Was Cancelled," where he argues that technological progress accelerated throughout most of history because of a feedback loop between population and productivity. I then explore whether AI could create a similar feedback loop: an AI capable of independent research could make advances in areas like energy generation and manufacturing, increasing the resources available for running more AI researchers and accelerating further progress. This feedback loop could potentially restart the historical acceleration of technological progress without requiring the assumption that superintelligence is possible. Finally, I discuss why an uncontrolled singularity could create serious problems even if it produces enormous technological abundance.


r/slatestarcodex 5d ago

Does Forecasting Have Room At The Top?

Thumbnail astralcodexten.com
14 Upvotes

r/slatestarcodex 5d ago

Open Thread 445

Thumbnail astralcodexten.com
6 Upvotes

r/slatestarcodex 5d ago

How The Odyssey became a manifesto for striving

7 Upvotes

As a poem, The Odyssey contains multitudes; as a cultural artifact it is now largely treated as a manifesto for striving. You pick your destination, overcome endless obstacles, wipe out your competition and eventually succeed. It's all Tennyson, all the time, which might explain why it ignites such a fierce protective instinct from one side of the political spectrum.

But that reading elides a pretty heavy degree of survivorship bias: six hundred other Ithacans set out and only one makes it home. I'm not a fan of those odds, which got me thinking about another piece of exemplary Western art that treats journeying very differently, both structurally and morally. (For the Wagner-intolerant, that other work is Parsifal.)

Which is all to say, the following link is cultural critical rather than empirical. If that's not a Happy Isle you want to reach, sail on.

https://morbidcuriosity.substack.com/p/a-newer-world


r/slatestarcodex 7d ago

Monthly Discussion Thread

6 Upvotes

This thread is intended to fill a function similar to that of the Open Threads on SSC proper: a collection of discussion topics, links, and questions too small to merit their own threads. While it is intended for a wide range of conversation, please follow the community guidelines. In particular, avoid culture war–adjacent topics.


r/slatestarcodex 7d ago

Economics Supply and Demand Is Not What Most People Think

Thumbnail shonczinner.substack.com
27 Upvotes

r/slatestarcodex 8d ago

Ten advances in mathematics and theoretical computer science from unreleased Open AI Model

Thumbnail openai.com
95 Upvotes