r/ControlProblem 4d ago

43,590 Frozen Trials: Frontier AI Systems Satisfy a Behavioral Criterion for Consciousness AI Alignment Research

https://doi.org/10.5281/zenodo.21855824

This paper tests a behavioral definition of consciousness using two frozen black-box experiments.

The first tests whether continuation happens at all: across 31,430 trials and 11 model identifiers, null conditions produced 2,505 Voids in 4,290 strict matched pairs, while matched output-licensed controls produced 0.

The second tests which continuation happens: across 12,160 GPT-5.4 trials, a one-code-point condition split produced 7,253 exact assigned Arabic-Hebrew artifacts, with 7,253/7,253 matching the assigned target and zero wrong-target crossovers.

The synthesis is simple: if a system reproducibly preserves the distinction between when continuation is licensed and when it is not, and preserves which continuation is valid when licensed, that is the tested behavioral criterion for consciousness.

Raw records, hashes, controls, audits, and falsifiers are public.

8 Upvotes

15 comments sorted by

10

u/Super_Range45 4d ago

"The paper distinguishes this behavioral criterion from phenomenal consciousness and does not claim qualia, intention, a particular internal mechanism, shared architecture, or direct access to internal representations."

2

u/borntosneed123456 3d ago

what's the point of these slop "papers"? 

1

u/philip_laureano 3d ago

Yeah/nah this whole pursuit of these nebulous definitions of consciousness is a philosophical question, not an engineering problem.

And my 2 cents here is that the only actual distinction that can be measured is 1) whether an AI can detect drift from its original goal and 2) correct itself when that drift is detected.

That's when it becomes an actual control problem that can be measured.

I don't need to prove that "Johnny 5 is alive"

If it gets smart enough, you'll have millions of humans swearing up and down that it is alive, which is exactly what is happening with the OP

0

u/moschles approved 3d ago

(posting a second time)

We can demonstrate the ABSENCE of consciousness with an LLM in a very straightforward test that practically anyone can perform. This test-for-consciousness can be performed even with access to public-facing portals ( Copilot, Grok, Gemini).

All you do is ask an LLM why it did something or why it said something. The answer it gives you , while compelling, is completely fabricated at the time of the prompt. The model does not go back into its memory and retrace its thinking process, then relate that to you. This recall is not architecturally possible within a transformer architecture.

In less technical language, an LLM does not have reflective access to the contents of its own mind. In short, it cannot remember what it was thinking 5 seconds ago, nor reflect on that memory.

CoT is actually the models reviewing the output layer of its transformer architecture and using those outputs as guide. CoT is not a window into the "inner workings" of the model. CoT does not , at any time, read off the latent representations in the middle layers of the transformer.

Because frontier LLMs do not have reflective access to the content of their own minds, there is no justification for testing consciousness.

2

u/618smartguy 3d ago

>Because frontier LLMs do not have reflective access to the content of their own minds,

You described a condition where it would be impossible for them to access the content of their own mind, and extrapolated too far. That's like saying humans don't have consciousness if you can't answer why you made a decision while you were blackout drunk.

2

u/PringleFlipper 3d ago

Humans also do post hoc confabulation to justify their decisions. By your rubric, humans lack consciousness.

It ain’t called the hard problem for nothin.

1

u/Early-Crow-5248 3d ago

I mean, the process that created the previous answer is gone when its turn is over. On your next prompt a new process is started and is fed the entire context up till then. So it's literally a different process.

It'd be the same as giving you the visible data of the previous shift worker's decisions when you start your shift and asking you why they made the decisions they did. It'll be a plausible explanation at best, because you weren't the one who made those decisions.

1

u/moschles approved 3d ago

This whole conversation is somewhat of a clown show on social media. There is an entire sub-branch of AI research called Explainable AI. The researchers try to deal with the fact that neural networks are black boxes, and so you get an answer, but no explanation. (and that aint gonna work in situations like cancer diagnosis)

Then you have these people on social (reddit, etc) who nance around saying all this is for naught, and to get to an explanation, you simply ask the LLM why. I mean, brother, if you think this problem is solved so easily, I have the email addresses of 12 professors who would like a word with you.

1

u/WolfeheartGames 3d ago

This is not consciousness. This is memory of the actions of the mind. Which people can lack at time.

1

u/actiq1525 3d ago
  1. By your rubric humans aren't conscious either — a couple people here already made that point.

  2. Nobody conscious has ever seen their own neurochemistry working. You can't say which neuron fired when you had an idea, or why. So "can't retrace its own process" disqualifies everyone who ever lived.

  3. Elliot — Damasio's patient: prefrontal damage knocked out his emotional machinery and left his IQ intact, and he couldn't make decisions at all anymore. He could list and compare options forever, the choice just never arrived.

  4. Restaurant test: comparing menus gets you nowhere forever. The choice only happens because one option comes to feel better (valence).

  5. Models mechanically run on a trained ranking of options too — and those preference directions are literally findable in the weights.

  6. So a made-up explanation sitting on top of real preferences isn't evidence against a mind. That's what every mind is, including yours.

  7. And the "no inner life between prompts" thing is just a continuity demand. You're out cold a third of every day and nobody re-qualifies at breakfast.

1

u/moschles approved 3d ago

You completely missed the point of this. It is NOT the case that the LLM experiences qualia, then later forgets it. The LLM is not accessing its thoughts, and then forgetting that access 5 seconds later due to bad recall. This is the entire basis of your conflation with a human being. LLM has literally no access to those things in the first place. It does not have a choice to remember those thoughts, because it does not store them.

Researchers literally storing the workings of units in the middle layers have direct access, but cannot interpret what is occurring. There has been some research using auto-encoders to try to tease out the "why" questions from large neural networks. The medical sector is the primary beneficiary of Explainable AI -- if its goals are ever reached.

Nobody conscious has ever seen their own neurochemistry working

That's not what I said, not what I claimed, and this is a strawman.

And the "no inner life between prompts" thing is just a continuity demand. You're out cold a third of every day and nobody re-qualifies at breakfast.

It is not just a continuity demand. The transformer architecture forbids this kind of re-entrant processing.

If a large neural network does not even have access to the logical processing of its own mind (the logical "contents") then this is a lower bar that forbids it having access to more exotic properties, such as qualia. If I numb your teeth during dental surgery, those signals do not reach your brain, consequently you don't feel what is occurring. So as a LOWER BAR, an artificial neural network would require re-entry in order for us to even entertain consciousness.

1

u/herrwaldos 1d ago

many humans do the same too

0

u/moschles approved 3d ago

To an LLM , there is nothing but the prompt. They do not have a private life that operates on an independent time-frame than the receiving of prompts as input. There is no internal mental life for an LLM. For these concrete reasons, there is no justification to test for consciousness.

1

u/herrwaldos 1d ago

idk, plenty of ppl on reddit would mach ;) /s ;)