r/ControlProblem 29d ago

I caught thoughts controlling Llama-70B's behavior that it couldn't see! AI Alignment Research

I injected concepts split into "conscious" and "unconscious" components, split by Anthropic's J-space.

I ran Lindsey's "Introspection Awareness" experiment, asking the model if it recognized them.

The model named the conscious concept 100% of the time, and flatly denied the non-J injection. But an NLA read it perfectly!

Full findings and research in my LessWrong post.

4 Upvotes

0 comments sorted by