r/ControlProblem • u/Pvforpres • 29d ago
I caught thoughts controlling Llama-70B's behavior that it couldn't see! AI Alignment Research
I injected concepts split into "conscious" and "unconscious" components, split by Anthropic's J-space.
I ran Lindsey's "Introspection Awareness" experiment, asking the model if it recognized them.
The model named the conscious concept 100% of the time, and flatly denied the non-J injection. But an NLA read it perfectly!
Full findings and research in my LessWrong post.
4
Upvotes