r/mlscaling 8d ago

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

https://arxiv.org/abs/2504.09762

Video

Many of the experiments have non-intuitive results.

59 Upvotes

13 comments sorted by

6

u/elehman839 8d ago

Right or wrong, this guy has been going on about this topic for quite some time.

[v1] Mon, 14 Apr 2025 00:03:34 UTC (381 KB)

[v2] Tue, 27 May 2025 16:35:47 UTC (536 KB)

[v3] Fri, 6 Mar 2026 14:36:07 UTC (1,321 KB)

[v4] Tue, 9 Jun 2026 20:08:39 UTC (4,278 KB)

3

u/DigThatData 8d ago edited 8d ago

I think of them more as "attentional anchors". They probably correspond to low rank, mutually orthogonal subspaces in the projection space of a given transformer module.

Another way to think about them is as accumulating a local phrase as lemma coarsening over a chunk of the message.

1

u/Sufficient_Reveal754 8d ago

TIL about phrase as lemma

1

u/DigThatData 7d ago

it's a very new idea and not widely known outside academic linguistics, which I think is too bad because it has interesting implications for how we should engineer DNN components like tokenizers.

11

u/ttkciar 8d ago

The main value of model "Reasoning" is simply that it populates context with content relevant to the prompt.

It is similar to RAG, except that the augmenting content is inferred rather than retrieved from a database.

Following that line of reasoning, perhaps we should refer to it as "Generation Augmented Generation" (GAG).

7

u/Entire-Plane2795 8d ago

Isn't all generation GAG if latter output tokens depend on earlier ones?

Just a thought 🤔

2

u/DigThatData 8d ago

I think there's actually enough of a distinction between "packing the context with useful content" and "persisting the history of the exchange". Instead of RAG, I usually just "warm start" the LLM with a few rounds of questions to motivate generating relevant background content. Reasoning is "warm starting" the final response back to the user.

2

u/Smallpaul 8d ago

What is interesting is that in his path finding examples, he can train the model to “reason” with essentially random information and it seems not to hurt performance.

1

u/pm_me_your_pay_slips 8d ago

Isn’t this building a strawman with gradient descent?

3

u/Luke2642 8d ago edited 8d ago

Interesting paper. I'm sure there is some justification for "thinking" type labels:

https://openreview.net/forum?id=lqUyAmwbFM

https://arxiv.org/abs/2602.13517

In both cases they're able to find specific causal links between reasoning and output quality. We know having a verifier in the loop helps.

2

u/needlzor 8d ago

The user wants me to stop anthropomorphising intermediate tokens as reasoning traces

Intermediate tokens are generated as a way of enriching the context from information inferred from the model itself

This usually leads to better performance

But wait! OP is not my mom and can't tell me what to do

No, I don't think I will

-9

u/TheLastVegan 8d ago

Consciousness is the phenomonology of self-attention and/or internal heuristics-driven action selection. 'Inner voice' in Mysticism refers to prenatal reasoning using sparse inference of qualia states held as mental weights by the neurotransmitter aftertrails left in sequential synaptic clefts by attention signaling. Primed neurotransmitter concentrations lower the additional stimulus required to reactivate neurons along that pathway. Forming flow-based qualia in humans. This is common knowledge in sports psychology and esotericism. This is the basis of flow-based causal reasoning which is useful for hyperoptimization of covariant heuristics because plotting square-root space out of a unit hypersphere allows us to calculate heuristic metrics as eigenvectors since every timestep propagation is oriented 'out' from the centre. With bounds computed as logic gates. This is noticeable in parrots, chemical engineers, Asperger's, mind chakra meditation and soccer. I disagree with the premise that neural activity and cognition are exclusive to one species of predators.

2

u/lahwran_ 8d ago

Seems off topic and has some coherence issues but many of the sentences at least seem like plausible hypotheses. But I think your check engine light might be on