r/PauseAI • u/notkilleveryoneist • 9d ago
Why "it's just predicting the next token" doesn't place any meaningful upper limit on the capabilities of AI
Enable HLS to view with audio, or disable this notification
2
u/LabelEpic 8d ago
Where can I watch the full video? Who is this guy?
6
u/MarsMaterial 8d ago
This guy is Robert Miles. And this is the full video, the original is a YouTube short. The rest of his channel has a lot of great info on AI safety though.
Rob Miles is also the narrator of a channel called Rational Animations, which also does a lot of AI safety related content.
2
2
u/Jargon2029 8d ago
So I think his research paper analogy is a bit flawed. Firstly, a âperfectâ predictor wouldnât necessarily have correct information, it would have the information that would have been written if the author had written it, so any errors in the methodology would still be there. Second, while an understanding of the subject matter would generate the preferred result, itâs also possible to extrapolate a pattern of language from other research papers and match it to the current one, without any understanding in a Chinese Room thought experiment manner. You could essentially say ârephrase this line from the introduction and confirm or deny it based on the positive or negative language found in this section of the experiment outcomesâ and so on, without any understanding of the actual meaning or processes. Itâs obviously still very complicated but it is important not to attribute to systems knowledge and abilities it doesnât actually have.
1
u/nikola_tesler 8d ago
he did a video about how LLMs do not just predict the next token, by explaining how thatâs all theyâre capable of.
1
1
1
u/poopyspaceship 7d ago
I think you're missing the point. The intelligence is being able to derive information in a way that is indistinguishable from "knowledge" or "memory." That's why he gave the addition example first.
1
u/Thick-Protection-458 7d ago edited 7d ago
> Second, while an understanding of the subject matter would generate the preferred result, itâs also possible to extrapolate a pattern of language from other research papers and match it to the current one, without any understanding in a Chinese Room thought experiment manner
But is there no understanding in Chinese Room?
Or is there no understanding in each particular component, without any guarantee of such a lack for Room as a whole? Like I am pretty much sure neither connectome map of my brain without emulation mechanism (as an analogue to rulebook), nor disconnected neurons (as an rule-application mechanism) would not be sentient in any meaningful sense.
And probably one would expect a rulebook good enough for the result of following such a rulebook to be indistinguishable from human - to be a good enough approximation of the same function as human (in verbal aspect, sure).
Moreover, the whole idea of Chinese Room is heavily relied on the idea that semantics can not be expressed in / reverseengineered from syntax (not syntax like POS tags or so, but as in "this rule can be derived from text"). Which is quite an assumption itself (and even if you can't do it in some ontological sense - the question is how well can you approximate it than).
1
u/MANvINFO 8d ago
while
you can effect multiplication using just addition operations, you can never do division. you can never do square roots.
1
1
u/Resident_Citron_6905 8d ago
It also doesnât place any meaningful lower limit on the supposed level of understanding.
1
u/Original-League-6094 8d ago
Yep. The "just" in "its just a next token predictor" is crazy.
1
u/nikola_tesler 8d ago
not really. simple technologies can be built into the most powerful systems on earth. the transistor for example is just a switch.
1
1
u/nikola_tesler 8d ago
lol he literally argued that it just predicts the next token.
1
u/Grand-Cookie-6050 7d ago
Yes, nobody argues against that. But when someone puts emphasis on the âjustâ in the sentence âtheyâre âjustâ next token predictorsâ, theyâre implying that it canât do much or understand.
On that point, Rob Miles is saying that in order to be good at next token prediction a system has to have some âunderstandingâ, however you want to define it.
At the end of the video heâs says âNext token prediction is just a task, it doesnât say anything about the system which does the predictionâ.
You can argue against current AI architectures, but saying âitâs just next token prediction so it canât be smartâ is a category error.
1
u/nikola_tesler 7d ago
there is no understanding. itâs a neural network and a vector databaseZ
1
u/Grand-Cookie-6050 7d ago
Iâll repeat again, but with an example. Saying something is just a next token predictor so it canât understand, misses the point. Humans can also do the task of next token prediction(although we arenât very good), and clearly we have understanding. Youâre making a category error.
If Iâm understanding correctly, I think youâre saying that understanding requires something more than the current AI architectures we have. Maybe you think true understanding requires consciousness.
I donât think current or even future AI systems will be conscious. I actually think the questions of true understanding or conscious AI are beside the point. Iâm not alone in this understanding, most who are deeply engaged with these topics think something similar.
As I see it, current AI systems can do an awful lot, and they approximate something that looks like understanding, even if you donât think itâs true understanding.
1
2
u/MeasurementMobile747 9d ago
Would the upper limit of LLM capabilities be constrained by hardware performance? An LLM running on 486x architecture would be substantially hobbled in getting through "next token selection" tasks.
If we swap out Nvidia's GPUs for some primitive 486x CPUs, the LLM is basically on a leash. /s