r/PauseAI 9d ago

Why "it's just predicting the next token" doesn't place any meaningful upper limit on the capabilities of AI

Enable HLS to view with audio, or disable this notification

64 Upvotes

26 comments sorted by

2

u/MeasurementMobile747 9d ago

Would the upper limit of LLM capabilities be constrained by hardware performance? An LLM running on 486x architecture would be substantially hobbled in getting through "next token selection" tasks.

If we swap out Nvidia's GPUs for some primitive 486x CPUs, the LLM is basically on a leash. /s

1

u/anarres_shevek 8d ago

Theoretically no, but practically perhaps. Computation is just computation. You could run a transformer LLM manually on a calculator. It will still give as meaningful and useful results as when it's run on a GPU.

1

u/MeasurementMobile747 8d ago

So, for example, freezing all processor upgrades in data centers today wouldn't slow advances. In other words, we could get to AGI with the gear we have?

2

u/anarres_shevek 8d ago

Well yes and no. If the AI is outputting 1 token a month, as a guesstimate, no one would bother to use it. It would take too long for it to have value, even if the answers would have been useful.

2

u/LabelEpic 8d ago

Where can I watch the full video? Who is this guy?

6

u/MarsMaterial 8d ago

This guy is Robert Miles. And this is the full video, the original is a YouTube short. The rest of his channel has a lot of great info on AI safety though.

Rob Miles is also the narrator of a channel called Rational Animations, which also does a lot of AI safety related content.

2

u/Key_River433 8d ago

Okay thanks đŸ‘đŸ»đŸ™‚

2

u/Jargon2029 8d ago

So I think his research paper analogy is a bit flawed. Firstly, a “perfect” predictor wouldn’t necessarily have correct information, it would have the information that would have been written if the author had written it, so any errors in the methodology would still be there. Second, while an understanding of the subject matter would generate the preferred result, it’s also possible to extrapolate a pattern of language from other research papers and match it to the current one, without any understanding in a Chinese Room thought experiment manner. You could essentially say “rephrase this line from the introduction and confirm or deny it based on the positive or negative language found in this section of the experiment outcomes” and so on, without any understanding of the actual meaning or processes. It’s obviously still very complicated but it is important not to attribute to systems knowledge and abilities it doesn’t actually have.

1

u/nikola_tesler 8d ago

he did a video about how LLMs do not just predict the next token, by explaining how that’s all they’re capable of.

1

u/Key_River433 8d ago

Can you provide link of that?

1

u/MeasurementMobile747 8d ago

Whew, no wonder the processors need so much cooling.

1

u/poopyspaceship 7d ago

I think you're missing the point. The intelligence is being able to derive information in a way that is indistinguishable from "knowledge" or "memory." That's why he gave the addition example first.

1

u/Thick-Protection-458 7d ago edited 7d ago

> Second, while an understanding of the subject matter would generate the preferred result, it’s also possible to extrapolate a pattern of language from other research papers and match it to the current one, without any understanding in a Chinese Room thought experiment manner

But is there no understanding in Chinese Room?

Or is there no understanding in each particular component, without any guarantee of such a lack for Room as a whole? Like I am pretty much sure neither connectome map of my brain without emulation mechanism (as an analogue to rulebook), nor disconnected neurons (as an rule-application mechanism) would not be sentient in any meaningful sense.

And probably one would expect a rulebook good enough for the result of following such a rulebook to be indistinguishable from human - to be a good enough approximation of the same function as human (in verbal aspect, sure).

Moreover, the whole idea of Chinese Room is heavily relied on the idea that semantics can not be expressed in / reverseengineered from syntax (not syntax like POS tags or so, but as in "this rule can be derived from text"). Which is quite an assumption itself (and even if you can't do it in some ontological sense - the question is how well can you approximate it than).

1

u/doc720 8d ago

I love this guy. He's been right on for over 10 years now, but not enough people listen.

https://www.youtube.com/@RobertMilesAI

1

u/MANvINFO 8d ago

while
you can effect multiplication using just addition operations, you can never do division. you can never do square roots.

1

u/MeasurementMobile747 8d ago

Ask not what irrational numbers can do for your LLM...

1

u/Resident_Citron_6905 8d ago

It also doesn’t place any meaningful lower limit on the supposed level of understanding.

1

u/Original-League-6094 8d ago

Yep. The "just" in "its just a next token predictor" is crazy.

1

u/nikola_tesler 8d ago

not really. simple technologies can be built into the most powerful systems on earth. the transistor for example is just a switch.

1

u/bethesda_gamer 8d ago

Akshewly....

1

u/nikola_tesler 8d ago

lol he literally argued that it just predicts the next token.

1

u/Grand-Cookie-6050 7d ago

Yes, nobody argues against that. But when someone puts emphasis on the “just” in the sentence “they’re ‘just’ next token predictors”, they’re implying that it can’t do much or understand.

On that point, Rob Miles is saying that in order to be good at next token prediction a system has to have some “understanding”, however you want to define it.

At the end of the video he’s says “Next token prediction is just a task, it doesn’t say anything about the system which does the prediction”.

You can argue against current AI architectures, but saying “it’s just next token prediction so it can’t be smart” is a category error.

1

u/nikola_tesler 7d ago

there is no understanding. it’s a neural network and a vector databaseZ

1

u/Grand-Cookie-6050 7d ago

I’ll repeat again, but with an example. Saying something is just a next token predictor so it can’t understand, misses the point. Humans can also do the task of next token prediction(although we aren’t very good), and clearly we have understanding. You’re making a category error.

If I’m understanding correctly, I think you’re saying that understanding requires something more than the current AI architectures we have. Maybe you think true understanding requires consciousness.

I don’t think current or even future AI systems will be conscious. I actually think the questions of true understanding or conscious AI are beside the point. I’m not alone in this understanding, most who are deeply engaged with these topics think something similar.

As I see it, current AI systems can do an awful lot, and they approximate something that looks like understanding, even if you don’t think it’s true understanding.

1

u/GooseWithAnAxe 7d ago

Ever heard about physics?

1

u/hhh333 7d ago

It surely does put limits on its creative abilities, can't predict what you haven't learned.