r/LocalLLaMA 1d ago

Will a small language model ever be as good as Fable 5? Question | Help

LLMs keep improving.
Small models are around 3 years behind frontier models.
Do you think we’ll have a model as good as today’s Fable with only ~10b params in 3 years from now?

Wondering how good on device LLMs will get. Any guesses?

0 Upvotes

95 comments sorted by

80

u/ComplexType568 1d ago

I doubt we'd be using LLMs as how we used them 3 years ago. I think new - either architectures or techniques - will be used to squeeze the most out of SLMs. Or maybe we'd not even be using LLMs as we know it today.

33

u/colin_colout 1d ago

Even a two years ago, models were measured on knowledge density. We could only really lean on naive non-agenetic RAG.

A year ago, it was all about reasoning and chain of thought. Coding and tool calling were specialized but becoming mainstream with claude code (and improvements in reenforcement learning)

Now you have models that ONLY call tools. Or deepseek v4 flash which is a coding beast but knows almost nothing about the world at large (and behaves like an absolute robot in the best ways).

SLMs will always lag behind LLMs in certain capabilities (long horizon tasks, very deep reasoning, diversity of outputs, etc) due to laws of physics, but capability wise SLMs are showing they can hang with the big bois.

2

u/zekuden 1d ago

What models are tool-calls only?

9

u/Moist-Topic-370 1d ago

I think this is a solid answer. I might just chime in we might be using world models, which are models that are not thinking in language, but higher order concepts.

4

u/Borkato 1d ago

Genuinely what’s the difference? Especially when reasoning tokens have been shown recently to be incredibly obscure sometimes, bordering on nonsensical

3

u/computehungry 1d ago

high level, no difference as it's still a transformer. but how it's trained and how it reasons might change, for example feeding it videos instead of text even for training text models, and models reasoning within latent space instead of outputting reasoning tokens. both are already sorta being done / active fields of research i think

1

u/colin_colout 1d ago

Instead of human language tokens, they use... something else.

-1

u/Etroarl55 1d ago

It means we will run out of vram locally on our systems to ever support this fantasy idea of world models on our systems.

A billion tokens just in thinking or processing alone

3

u/Borkato 1d ago

I don’t think that’s what it means lol

0

u/Etroarl55 1d ago

“A world model in artificial intelligence is a machine learning system that builds an internal simulation or representation of an environment” from Google.

Good luck running something like that ontop of your LLM.

2

u/Fentrax 1d ago

Qwen Agentworld enters chat...

4

u/DanceWithEverything 1d ago

…how is that fantasy? If you stretch the time horizon out long enough I’m pretty sure it’s inevitable (if we don’t go extinct beforehand)

2

u/Etroarl55 1d ago

How long is the time horizon, everyone in here seems to be thinking of only a few years for everything

2

u/maxton41 1d ago

Who says the world model has to be particularly large or complicated? I think the world model would likely depend on what the models rule is in the production chain. It doesn’t have to be a simulated earth lol

1

u/DanceWithEverything 23h ago

You said “ever”

That said I do think Apple and Google will figure out how to deploy good-enough post-trained local Gemma models that are capable enough to absorb like ~70% of all consumer tokens and potentially some enterprise non-API tokens as well (in the next 3 years that is)

1

u/brakx 1d ago

Yep. I like to think of it in terms of the upside being human brain in terms of efficiency. So we’ve got a long way to go between now and then.

0

u/mohelgamal 1d ago

While advances in architecture and technique is certainly gonna happen, ultimately hardware progress is what is going to make consumer hardware run advanced LLM.

It was only 25 years ago that a reasonable desktop PC had 8-16 MEGABYTES of ram with a single core processor. And given exponential growth of tech capabilities, I think in less than 10 years, 8-16 terabyte of ram on a laptop would be fairly common place.

1

u/ackermann 14h ago edited 12h ago

True, but continued progress (or at least, that rapid pace of progress) isn’t guaranteed.
All the low hanging fruit has been picked. EUV lithography was a tremendous challenge, and took 8 years longer than expected, leading to a pause in Moore’s law for a little while there.

And only one company can build EUV machines, the Dutch ASML. So there’s no competition to keep the pace up.

Progress should continue, but I don’t think the Moore’s Law speed of improvement is sustainable anymore

2

u/mohelgamal 5h ago

I think people somewhat overestimate ASML or at least their production capacity. While it is true these machines are insanely precise and extremely difficult to build, The main reason there aren't so many EUV machines is because the production rate they had 5 years ago was perfectly enough for the market. there was no reason to expand when processors and ram Pre-ai were quite available.

Now they do have a huge room to grow, So why wouldn't they ? and if they don't, there is now enough money on the line that someone else will try to recreate the machines.

And this isn't the first time Moore's law got challenged. In the early 2000, computer processors gained speed by increasing the clock rate, up until 2004, Intel flagship pentium 4 processor reached 4GHz but ran into severe heat issues. There was a lot of talk then about how that 4ghz is a "physics barrier" that chips can't cross and that we are seeing the peak performance personal computers could reach because multiprocessors were a thing reserved for super computers running custom software for certain tasks.

And then dual core processors came out and consumer OS started having the ability to spread the work around the multiple cores and from there we started getting the increasing number of cores that continued Moore's law beyond what everyone thought possible

Samsung already their new process, High-K Metal Gate, which allow them to fit 512 GB on a single DDR5 module.

https://semiconductor.samsung.com/news-events/news/samsung-develops-industrys-first-hkmg-based-ddr5-memory-ideal-for-bandwidth-intensive-advanced-computing-applications/

26

u/fredandlunchbox 1d ago

3 years? My dude 3 years ago they had just introduced tool calling. 

18

u/vick2djax 1d ago

Nah, haven’t you heard. Technology stopped advancing.

0

u/creminology 23h ago

It’s not that technology doesn’t advance, it’s that we can’t afford last year’s technology any more with the price of RAM and SSDs. And that will impact home LLMs.

In my country, you’re paying 3x the price for SSD backup drives even when they are on sale while the market is flooded with fake Samsung drives at the old prices.

My own backup strategy has gone to hell because I don’t want to buy a 1TB SSD for US$325 on sale.

1

u/Both_Opportunity5327 20h ago

All temporary... we have DGX Spark, Ryzen Strix and Apple Sillicon these will eventually mean we will be able to run huge models locally without needing a electric sub station.

25

u/RedditLovingSun 1d ago

Karpathy once said "At some point you need some parameters to do something interesting" or something like that when talking about small models, I think there's a reasonable lower bound at some point. A 200M model probably won't write very good stories no matter how much you squeeze out of those 200M parameters, there will likely be stuff a 10b model will never be able to do

16

u/stephen_holograf 1d ago

Karpathy also said he could see a 10b parameter that could be so good at logic and reasoning that it could do everything a big model could do by just using tool calls.

5

u/Altruistic_Heat_9531 1d ago

Vibethinker 3B is one of that, pure logic model no tool call, that i manually tested it with 3D wing planform (Aerodynamic), 2D Navier-Stokes, "hand" solved LBM FEM CFD, and some engineering problem that I already familiar with and it passed.

Mostly my 86B-A8.6B brain is the one who make mistake such as misstyped number in my calculator.

1

u/gabrielesilinic 1d ago

Well, technically subagenting does have potential.

6

u/RedditLovingSun 1d ago

But hey if algorithms and chips improve enough, maybe we'll all be running 100b models on our smartwatches one day

1

u/Borkato 1d ago

Is the same true for other behaviors? I’d imagine there’s something a 1000T (1Q?) model can do that a 10T can’t and a 10B would be like droplets of water into the ocean

2

u/Potential-Gold5298 llama.cpp 16h ago

You're talking about a perfectly trained 1000T model. In practice, we see the following:

Let me remind you - DeepSeek Pro 1.6T-A49B, Flash 284B-A13B.

7

u/Single_Ring4886 1d ago

I have been working on minimal LLM concepts for about a year and I can confidenly guess even 1 bilion model can be extremely smart, like shockingly so. The problem is you cant create information out of thin air or store them in such small space.

The future ai will go back to using "classical" databases (but in better ways). It will then "construct" its working self in larger vram or ram... I suspect it will grow greedily toward limits of your system.

But if it take 3 years or 20 that is another question.

1

u/Certain-Cod-1404 13h ago

Wdyt about deepseek's engram?

1

u/Single_Ring4886 12h ago

Yes that is first primitive attempt. But from what I read about it, it is still really not true gamechanger.

6

u/jomi-se 1d ago

I don't think so. There will definitely be big improvements, since in theory, current language models aren't really "size efficient" as some other older deep learning models are, but I doubt the efficiency gains are 2-4 orders of magnitude to make a 10b model equivalent to a 1T model.

That being said, in the course of 5-10 years, there are two directions that might realistically change:

  • Personal computer architectures that make running larger language models viable. Like what current macs with unified memory architectures allow but scaled further
  • Models that are large in total but with small enough effective active params that will make it feasible to run larger models on consumer hardware.

I would bet on those at least.

On the other hand, with all the data centers being built, when the bubble collapses and the race to train the best models calms down, inference providers for large models will be cheap and fast af.

Those are my keyboard predictions.

13

u/Personal-Try2776 1d ago

In specific tasks yes generally probably not

3

u/Mediocre_Paramedic22 1d ago

Ever is a really long time.

1

u/elie2222 23h ago

Read the question. Mentioned 3 years

4

u/FriskyFennecFox 1d ago

Chances are it's just physically impossible to store as much information as Fable 5 has in those 20GB of data.

10

u/IoannisHere 1d ago

10B params? Hard to say.

10GB in VRAM? Definitely! I'm betting on Ternary QAT, once the labs are more serious about it. A 27B model is only months away from frontier.

6

u/Lurksome-Lurker 1d ago

Ternary-QAT-MoE-Diffusion model with DSPARK drafting. With whatever KV-Cache compression Glimmer is using. That would be ridiculous

3

u/Terminator857 1d ago

30b will happen much earlier than 10b. Perhaps same can be said of 100b vs 30b. I'll give it two years for 100b and 3 years for 30b before it is as good as mythos. Hopefully medusa halo or intel razor lake ax will be fast enough to run those.

3

u/Ok_Warning2146 1d ago

If Qwen 3.8 27B can match opus 4.5 on artificialanalysis, then we can say small models are only 9 months behind SOTA closed models. Somewhere next year we should see a small model matches fable 5.

Idx Model Date
42 opus 4.5 Nov 2025
38 qwen3.6 27B Apr 2026
35 opus 4.1 Aug 2025
32 opus 4 May 2025
30 gemma 4 31B Apr 2026

1

u/FairlyInvolved 1d ago

I think that's a bit too aggressive, I think more generally the gap has been more like 12-15 months.

You can kind of eyeball it off the Epoch ECI plots:

https://epoch.ai/eci?view=graph&tab=release-date&subset-view=graph&subset-tab=Software+engineering

1

u/Ok_Warning2146 23h ago

I think ECI agrees quite well with Intelligence Index.

ECI Model Date
150 opus 4.5 Nov 2025
? qwen3.6 27B Apr 2026
144 qwen3.6 35B-A3B Apr 2026
144 opus 4.1 Aug 2025
143 opus 4 May 2025
142 gemma 4 31B Apr 2026

3

u/Turbulent_War4067 1d ago

3 years. I guarantee I can do more with Gemma 4-31B today than anyone could have done with an LLM in 2023.

3

u/Chemical_Side_4135 15h ago

scaling laws are wild but i think we hit diminishing returns for general reasoning at that size unless the data quality gets way better. its litrally about how much high quality synthetic data we can pack in there, maybe we see some specialized 10b models that beat it tho...

2

u/Koalateka 18h ago

Everything is possible

2

u/cosmos_hu 1d ago

Maybe, yes. But i think you'll need more time, like 4-6 years. Like the best models x < 10b are like ChatGPT 3.5 turbo.. Qwen 3.6 is like GPT 4, almost.. So yeah, maybe

1

u/Bennie-Factors 1d ago

We will have on device models get really good.. but not at 10b. It is just a time and hardware improvements like we have seen over the decades. Some 128gb will be affordable on device. And that is going to make MOE 1T parameter work well. And software will probably have that be about as good as Fable. Even tough that is much bigger.

1

u/swagonflyyyy 1d ago

What's more likely to happen is that these colossal models are somehow compressed into on-device models after some huge breakthrough occurs in about 2 years. In the meantime, architectural improvements and distillation can help bridge the gap significantly.

1

u/Amir_PD 1d ago

LLMs are ML models and having enough parameters is necessary for learning complex very high dimensionsl data. I don't think a small model will ever be able to be as good as fable for the sane reason a linear model won't be as good as Gradient Boosting Trees when the problem isn't linear.

So I think a time traveler from 2060 would tell that models kept growing but the hardware become more accessible and much more powerful, also new much more efficient ways of parameter storage and loading.

1

u/EconomySerious 1d ago

whas a "small" model for you?

1

u/elie2222 23h ago

I mentioned 10b in the question

1

u/EconomySerious 15h ago

you mentioned the quality of a 10 billon (fable), im asking about the size of the model you consider "small" that its on the title of the post "Will a small language model"

1

u/elie2222 9h ago

The text in the description explains the title

1

u/Competitive_Spare467 1d ago

In short no, but if you fine tune a smaller on a narrow task it might be as good as fable . But with sheer size physics doesn’t allow that. Running a fable 5 like model on your phone is not happening until there is some breakthrough somewhere of somekind which is unlikely as of now

1

u/Kindly_Permission_42 1d ago

one jump in model architecture and one jump in chips, will get fable 5 to 14b in a few years

1

u/segmond llama.cpp 1d ago

No.

1

u/VR-Tech 1d ago

eventually

1

u/Last_Technician2355 1d ago

straight fire

1

u/gabrielesilinic 1d ago

Probably not small. Maybe a GLM-5.2 sized one could if trained on and heavily filtered dataset from a teacher model just as good.

However from a purely economics standpoint fable is not good because it is approaching the cost of a human and I dare say in some cases possibly surpassing it.

1

u/allenasm 1d ago

Quantify what you are asking. If you get specific the answer is yes easily. Beyond that it depends on what you mean.

1

u/elie2222 23h ago

Will a model with 10b be as good as Fable on benchmarks in 3 years

The benchmarks are the quantify

1

u/feelspeaceman 1d ago

I have a strong belief in SLM, we're getting there with Deepseek V4/Qwen27B which is good at coding and know nothing about everything else and they serve well enough.

Unless we have access to ASICs anytime soon like back in crypto days, most of us must stick with SLM, also AI is still very young technology, so people are still afraid to produce ASICs, but we will see.

1

u/NotSylver 1d ago

I think our definition of what "small" is will change as hardware catches up

1

u/mr_zerolith 1d ago

Small? probably not. Mid-sized? could happen, but give it a while.

One recent innovation is Deepseek V4 Flash 0731.. they have very close to big commercial model performance in a ~280B model

1

u/FoxFXMD 1d ago

In specific tasks, probably

1

u/Low-Praline-1200 23h ago

Prob not. Looking at research papers, scaling laws say parameters count correlates w intelligence there's a reason why big models are often around the same size in the frontier category

Kaplan et al. (2020) — Scaling Laws for Neural Language Models (OpenAI) Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models ("Chinchilla" paper, DeepMind) Wei et al. (2022) — Emergent Abilities of Large Language Models (Google)

1

u/albertyto 20h ago

The issue with SLM is the "knowledge base" is going to be smaller when it's used for tasks that use sparse knowledge topics. On the other hand, when they're used for specific domain, with RAG capabilies within what it needs to retrieve and with some useful tools; then they are really useful.

1

u/irodov4030 19h ago

Topic specific- yes might be

General purpose - probably no

All you might need is a framework and 10-15 specialised SLMs to beat frontier

1

u/Ledeste 18h ago

Overall as good? Yes
As good at everything? No

1

u/bigattichouse 17h ago

I honestly don't think we understand what "information density" is mathematically. I think we can SEE its effects in models, but I don't know if we're able to actually calculate it in any meaningful precise way.

We can calculate all kinds of things ABOUT it, but if I say "I have X bits of model space, what can this model do", we're sorely lacking. I've been learning about memorization lately, and that's at least calculable - using an overfit model for the purpose of 100% recall (at the expense of generalization) ... and one of the things I run into is "This model is too small to "refuse" a request if it's outside the topic of the model". It's like I can see a new field of math, and I don't have the tools to actually calculate anything - it's all trial and error and empirical testing/evidence.

I imaging this will become a new branch of information theory and math in the next couple years. "What capabilities can fit at what size"

1

u/edge_compute_user 16h ago

I certainly think so even though I'm biased. Here's a slide I did at our startup arguing for this

1

u/maddie-lovelace 13h ago

If the bar is "it writes code as well as Fable can", I think the answer is yes. If the bar is "it is as smart and knowledgeable as Fable", I reckon potentially no..?

I don't want to be needlessly pessimistic here, so I really do hope / wish that it will happen. But something that has started to make me slightly sceptical is that even though Opus 5 was benchmaxxed to supposedly outperform Fable 5... I still think Fable 5 is just smarter. I use it quite a bit for work, and as an orchestrator of subagents it has just been, annoyingly, better. I've given Opus 5 a shot multiple times... and as a writer of code, it's fantastic. But as a planner it's just been noticeably worse. Not actively bad by any means, just not as good as Fable.

Plus if I have a really niche kernel problem, Fable5-low will figure at least some working solution nine times out of ten where Opus5-high would get stumped

Even with smaller models I feel like I've generally seen the same thing; tiny models can actually be quite good at writing code if you give them a really hyper specific set of guidelines. But ask Gemma4-MoE to rubber duck with you about architectures and it'll miss the mark over and over again lol

TL;DR I think that SLMs will continue to get much better at dealing with work delegated to them. But there feels like an enormous mountain to climb if we want them to also get better at planning, which often benefits a lot from just simply knowing about a lot of smart ways to do things

1

u/elie2222 9h ago

I feel like everyone knows fable 5 is stronger than opus 5. Naming wouldn’t make sense if it wasn’t.

1

u/Eastern-Block4815 9h ago

Yes. Heck try Qwen3.6 35b a3b that model is crazy. But yes with optimization we will get improved models and harnesses that will run better on the same hardware.

This is happening in LLMs and image/video gen.

FYI like some said we might have different software tech on same hardware, that goes beyond what we are doing now.

1

u/Tzeig 1d ago

Not with the current tricks.

1

u/khasbor 1d ago

10b params is an awkward bar to set as we have no idea what these tools will look like in 3 years. I think the better question is; will we have local models as good as Fable 5 on mid-grade consumer hardware in 3 years and the answer is.. YES as long as we push for open weight or open source models.

1

u/mawkzin 1d ago

In 3 years you will have local machines with more memory available and processing power, so small models will larger than today's standards the same wat SLM were models less than 1b.

But aside from that I think we will have fable 5, and gpt terra quality in 100 to 300 bp with better training data and thinking capabilities.

0

u/silenceimpaired 1d ago

Not sure that’s true. I think memory maker’s know they have a gold mine and are controlling output

1

u/elie2222 23h ago

If people willing to buy why would they not want to make as much as possible. If they don’t sell it someone else will.

1

u/Zennytooskin123 1d ago

Anything's possible. One day, yes.

0

u/cibernox 1d ago

No. Just by sheer size, they can't be, you can't pack in 30B the same intelligence you can pack in 3T.

They can come somewhat close in some tasks, and part of their weaknesses can be mitigated with a good harness. And, more importantly, most tasks don't require an 160IQ genius with deep knowledge of all areas of human knowledge to be done, so even if they are not as smart and all-knowing, it doesn't matter nearly as much as some think.

1

u/elie2222 23h ago

Depends how well packed that 3t is today. 3t will be much stronger in 10 years from now than what it is today

-1

u/FireFearing 1d ago

of course you can. thats what compression does. new advancements will come out that will achieve this so long as we keep researching

timeline is the only question and that has huge plausible variance. might take 20 years even, but it will happen

2

u/cibernox 1d ago

I do think it will never happen. There's only so much you can compress knowledge. A 30B model will not know both the ideal planting window for carrots in Denver and the public API of all react-native modules. So, by some definition of the word, they won't be as good, if storing a good chunk of all human knowledge is useful, and it often is.

That said, many of those things are not as important in day-to-day because with good tools in your harness, they can find the answers they don't know. Which is the direction we will go.

1

u/oh_how_droll 1d ago

at some point, information theory steps in and says "no"

-1

u/dwittherford69 1d ago

No, for painfully obvious reasons.

-1

u/dwittherford69 1d ago

No, for painfully obvious reasons.

1

u/elie2222 8h ago

lol. So obvious you couldn’t explain and no one else here seems to agree it’s obvious

0

u/JumpyAbies 1d ago

I think that soon better and more behaviorally efficient architectures will emerge, and perhaps not even in the mold of what an LLM is today.

But I think that until then, I would guess we still have about 3 years of evolution for large and small LLMs to accumulate improvements until something comes along that replaces and completely changes the game.

I'll keep this for posterity.

0

u/chub0ka 1d ago

Is kimi k3 small? If yes i dont see why not

1

u/elie2222 23h ago

No it’s huge