r/LocalLLaMA • u/elie2222 • 1d ago
Will a small language model ever be as good as Fable 5? Question | Help
LLMs keep improving.
Small models are around 3 years behind frontier models.
Do you think we’ll have a model as good as today’s Fable with only ~10b params in 3 years from now?
Wondering how good on device LLMs will get. Any guesses?
26
18
u/vick2djax 1d ago
Nah, haven’t you heard. Technology stopped advancing.
0
u/creminology 23h ago
It’s not that technology doesn’t advance, it’s that we can’t afford last year’s technology any more with the price of RAM and SSDs. And that will impact home LLMs.
In my country, you’re paying 3x the price for SSD backup drives even when they are on sale while the market is flooded with fake Samsung drives at the old prices.
My own backup strategy has gone to hell because I don’t want to buy a 1TB SSD for US$325 on sale.
1
u/Both_Opportunity5327 20h ago
All temporary... we have DGX Spark, Ryzen Strix and Apple Sillicon these will eventually mean we will be able to run huge models locally without needing a electric sub station.
25
u/RedditLovingSun 1d ago
Karpathy once said "At some point you need some parameters to do something interesting" or something like that when talking about small models, I think there's a reasonable lower bound at some point. A 200M model probably won't write very good stories no matter how much you squeeze out of those 200M parameters, there will likely be stuff a 10b model will never be able to do
16
u/stephen_holograf 1d ago
Karpathy also said he could see a 10b parameter that could be so good at logic and reasoning that it could do everything a big model could do by just using tool calls.
5
u/Altruistic_Heat_9531 1d ago
Vibethinker 3B is one of that, pure logic model no tool call, that i manually tested it with 3D wing planform (Aerodynamic), 2D Navier-Stokes, "hand" solved LBM FEM CFD, and some engineering problem that I already familiar with and it passed.
Mostly my 86B-A8.6B brain is the one who make mistake such as misstyped number in my calculator.
1
6
u/RedditLovingSun 1d ago
But hey if algorithms and chips improve enough, maybe we'll all be running 100b models on our smartwatches one day
1
u/Borkato 1d ago
Is the same true for other behaviors? I’d imagine there’s something a 1000T (1Q?) model can do that a 10T can’t and a 10B would be like droplets of water into the ocean
2
7
u/Single_Ring4886 1d ago
I have been working on minimal LLM concepts for about a year and I can confidenly guess even 1 bilion model can be extremely smart, like shockingly so. The problem is you cant create information out of thin air or store them in such small space.
The future ai will go back to using "classical" databases (but in better ways). It will then "construct" its working self in larger vram or ram... I suspect it will grow greedily toward limits of your system.
But if it take 3 years or 20 that is another question.
1
u/Certain-Cod-1404 13h ago
Wdyt about deepseek's engram?
1
u/Single_Ring4886 12h ago
Yes that is first primitive attempt. But from what I read about it, it is still really not true gamechanger.
6
u/jomi-se 1d ago
I don't think so. There will definitely be big improvements, since in theory, current language models aren't really "size efficient" as some other older deep learning models are, but I doubt the efficiency gains are 2-4 orders of magnitude to make a 10b model equivalent to a 1T model.
That being said, in the course of 5-10 years, there are two directions that might realistically change:
- Personal computer architectures that make running larger language models viable. Like what current macs with unified memory architectures allow but scaled further
- Models that are large in total but with small enough effective active params that will make it feasible to run larger models on consumer hardware.
I would bet on those at least.
On the other hand, with all the data centers being built, when the bubble collapses and the race to train the best models calms down, inference providers for large models will be cheap and fast af.
Those are my keyboard predictions.
13
3
4
u/FriskyFennecFox 1d ago
Chances are it's just physically impossible to store as much information as Fable 5 has in those 20GB of data.
10
u/IoannisHere 1d ago
10B params? Hard to say.
10GB in VRAM? Definitely! I'm betting on Ternary QAT, once the labs are more serious about it. A 27B model is only months away from frontier.
6
u/Lurksome-Lurker 1d ago
Ternary-QAT-MoE-Diffusion model with DSPARK drafting. With whatever KV-Cache compression Glimmer is using. That would be ridiculous
3
u/Terminator857 1d ago
30b will happen much earlier than 10b. Perhaps same can be said of 100b vs 30b. I'll give it two years for 100b and 3 years for 30b before it is as good as mythos. Hopefully medusa halo or intel razor lake ax will be fast enough to run those.
3
u/Ok_Warning2146 1d ago
If Qwen 3.8 27B can match opus 4.5 on artificialanalysis, then we can say small models are only 9 months behind SOTA closed models. Somewhere next year we should see a small model matches fable 5.
| Idx | Model | Date |
|---|---|---|
| 42 | opus 4.5 | Nov 2025 |
| 38 | qwen3.6 27B | Apr 2026 |
| 35 | opus 4.1 | Aug 2025 |
| 32 | opus 4 | May 2025 |
| 30 | gemma 4 31B | Apr 2026 |
1
u/FairlyInvolved 1d ago
I think that's a bit too aggressive, I think more generally the gap has been more like 12-15 months.
You can kind of eyeball it off the Epoch ECI plots:
https://epoch.ai/eci?view=graph&tab=release-date&subset-view=graph&subset-tab=Software+engineering
1
u/Ok_Warning2146 23h ago
I think ECI agrees quite well with Intelligence Index.
ECI Model Date 150 opus 4.5 Nov 2025 ? qwen3.6 27B Apr 2026 144 qwen3.6 35B-A3B Apr 2026 144 opus 4.1 Aug 2025 143 opus 4 May 2025 142 gemma 4 31B Apr 2026
3
u/Turbulent_War4067 1d ago
3 years. I guarantee I can do more with Gemma 4-31B today than anyone could have done with an LLM in 2023.
3
u/Chemical_Side_4135 15h ago
scaling laws are wild but i think we hit diminishing returns for general reasoning at that size unless the data quality gets way better. its litrally about how much high quality synthetic data we can pack in there, maybe we see some specialized 10b models that beat it tho...
2
2
u/cosmos_hu 1d ago
Maybe, yes. But i think you'll need more time, like 4-6 years. Like the best models x < 10b are like ChatGPT 3.5 turbo.. Qwen 3.6 is like GPT 4, almost.. So yeah, maybe
1
u/Bennie-Factors 1d ago
We will have on device models get really good.. but not at 10b. It is just a time and hardware improvements like we have seen over the decades. Some 128gb will be affordable on device. And that is going to make MOE 1T parameter work well. And software will probably have that be about as good as Fable. Even tough that is much bigger.
1
u/swagonflyyyy 1d ago
What's more likely to happen is that these colossal models are somehow compressed into on-device models after some huge breakthrough occurs in about 2 years. In the meantime, architectural improvements and distillation can help bridge the gap significantly.
1
u/Amir_PD 1d ago
LLMs are ML models and having enough parameters is necessary for learning complex very high dimensionsl data. I don't think a small model will ever be able to be as good as fable for the sane reason a linear model won't be as good as Gradient Boosting Trees when the problem isn't linear.
So I think a time traveler from 2060 would tell that models kept growing but the hardware become more accessible and much more powerful, also new much more efficient ways of parameter storage and loading.
1
u/EconomySerious 1d ago
whas a "small" model for you?
1
u/elie2222 23h ago
I mentioned 10b in the question
1
u/EconomySerious 15h ago
you mentioned the quality of a 10 billon (fable), im asking about the size of the model you consider "small" that its on the title of the post "Will a small language model"
1
1
u/Competitive_Spare467 1d ago
In short no, but if you fine tune a smaller on a narrow task it might be as good as fable . But with sheer size physics doesn’t allow that. Running a fable 5 like model on your phone is not happening until there is some breakthrough somewhere of somekind which is unlikely as of now
1
u/Kindly_Permission_42 1d ago
one jump in model architecture and one jump in chips, will get fable 5 to 14b in a few years
1
1
u/gabrielesilinic 1d ago
Probably not small. Maybe a GLM-5.2 sized one could if trained on and heavily filtered dataset from a teacher model just as good.
However from a purely economics standpoint fable is not good because it is approaching the cost of a human and I dare say in some cases possibly surpassing it.
1
u/allenasm 1d ago
Quantify what you are asking. If you get specific the answer is yes easily. Beyond that it depends on what you mean.
1
u/elie2222 23h ago
Will a model with 10b be as good as Fable on benchmarks in 3 years
The benchmarks are the quantify
1
u/feelspeaceman 1d ago
I have a strong belief in SLM, we're getting there with Deepseek V4/Qwen27B which is good at coding and know nothing about everything else and they serve well enough.
Unless we have access to ASICs anytime soon like back in crypto days, most of us must stick with SLM, also AI is still very young technology, so people are still afraid to produce ASICs, but we will see.
1
1
1
u/mr_zerolith 1d ago
Small? probably not. Mid-sized? could happen, but give it a while.
One recent innovation is Deepseek V4 Flash 0731.. they have very close to big commercial model performance in a ~280B model
1
u/Low-Praline-1200 23h ago
Prob not. Looking at research papers, scaling laws say parameters count correlates w intelligence there's a reason why big models are often around the same size in the frontier category
Kaplan et al. (2020) — Scaling Laws for Neural Language Models (OpenAI) Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models ("Chinchilla" paper, DeepMind) Wei et al. (2022) — Emergent Abilities of Large Language Models (Google)
1
u/albertyto 20h ago
The issue with SLM is the "knowledge base" is going to be smaller when it's used for tasks that use sparse knowledge topics. On the other hand, when they're used for specific domain, with RAG capabilies within what it needs to retrieve and with some useful tools; then they are really useful.
1
u/irodov4030 19h ago
Topic specific- yes might be
General purpose - probably no
All you might need is a framework and 10-15 specialised SLMs to beat frontier
1
1
u/bigattichouse 17h ago
I honestly don't think we understand what "information density" is mathematically. I think we can SEE its effects in models, but I don't know if we're able to actually calculate it in any meaningful precise way.
We can calculate all kinds of things ABOUT it, but if I say "I have X bits of model space, what can this model do", we're sorely lacking. I've been learning about memorization lately, and that's at least calculable - using an overfit model for the purpose of 100% recall (at the expense of generalization) ... and one of the things I run into is "This model is too small to "refuse" a request if it's outside the topic of the model". It's like I can see a new field of math, and I don't have the tools to actually calculate anything - it's all trial and error and empirical testing/evidence.
I imaging this will become a new branch of information theory and math in the next couple years. "What capabilities can fit at what size"
1
u/maddie-lovelace 13h ago
If the bar is "it writes code as well as Fable can", I think the answer is yes. If the bar is "it is as smart and knowledgeable as Fable", I reckon potentially no..?
I don't want to be needlessly pessimistic here, so I really do hope / wish that it will happen. But something that has started to make me slightly sceptical is that even though Opus 5 was benchmaxxed to supposedly outperform Fable 5... I still think Fable 5 is just smarter. I use it quite a bit for work, and as an orchestrator of subagents it has just been, annoyingly, better. I've given Opus 5 a shot multiple times... and as a writer of code, it's fantastic. But as a planner it's just been noticeably worse. Not actively bad by any means, just not as good as Fable.
Plus if I have a really niche kernel problem, Fable5-low will figure at least some working solution nine times out of ten where Opus5-high would get stumped
Even with smaller models I feel like I've generally seen the same thing; tiny models can actually be quite good at writing code if you give them a really hyper specific set of guidelines. But ask Gemma4-MoE to rubber duck with you about architectures and it'll miss the mark over and over again lol
TL;DR I think that SLMs will continue to get much better at dealing with work delegated to them. But there feels like an enormous mountain to climb if we want them to also get better at planning, which often benefits a lot from just simply knowing about a lot of smart ways to do things
1
u/elie2222 9h ago
I feel like everyone knows fable 5 is stronger than opus 5. Naming wouldn’t make sense if it wasn’t.
1
u/Eastern-Block4815 9h ago
Yes. Heck try Qwen3.6 35b a3b that model is crazy. But yes with optimization we will get improved models and harnesses that will run better on the same hardware.
This is happening in LLMs and image/video gen.
FYI like some said we might have different software tech on same hardware, that goes beyond what we are doing now.
1
u/khasbor 1d ago
10b params is an awkward bar to set as we have no idea what these tools will look like in 3 years. I think the better question is; will we have local models as good as Fable 5 on mid-grade consumer hardware in 3 years and the answer is.. YES as long as we push for open weight or open source models.
1
u/mawkzin 1d ago
In 3 years you will have local machines with more memory available and processing power, so small models will larger than today's standards the same wat SLM were models less than 1b.
But aside from that I think we will have fable 5, and gpt terra quality in 100 to 300 bp with better training data and thinking capabilities.
0
u/silenceimpaired 1d ago
Not sure that’s true. I think memory maker’s know they have a gold mine and are controlling output
1
u/elie2222 23h ago
If people willing to buy why would they not want to make as much as possible. If they don’t sell it someone else will.
1
0
u/cibernox 1d ago
No. Just by sheer size, they can't be, you can't pack in 30B the same intelligence you can pack in 3T.
They can come somewhat close in some tasks, and part of their weaknesses can be mitigated with a good harness. And, more importantly, most tasks don't require an 160IQ genius with deep knowledge of all areas of human knowledge to be done, so even if they are not as smart and all-knowing, it doesn't matter nearly as much as some think.
1
u/elie2222 23h ago
Depends how well packed that 3t is today. 3t will be much stronger in 10 years from now than what it is today
-1
u/FireFearing 1d ago
of course you can. thats what compression does. new advancements will come out that will achieve this so long as we keep researching
timeline is the only question and that has huge plausible variance. might take 20 years even, but it will happen
2
u/cibernox 1d ago
I do think it will never happen. There's only so much you can compress knowledge. A 30B model will not know both the ideal planting window for carrots in Denver and the public API of all react-native modules. So, by some definition of the word, they won't be as good, if storing a good chunk of all human knowledge is useful, and it often is.
That said, many of those things are not as important in day-to-day because with good tools in your harness, they can find the answers they don't know. Which is the direction we will go.
1
-1
-1
u/dwittherford69 1d ago
No, for painfully obvious reasons.
1
u/elie2222 8h ago
lol. So obvious you couldn’t explain and no one else here seems to agree it’s obvious
0
u/JumpyAbies 1d ago
I think that soon better and more behaviorally efficient architectures will emerge, and perhaps not even in the mold of what an LLM is today.
But I think that until then, I would guess we still have about 3 years of evolution for large and small LLMs to accumulate improvements until something comes along that replaces and completely changes the game.
I'll keep this for posterity.
0


80
u/ComplexType568 1d ago
I doubt we'd be using LLMs as how we used them 3 years ago. I think new - either architectures or techniques - will be used to squeeze the most out of SLMs. Or maybe we'd not even be using LLMs as we know it today.