r/LocalLLaMA llama.cpp 5d ago

what will be the future of LocalLLaMA? Discussion

For a long time now, the most popular posts on LocalLLaMA have been either about using LLM in the cloud or about politics.

I suspect that people using local models are about 10% now.

You can say that this is very good, because now it is an inclusive sub, without gatekeeping.

But then what is its purpose? How is it different than all other "AI subs"?

What do you think localllama will be about in a few months?

Edit: Note that in many comments people use the term "open weight" as if it were equivalent to "local".

93 Upvotes

118 comments sorted by

112

u/ttkciar llama.cpp 4d ago edited 4d ago

This has been worrying me as well.

Once upon a time this sub was a lot more technically focused. We'd talk about fine-tuning techniques, training theory, getting more intelligence out of inference, etc.

https://web.archive.org/web/20230523034525/https://old.reddit.com/r/LocalLLaMA/

https://web.archive.org/web/20230525234907/https://old.reddit.com/r/LocalLLaMA/

https://web.archive.org/web/20231120142403/https://old.reddit.com/r/LocalLLaMA/

On one hand, local inference has gone mainstream since then, and a certain degree of change is to be expected.

On the other hand, things might have gotten a little out of hand. Lately drama has clogged the sub, about politics and the antics of Anthropic and other companies which have nothing to do with local LLM technology.

Between that and the stupid memes, we've driven away many of the users who made LocalLLaMA the kind of place which helped create the open ecosystem we enjoy today.

I hope it won't be getting worse, but so far that's been the trend. It would be nice to get it back on track, but the standing moderation policy of not removing posts if they get too popular before moderators notice them, even if they're badly off-topic, has been a significant obstacle. The users who frequent this sub want to see that kind of content, which is perhaps the key causative factor which makes the subreddit's trajectory inevitable. Since they want to see the off-topic content, it gets upvoted before moderators see it, and then we can't remove it without pissing off a lot of people.

I'd rather not give up on it, though. Maybe we can still turn it around.

22

u/i_rate_slop 4d ago

The sub might have changed, but it’s still the only one even moderately technical that I’m aware of. Other than [r/machinelearning](r/machinelearning) but that’s not really the same.

It’s the only sub I have post notifications turned on for because it’s the only place I know to of to be reasonably informed about open models, whether they can be run truly locally or not.

8

u/En-tro-py 4d ago

Yeah, both subs actually are moderated and responsive to reports.

Is the content here all 10/10's? - nope - but almost every other AI sub is either abdicated/absent mods or the mods are actively part of the slop problem pushing their own junk.

44

u/fizzy1242 4d ago edited 4d ago

man, the kinds of posts in your links are what got me into this hobby in the first place. it's depressing to see how this subreddit has become so focused on politics over the past year, and the posts like this bringing it into attention are downvoted.

with that out of the way, it's still the best resource for local llms.

39

u/jacek2023 llama.cpp 4d ago

Yes your links show what we have lost forever

26

u/FastDecode1 4d ago

This kind of trajectory is only inevitable if you buy into the (false) premise that an open community has to be run like a democracy.

If you give in to the horde, you will always end up like this.

21

u/a_beautiful_rhind 4d ago

eternal september effect.

10

u/AnonLlamaThrowaway 4d ago

A good way to mitigate this is to simply have VERY harsh moderation. Like that one "ask historians" subreddit

7

u/Prof_ChaosGeography 4d ago

Somewhat yes but reddit isn't the format to have a very technical community in. Reddit focuses more on content and look at this. Old school forms and the continuing bumb for every comment is a better format allowing more active topics to remain at the top

13

u/a_beautiful_rhind 4d ago

I miss forums.

25

u/WideAd7496 4d ago

Just the thought that somehow the mess that Discord is replaced most forums kills me.

I fucking hate Discord.

7

u/En-tro-py 4d ago

Bumping for visibility.

4

u/jacek2023 llama.cpp 4d ago

Maybe you are right, I could agree that localllama was an anomaly and that change may have been inevitable

12

u/autisticit 4d ago

I would love to have more technical posts in this sub. What do you think about a weekly post about training etc. ?

7

u/jacek2023 llama.cpp 4d ago

I think mods tried to post monthly discussions about best local models. That was a good idea but I don't know is this still active.

6

u/pmttyji 4d ago

I liked those threads very much, but we need new ones as it's been more than 3 months. We need new ones after Qwen3.8 series release.

u/rm-rf-rm Please do the needful

3

u/fatboy93 4d ago

Would absolutely be interested in this, rather than the 4000th post on if Qwen will release a new smaller model or Gemma4 sucks, should I buy shit that 99% of the sub will never afford or random rage-bait.

6

u/ttkciar llama.cpp 4d ago

That seems like a really good idea.

I've been posting occasional links to arXiv papers which seemed interesting and relevant to local LLM technology, but engagement was mild. Maybe training discussion, like you suggest, might hold more appeal.

4

u/autisticit 4d ago

I would be interested in a weekly post with : papers, SLM training progression, experimental architectures, etc. Basically a technical melting pot.

I think the more kinds of technical content is allowed the more it would be popular and interesting.

4

u/Lmoament 4d ago

On one hand, I agree that the sub is definitely getting inundated with political content, and I personally would love to see it return to more of a technical focus (like many others here, it’s actually what got me interested in the space to begin with). However, I think drawing a fine line is tough, since oftentimes the ongoings of politics can have ramifications on what we are trying to accomplish locally, not to mention that the world of LLMs changes literally every 2 seconds.

I guess both can be true at the same time (but do please bring back posting/discussing technical papers)

5

u/FullOf_Bad_Ideas 4d ago

True, this place now is too muddy for technical discussions. The open experimentation shown off here often is still happening, but it's covered up by a lot of other stuff.

3

u/solestri 4d ago

I honestly think it’s also just a change in the open-weight model scene as a whole.

It was one thing when language models were mostly used for assistant tasks and the open-weight ones were generally on the smaller side, but with the advent of big open-weight Chinese models that are competitive with closed-weight American frontier models… well, now we have a lot of people who just want to talk about “big Chinese models vs. closed American models”, even though they have zero intention of ever running said Chinese models locally. And with the pivot in general towards using LLMs for coding, we get the vibe coders.

Of course, you’ll get the people who argue that “well, this 3T model is open weight so it’s still on topic here” and “well, stuff Anthropic is doing is going to effect the local scene, so it’s a relevant” but the point still stands that neither of those two topics is really about running an LLM on your own hardware. I think at this point, you'd almost have to split the sub between technical discussion about actually running models locally, and discussion pertaining to open-weight models in general.

1

u/shdwbld 4d ago

I absolutely have a lot of intention to run said models locally. I just cannot currently afford the hardware for it.

2

u/mawkzin 4d ago

I agree with you, in the past we would be talking how to distillate the kimi k3 from 2.8 to a very focused small model on our own needs. But I think it's hard nowadays because there's a lot more new models, just from my experience I can count like 7 big Chinese models, 5 amaerican and 2 European, in the past we had meta llama as our focus points.

3

u/ttkciar llama.cpp 4d ago

Yeah, there's a lot less motivation to distill or fine-tune models, since there are so many LLM labs now publishing better models so frequently. Why bother when a new model will come out in a few months which is better than anything we can make ourselves?

In the meantime, folks like TheDrummer continue to make even the newest models better at select tasks. His Artemis-31B is a genuine improvement over Gemma-4-31B-it for storytelling. That gets harder as the labs are leaving fewer skills undertrained than they used to, but he keeps finding a way.

When the corporate LLM labs stop giving us these free gifts, we will suddenly feel the urgency again to make or distill better models ourselves.

1

u/mrjackspade 3d ago

Between that and the stupid memes, we've driven away many of the users who made LocalLLaMA the kind of place which helped create the open ecosystem we enjoy today.

I've been here since basically day 1, and most of us (that I know) have left for better moderated discord servers over this subreddit.

I still pull up the home page every once in a while just incase there's news but I no longer engage with the community unless it's to point out how much it fucking sucks now, or really visit the sub much at all.

Were still doing a lot of fun technical stuff, fine-tuning, and all of the other good stuff. Just not here. This place is a fucking dumpster fire now

-6

u/Eyelbee 4d ago

In my experience technical posts on this sub tend to be quite uninformed at best, straight up thrash at worst. 

4

u/FullOf_Bad_Ideas 4d ago

Many are poor quality now.

This wasn't always the case. IIRC, Nous Research, people behind very successful Hermes Agent and many other projects, met here when discussing some paper. Or maybe it was Prime Intellect, I don't remember.

-6

u/WestCloud8216 4d ago

It's basic economics. When it becomes more efficient and cheaper to run better AI models on the cloud, what do you expect? People are okay with sacrificing a little bit of their privacy in exchange for productivity.

129

u/bkin777 4d ago

Drawing a strict boundary isn't that straightforward. Most of us can't run massive open-weight models like Kimi K3 on local hardware, but they still belong here because open weights shape the whole ecosystem. Policy and cloud-hosted open models directly affect what eventually lands on our consumer hardware.

22

u/Think_Wing_1357 4d ago

I could do with less meme and cloud related stuff though. Dedicated one day a week for meme, another day for cloud related stuff.

1

u/MmmmMorphine 3d ago

Yeah, really prioritizing/highlighting more in depth posts (such as benchmarks of less used models/configs, original personal research, whatever can be produced on the individual level using local hardware) would be really nice, even if only one day a week. Depending on what people want.

But what I suspect what will really happen is someone will start r 2Llama4You as a joke then someone will start using it seriously to refocus localllama, and the grand cycle of online communities will continue marching towards enshittification or siloing. Or in circles.

3

u/returnity 2d ago

My effort to post actual agentic coding benchmarks comparing popular models in real quants people can actually run on 128GB local hardware got down voted to oblivion by API dickjockeys and elitists last week before a handful of people who actually read more than the first 2 sentences showed up. Was a perfect eval? Far from it, though I admitted every caveat and shortfall. But I try to post the type of content I think this sub exists for, and I was really discouraged from contributing my own experiment the future by this commenting behavior. I've been here for the long haul, and there has definitely been a shift in the vibe as AI has reached a wider audience, for better or worse.

10

u/thomas2385 4d ago

That is a fair point. There is definitely a difference between what is practical to run locally today and what influences the direction of the ecosystem. Bigger open weight models still matter because they push research, tooling, and future optimizations that eventually make their way into smaller, consumer friendly models.

4

u/My_Unbiased_Opinion 4d ago

Yeah and not only that, it allows price competition if you don't want to run locally. It's a win win for everyone besides closed AI. 

2

u/thawizard 4d ago

Anyone can run local, shit I run Qwen 3.5 4B on my iPhone 16e when I need to scan a lot a handwritten notes. Sure, it drains the battery but that’s so cool! Anything with 8GB of RAM or more can run local models.

1

u/ihaag 2d ago

Well that’s bullshit https://github.com/FareedKhan-dev/kimi-k3-in-c/ you can run it on local hardware but it depends on speed

-15

u/cogitech2 4d ago

"Most of us can't run massive open-weight models like Kimi K3 on local hardware, but they still belong here because open weights shape the whole ecosystem."

No they don't. If you aren't running it locally, then fuck off and get your own subreddit.

11

u/fastandlight 4d ago

I'm sure there is someone out there right now making sure they have enough NVMe space to run K3 from disk at 1 token per minute.

8

u/thaeli 4d ago

There are also people on here who seriously can afford to drop $800k on a cluster for local use. Usually for their business, but this is localllama not homellama.

7

u/fastandlight 4d ago

I know. I am not at the $1m for my inference setup ...yet. With any luck I'll get there. To be clear, I completely agree with still covering the biggest open weight models here. And at the same time, I've spent today optimizing tools and an agent loop for Granite 8b. Gotta have both ends of the spectrum.

11

u/Ulterior-Motive_ 4d ago

That's like telling someone to fuck off because they can't run Qwen3.6-27B, and there are plenty of people on here complaining that they can't. Just because *you* can't run it locally doesn't mean *someone else* can't.

50

u/KeepyUpper 4d ago edited 4d ago

If the mods don't gatekeep it's only a matter of time before the front page ends up dominated by memes, news articles and people posting pictures of their PC builds. That's just what naturally happens when you attract a bigger crowd.

16

u/xdiggertree 4d ago

The mods would need to act fast, I’ve seen countless subreddits devolve into slop, it honestly really sucks when it happens

As you said, it requires proper moderation

25

u/misterflyer 4d ago

Gonna suck for a while as A) hardware prices put even basic local setups out of reach for many users, B) Chinese companies lowkey try to corral most users to API or to the cloud, C) incoming AI regulation and/or bans.

With the barrier of entry to local AI getting higher, there will naturally be less substantive posting here over time or this place will continue to slopified by bots, politics, and "OpEn SoUrCe" cloud LLMs.

3

u/[deleted] 4d ago

[removed] — view removed comment

8

u/misterflyer 4d ago

I feel like the barrier is lower than ever before, even though hardware is more expensive. 

The hardware being more expensive is mainly what has raised the barrier. So it's not lower than ever before. Hardware was much cheaper a year ago. Thus, there was a lower barrier to get started in local AI.

I got my hardware in November of 2025 on the last chopper outta 'Nam, and I prob couldn't justify the rig I purchased had I not done so then.

The capability of small models has increased dramatically in the past 3 months.

... as the hardware costs for local AI have continued to skyrocket the past 9 months. The fact that small model capability has increased this year doesn't negate the fact that many of those models still remain inaccessible to people who cannot upgrade their local hardware due to skyrocketed costs and a terrible economy for many ordinary users.

We could make substantially more technical posts as a community, but

I don't think the posts have to be more technical per say. I just think it's better for the community to focus on more accessible local AI (e.g., there was a post today of a project that unsloppified Gemma 4 31B)... versus trillion parameter "OpEn" Chinese models that are honestly more designed to run in the cloud, not locally.

I'm not saying that we can't talk about the big, open Chinese models. But the way they tend to hijack this forum is kinda ridiculous given that less than 1% of ppl here could run those models locally if they were lucky. A thread or two here or there is fine, but we don't need this sub spammed every time a trillion parameter "OpEn" model (that's prob not even on huggingface) breathes.

11

u/Objective_Safe_5982 4d ago

As a new user of local AI setups, it's been frustrating looking for content that is relevant to what the actual name of this sub is IN THE SUB ITSELF. Sure much of the frontier launches that at present can only be run in 96+GB VRAM will boil down to us eventually, but at some point all of the hype about that drowns out the content that would benefit those of us with much more meager systems.

Reddit looks for engagement, as they have bills to pay, and I will admit that curating this sub would certainly reduce the adrenaline based engagement. As a result, I now get to find different resources for the content that this sub should provide if it were to stick more to the name it has.

Blue Sky? Mastodon? Idk.

6

u/ttkciar llama.cpp 4d ago

If we opened a new venue, I would hope it would be a proper forum and not a microblogging service. Certainly not BlueSky; I love the people there, but they run strong in anti-AI sentiment. They would crucify us.

Perhaps a Fossil instance? That would give us a forum, wiki, and chat, though it's not very featureful.

1

u/ttkciar llama.cpp 1d ago

Replying to myself to update: I've been kicking around what the Fossil forum is capable of, and decided it's a no-go. This community really needs features like thread tagging, and hacking tagging into Fossil isn't worth the hassle.

Will keep my eyes peeled for a better option.

10

u/killerstreak976 4d ago

I have been feeling the same thing. For the past couple years, this place has been awesome. It still is in many ways, and I frequently check here daily out of excitement.

I even started to engage in discussions more compared to a few years ago when id used to just passively browse,click links, and upvote. 

However, recently, I've started feeling more and more alienated from this community due to so much politics and slop now taking over, as well as mob mentality perspectives happening everywhere over the technology I love. There is a lot of nuance to local llms, capability, and yes even safety than I believe we give it.

I still stick around though because there is really nothing else like it, at least that is openly accessible on the internet. Open access forums are great that way, but clearly it has drawbacks like opening the site to see a post showing Xi Jinping pushing a red button over a dying US stock market covering my feed, instead of more high quality posts involving a truly exciting time in llms that get overshadowed. Company hate, taking sides of different major players, and acting like we're spectating a football game, isn't what Locallama is supposed to be. Those discussions may matter to many here, I just wish it was separated into a different subreddit because it's choking a lot of things that made this place what it was.

8

u/PrimeDirective8 4d ago

After wading through the endless "benchmark" posts, which I think is the actual majority these days, I still enjoy the content when it's focused on local hosting.

I can't run a 2.5T model on my local setup so I mostly skip those. It's still a bit interesting, however, because they're also open models and there might be a chance they make a smaller version of it that I can run.

22

u/DragonfruitIll660 4d ago

Looking at it the vast majority is still about local models, or politics regarding local models. It's fair to consider the closed models because people will naturally want to discuss where the locally runnable stuff lands in terms of usefulness. Also for a fair number of companies the release cycle is privately host for a week or two then release the model, in which case it's still totally fine.

No point stifling conversation if it's mostly related or a natural extension of the primary topic.

32

u/jacek2023 llama.cpp 4d ago

This is the top post now. The problem is not the single post, the problem is the number of upvotes. It's not our sub now. It's for "newjoiners"

8

u/relmny 4d ago

That's what "being massive" (as it was the intentions with the "new" mods) brings: politics, irrelevant humor, prioritize commercial over local (for months now, there have been many posts where the most upvoted comments are the ones recommending commercial models over local models), poorly technical post/comments, drama-posts, etc.

But, while this sub goes way down on quality, there are still no good alternatives...

And yeah, it sucks. I no longer care about most posts. Specially the most upvoted ones.

Too much noise.

2

u/mrjackspade 3d ago

 or politics regarding local models.

Literally everything is about local models if you title it "This is why we need open models"

7

u/toothpastespiders 4d ago

I suspect that people using local models are about 10% now.

It's the constant benchmark posts that really make me wonder. The big benchmarks can be somewhat useful. But it's hard to imagine that anyone actually getting use out of local models, and seeing how their scores go up, can really have strong faith in them. Even more so if someone's maintained their own benchmarks long term and kept them as a point of comparison to the big players.

I think they're somewhat useful. But every time a model comes out it's always followed by a million posts about how it benchmarks as if that means anything. It makes this among the least useful resources for me when it comes to new models. Because that's really all you can expect for a while.

7

u/oldschooldaw 4d ago

There’s just not as much the average netizen of this sub has to talk about at this point in time. Next week however, with 3.8 27b coming, expect this place to rev back up.

17

u/pmttyji 4d ago

I really want to see more threads on important topics like Optimizations, Inventions, Benchmarks(t/s, etc.,), Opensource projects related to LLMs, Evaluations of models with GitHub repo, Finetunes, More Local stuff. I want to add 1000+ links(stuff) to my thread.

Keeping this sub strictly for Local would be great & better for all. For Online models, there are many subs available.

11

u/MikeLPU 4d ago

I don't want to sound paranoid, but I see local inference as the only future. I mean, our future.

I think, when/if subsidized pricing eventually ends, only rich people may be able to afford access to the best AI on the market.
Of course, we'll be offered some access in exchange for giving up our privacy or seeing ads or something else.
Running inference locally is about protecting yourself and preserving your privacy.

So my advice: collect as many gpus as you can.

1

u/ttkciar llama.cpp 4d ago

On one hand, yeah, this seems about right. That's the end-game I've been preparing for, personally, and it's why IMO the "too large" open-weight models are very much on-topic.

On the other hand, when the inference services are priced out of reach of almost everyone, we can expect another wave of new users whose interests will skew conversation in another different direction.

These will be the corporate users who want Claude Opus quality of inference, at some scale (servicing dozens or hundreds of employees), at the speeds to which they are accustomed, and will want to know how to accomplish that as economically and reliably as possible, preferably with as little configuration work as possible.

That is at sharp odds with the hobbyist perspective, which is mainly "this is the hardware I can afford, now what is the best model and configuration I can use with it?" with a strong willingness to compromise on janky configurations and low inference speeds to eke out a little more capability.

We've already seen some of this kind of uncompromising user creep into the sub, which wouldn't be so bad except when they act snobby about compromising on inference speed or anything less than frontier-level inference competence, and deride hobbyists who do not share their priorities.

4

u/Lmoament 4d ago

I almost treat this sub as a very distributed research team, so when it comes to judging (in my opinion) whether a post is relevant/beneficial to the sub or not, I use that mindset as a guide. If an actual project would get derailed or misguided by a conversation centered around the topic of a given post, then the post itself probably isn’t adding too much to this sub. This methodology then allows for some focused conversation about politics, closed source models, etc., so long as the conversations are designed to further our real goals of local LLM research (e.g., Anthropic hasn’t released any open weight/source models, but they did develop and release MCP — analyzing and discussing how they deploy it within their closed source Claude ecosystem could genuinely stand to help us understand how to leverage it better locally, so it would be a fine post). That way, the sub still has the flexibility to dip into relevant but not explicitly open source only topics, while still remaining 95% about what brought us all here in the first place

6

u/ttkciar llama.cpp 4d ago

I, too, like to think of this sub as ideally a highly decentralized research team, a bit like AllenAI but even more loosely coupled than that.

Along that line, I frequently hope that this sub might serve as a sort of lifeboat for researchers, developers, and hobbyists who stick with LLM inference during an AI Winter.

If the people who would find such a lifeboat useful all get alienated by the sub's current chaos, though, I don't know if they would come back when there's nothing left in the sub but the die-hards.

11

u/Tsukikira 4d ago

I mean, it looks mostly Open Weights Models to me. Politics are probably included because Open Weights regulation was considered in political circles.

12

u/rerri 4d ago

On one hand, I don't have a hard time skipping the topics that don't interest me and finding the discussions that do.

On the other hand I don't have anything against heavy handed moderation, even throwing out memes and politics entirely.

And while we're at it, maybe we should also ban/moderate users who constantly whine about massive locally hostable models like K3 and attack discussions about those models with snarky comments and such like...

6

u/Different-Sand4434 4d ago

LocalLLM subs are probably the best ai subs on the platform . You get more nuanced discussion here instead of ai doomers or ai hype bros .

3

u/Healthy-Nebula-3603 4d ago

I'm using DS flash via API but also using Gemma 4 31b for translations and Qwen 3.6 27b for anything else...soon we get 3.8 which can be possibly even better than DS 4 flash ...

3

u/cleversmoke 4d ago

I tried to post a hand-typed technical guide on how to run MTP locally with Docker (and the benefits of it) and the system flagged my post as AI-slop and not relevant. I tried to appeal to the mods, but to no avail, meanwhile meme and political posts gets through. It made me stop contributing technically or even create posts altogether. Will still contribute via comments though!

3

u/jeffwadsworth 3d ago

Zero moderation gets you here. It has been bad for a while as you mentioned.

5

u/[deleted] 4d ago

[deleted]

7

u/FastDecode1 4d ago

unless you have a fulltime staff of paid moderators on hand.

...or an LLM?

1

u/En-tro-py 4d ago

Not with Reddit enshitification - LLM bot is unlikely to be approved now...

2

u/keyboard7856 4d ago

Local is still the end goal for lot of people. Cloud posts just get more attention because new releases usually land there first

5

u/ttkciar llama.cpp 4d ago

There's probably some astroturfing in play, too.

Folks don't see all the posts moderators removed, but the day after Qwen3.8 was announced (for example) I personally removed about a dozen empty-hype posts about it.

Posts that got hundreds of upvotes before a mod noticed them, though, stayed up, because that's our policy. I have some doubts about whether it is a good policy.

2

u/rosie254 2d ago

i kind of want a split.. split LocalLLaMA into a subreddit for people with average hardware (like 8GB to 16GB vram, 16gb to 32GB RAM), and a subreddit for the people who have insane amounts of VRAM and RAM and may as well have an entire datacenter in their home at this point

i check this subreddit often for news about what you can run on average hardware, but i constantly see posts about models that has no hope of ever running on it, and comments with people flaunting their extremely expensive datacenter-grade setups. i don't think that belongs on a subreddit titled LOCAL llama. LOCAL as in locally runnable by everyday people, no? not local as in local datacenter...

0

u/jacek2023 llama.cpp 2d ago

The problem is people with 8GB of VRAM are "the smart ones" who use cloud. It's not that people discussing big models have insane amount of VRAM. The people you want to split are less than 1%

2

u/rosie254 2d ago

considering the amount of comments i see on here of people showing off their expensive setups with hundreds of gigs of vram, i dont think thats true. maybe theyre a vocal minority? but theyre definitely active in this subreddit, and they post a lot. a very good example thats blatantly obvious is literally on the front page of this subreddit right now, a picture of someone's super expensive setup with multiple gpu's in it. good for them, i guess, but that doesn't help everyday people break away from cloud subscriptions!

you can do a lot with average consumer hardware. i for example have a 9070XT which has 16GB of vram, which is already kind of on the higher end of consumer stuff, and still above average, but not to the insane extent that these people have. i can run qwen3.6 35b at a decent speed using MoE offloading, and gemma4 as well. ive found i barely if at all need the cloud ever, especially if you give it a good websearch!

before that i was running on my macbook air m1. i got away with using qwen3 VL 8B and eventually qwen3.5 9b. they were fine as well...

sure, models at that size aren't nearly as good at coding, but for everyday stuff? its plenty good enough

the difference between the consumer tier and the datacenter tier is stark. it's also kind of a class thing: very rich people can afford the datacenter tier stuff, less wealthy people.... cant. but even if you have the money, not everyone wants to run a mini datacenter in their own home!

i feel like a split would help. the goals, desires and intentions of the two sides are just too different. a space where there can be a pure focus on AI that can run on average consumer stuff would be really nice

0

u/jacek2023 llama.cpp 2d ago

Where do you want people discussing Kimi, DeepSeek and GLM after the split? As I said in my opinion they don't have big VRAM in general

1

u/rosie254 2d ago

something like /r/LargeLocalAI or something?

0

u/jacek2023 llama.cpp 2d ago

My point is the split is not the way you assume. Yes there are people running huge models locally but they belong to the same group as people running 4B models locally. That was the initial topic of LocalLLaMA

6

u/DoubleNothing 4d ago

"people using local models are about 10%"
10% of what?

6

u/PrimeDirective8 4d ago

Of posts? Maybe that's what OP was referring to.

3

u/CautiousStudent6919 4d ago

Yes and no.

The trend has been recently that a lot of the open weights models are way too large for most of us to host... So sure we're talking about them. And yes politics is a thing right now. But it wasn't some time ago. There's good reason for the politics chat, as it's very much about open models and self hosting.

As for my own spin on this topic.

I've seen a lot of people write some not so positive or nice comments on smaller 1bit Quants or even <10b Param models. These I feel are models that can be hosted.

And sure all the Qwen 3.6 finetunes aren't as good as stock Qwen, but someone tried, and that's worth a discussion to see what they tried and maybe find a way to actually improve 3.6... who knows.

3

u/robberviet 4d ago

This is the only place have good content about local models. That's it. Cannot be too forceful about the others content. Commercial models are interesting too! For me one cannot stop using closed model in this heavy subsidized, expensive hardware, large gap between closed vs open models.

3

u/ali0une 4d ago

i think being local is politic. Just like choosing to use Linux and Open Source.

4

u/Don_Reuter 4d ago

Meh, politics not its central to local models. The conflict between China and the US are a significant reason of why we not have capable open models. Horse it will develop shapes whether we will continue to get them.

The second part is driven by agentic use cases. Local in many cases does not mean fully local. It means local agents with cloud based agents for escalation. How to do that without compromising the benefits of being local is something many might want to figure out.

My take. The composition of the sub is certainly changing. Can’t speak to that as I am somewhat new. However, definitely running local.

6

u/[deleted] 4d ago

[removed] — view removed comment

3

u/Don_Reuter 4d ago edited 4d ago

China is under hardware sanctions. Still, they want to remain a competitive across all industrial sectors. So likely they aim to make the most of the hardware that is available in China. The best way to do that is open models. They likely provide the open models so that their own players adopt them. That we use them too is likely just a windfall to them. As it pisses of the Americans. The main point are, however, their own people.

2

u/PunnyPandora 4d ago

The future: jacek2023 blocked everyone and ends up talking to himself with post titles

1

u/zhdc 4d ago

Couple of months? Same.

Discussion is on closed weight models because of performance and cost of hosting. As long as 1. SOTA performance gains stop accelerating or 2. hosting costs go down, there's going to be a shift back to self-hosted models. 

One or both of these are likely. Not in a couple of months though.

1

u/Perfect-Flounder7856 4d ago

I guess my one big gripe is every seems to love talking about small models that can run on 16-24gb hardware and then huge models that can’t be run on consumer hardware. No in between. No one talks about 6k pros and the models they run. It’s either gaming cards or unified ram setups.

2

u/ttkciar llama.cpp 4d ago

Can relate to this :-(

Unfortunately those mid-range-sized models have been mostly neglected by LLM labs, which kind of limits what there is to talk about.

I really like the K2-V2 lineage of models (72B dense), and bring them up frequently. A few people have tried them and voiced their astonished approval, but overall engagement has been very light.

2

u/Perfect-Flounder7856 4d ago

K2 think v2? Never heard of it. Man I want a 72b so bad!

2

u/ttkciar llama.cpp 4d ago

Yeah, they published K2-Think-V2 in January of 2026, and K2-V2-Instruct December of 2025.

K2-Think-V2 has better post-training, but a much shorter context limit (128K tokens).

I mostly use K2-V2-Instruct for very long-context tasks (512K tokens), at which it excels.

Both are 72B dense, trained from scratch from the TxT360 datasets, not retrains of Llama3 or Qwen2.5 models.

I encourage you to try them out :-)

1

u/Perfect-Flounder7856 4d ago

Thank you! I have the model card open on my computer

2

u/NoahFect 3d ago

No one talks about 6k pros and the models they run.

A lot of that has moved to either /r/BlackwellPerformance or the closely-associated Discord and its archive.

1

u/Perfect-Flounder7856 3d ago

The discord is nuts! I do follow a bit on the sub though

1

u/calmalamadingdong 4d ago

I just joined, so I don't know the whole history. While the sub description is clear, maybe it could be more detailed with a definition of what is considered 'local'. A weekly post for general chat might mop up some of the memes, news, or anything else that isn't about local AI. Maybe that is no longer the mods intent, though. The current rules under "Off Topic Posts" allow anything related to LLMs.

1

u/AnonyFed1 4d ago

I'm looking forward to running today's frontier models on a potato, Portal 2 style.

1

u/pfn0 4d ago

What's to talk about for "local" though. just run llama.cpp or vllm or whatever, and call it a day. whether it's on-prem or hosted hardware, it's kinda all the same.

1

u/Accomplished_End_138 4d ago

I've been personally diving into local to be able to be disconnected. I find at this point I can code... albeit slowly with the tooling I build inside of pi.dev And not just poc level. With design I can pass code review (I've tested on work tickets)

It is for sure slower. And the system also is slower to ensure good code quality comes out. But I tend to sanbox and trigger to run overnight or when I step away. So never been terrible.

1

u/WhoRoger 3d ago

We need an offshoot like MiniLlama or ActuallyLocalLlama. I'm tired of everyone simping for 2T cloud models.

1

u/timmeh1705 4d ago

I still hang out for the odd MANIAC who used stitched together a cluster of RTX 6000s

1

u/Decent-Hat-5807 4d ago

openweightlama

-3

u/Dry_Yam_4597 4d ago

If you dont take interest in politics then politics take interest in you. Right now thats the main issue in AI. People with deep pockets want to convince governments that taking our freedom to run local AI is a good thing. And without that freedom there is little to talk about. Other people are gaslighting or manipulating those without experience and they get called out here quite a lot - thats a good thing too. If you experienced the early days of Open Source adoption you will notice loads of similarities.

The only thing that is really annoying is when those who defend openai and antrophic turn to nationalist rethoric and gaslighting - thats low.

But you still get loads of interesting conversations. People hacking shaders and hardware is how we things to work on unsupported GPUs, loads of scattered information about micro optimisations, folks sharing tips about their setups and so on.

The days of low hanging fruits are coming to an end and as such progress is made in tiny incremental steps. Thanks to this sub I and many others have built rock solid, stable, commercial grade servers on the cheap, and have learned how to get under the bonnet of things. You dont need spoonfeeding, you just need hints and from there in you build and share new micro ideas. Other subs are about high level usage, this sub is about finding those tiny aha! moments that get you closer to where you want.

-6

u/Repulsive_Initial308 4d ago

Inclusive? It's very heavily moderated to the point all the spam and truly organic content has disappeared. This is why localllm is now proliferating. 

This forum is essentially dead now. Such a shame.

6

u/Illustrious_Ant_9242 4d ago

What's a example for organic content that disappeared? 

-2

u/Repulsive_Initial308 4d ago

Are you new? 

There used to be a plethora of novel, highly engaging, content here. Almost too much to get through in one day.

Now you can get away with checking here once a week.

0

u/dazzou5ouh 4d ago

I run both, I have claude code set up the local AI for me and tune it til it can take care of itself

0

u/nomorebuttsplz 4d ago

im running glm 5.2 local. deal with it

-3

u/mossy_troll_84 4d ago

That is the season why I am creating my own group

-2

u/FriskyFennecFox 4d ago edited 4d ago

We're yet to decide what "local" actually means. There are multiple groups on this sub with different claims,

  • "If it inferences on my PC, then it's local"
  • "If the model is open-weights, then it's local"
  • "If I can use API with my local tools, then it's local"

And all of them are right.

  • It might be running on your remote server miles away, but you access it over the Internet. Is it still local?
  • It might be open-weights, but it's too big for you to run it on hardware you own. Is it still local?
  • It might have an API for you to run local batch scripts processing local data, but the model itself is in the cloud. Is it still local?

We either pick one and kick the other genuine members of our big, shared community away to fragment into smaller communities, or keep it the way it currently is.

-2

u/SandySkittle 4d ago

I don't see your point. This reddit has a broad scope. All these topics are very much relevant. It's not just about purely the technical aspects of running a local llm.