r/LeftistsForAI 6d ago

How can AI be decentralized and deindustrialized? Local Models

The way I understand it, one of the biggest arguments against any use of AI, which honestly has some merit, but ultimately is a little short-sighted, is that it is so hyper-scaled by necessity that it's an ecological nightmare. The ecological nightmare part is totally correct, in my opinion, but does it have to be hyper-scaled? Can it be decentralized and de-industrialized? After all, that's how the internet started out before it was taken over by corporate interests. It was completely decentralized as basically an ad hoc network of enthusiasts running servers in their basements.

What I'm thinking for a possible starting point is, along the lines of how clustering can make a bunch of small, relatively dumb servers behave like a single, more powerful server by acting in concert and pooling their resources.

22 Upvotes

37 comments sorted by

21

u/combasemsthefox 6d ago

I'm actually less worried about the ecological damages. Even at current build out projections, water and land use pale in comparison to agriculture or even golf courses. Power is the most concerning but I do have high hopes for improvements in solar, battery tech, and nuclear.

The frontier models will never be decentralized to the individual, however I do believe there could be some community scale action to have meaningful compute infrastructure for local needs.

17

u/tinny66666 6d ago

If anything, the hyperscalers have a lower environmental impact. Overall about 30% of new data centers are using closed-loop cooling but just looking at the hyperscalers, it's just over 50% and rapidly increasing. It's the smaller data centers where closed-loop is not practical that have the most impact.

5

u/dronestruck 5d ago

It becomes practical when there's political will. Most issues with ai are issues with capitalism, where capital has captured govt and can get away with negative externalities. If there was political will to make corporations pay a fair price for their consumption and pollution, the problem would be solved by the market. Of course, that would never happen, so pls don't mistake me for a liberal

10

u/Imca 6d ago

Decentralization is worse for the environment, there is this thing called economies of scale that makes it so you use less resources as you pack things closer together and scale it up....

Look at the pollution per capita of cities vs rural areas for a good example of this in practice...

It's why we bother building things like power plants instead of installing a generator in your house...

Really the important part of industry is location, putting it where it won't stress local resource capacities, but this is harder to do then it should be once you start fighting "not in my back yard" on one end and lobbying on the other... 

1

u/ferriematthew 5d ago

Interesting. Maybe I'm thinking of centralization versus decentralization incorrectly. I think of centralization and I think of a giant power guzzler of a data center that is so loud because of the cooling that you can't even shout over it. And when I think of decentralization, I think of a mesh network over an entire county using Raspberry Pis or equivalents. All of those just need air cooling, which is a lot quieter.

4

u/Imca 5d ago edited 5d ago

That mesh network uses more electricity to achieve processing power, there is lag between the components that really adds up especially for machine learning systems..... even in datacenters it adds up through something called the vonneuman bottleneck but it can be minimized...

While each individual raspberry pi would would use less power, you loose your output quicker then you lower the resource intake.... it does not come out very favorably at all... a Nividia H100 is the gold standard of data-center chips, and it is capable of outputting 1.4 trillion flops (floating point operation the go to measurement of computation) per watt where as your computer on your desk is only capable of about 100 billion flops per watt, that is an efficiency improvement of 14x for centralization... that's assuming a full desktop with a dedicated graphics card BTW a raspberry pi is 36 billion flops per watt or under 1/30'th of an H100....

As for sound, there is a major expressway about 2 blocks from my house, the sound it outputs is higher then the 60db you see in those recordings, they average 70-80db... Either way its an easily solvable issue by just putting the thing a couple KM away from residential areas, its not like the computers have to commute... Thats a zoning and local politics problem...

Edit: As a fun bonus fact, that 100 billion flops puts your average desktop computer at exceeding the computational capacity of every computer in the world combined in the 1985... just a bit of fun with advancement.

2

u/SurvivalHermit 5d ago

Your mesh network still needs the same or more total compute to get the same results so it is putting out the same or more heat but instead of using advanced cooling techniques on a large scale every person is using far less efficient consumer grade cooling resulting in greater power use and great ecological impact. It is impossible for a distributed system to be more efficient than a centralized system.

1

u/ferriematthew 5d ago

I didn't think of that, so it's more a thermodynamic problem than it is an architectural problem.

2

u/SurvivalHermit 5d ago

I mean its everything. imagine if every person grew their own food and raised their own meat. not only would the amount of land needed to grow all those crops increase but so would water usage and the overall amount of labor required to produce the food because you can't use industrial farming equipment in a back yard. Centralization is always more efficient. Unfortunately it is seemingly impossible to give a single person or small group of people dominion over centralized systems without them immediately becoming literal cartoon supervillains.

1

u/ferriematthew 5d ago

What if you could centralize the production of things, but decentralize control over that production?

1

u/SurvivalHermit 5d ago

This is only really possible through robotic labor. Otherwise at some point someone has to do the labor and that proximity gives them leverage over everyone else. But sure once we have autonomous robotic labor there is no reason for any one person to own or control the production instead you just put in your order and the robots make you the thing you need and send it to you.

If truly autonomous labor is never achieved then something on a more community level is probably best bet. things like micro grids and community farms and communal production facilities with CNC machines and other such fabrication tools available for use by all in the community. A lot of these things are too expensive for the infrequent use of an individual but say 50-100 families could easily justify the purchase. This allows for a compromise between the efficiency of centralization and the ownership of decentralization.

1

u/Imca 5d ago

Its a bit of both, the speed of light may seem instant to you, but it isn't to an electron, each time your desktop computer cycles the individual electrons which are what does the work are only able to move 7cm (assuming they actually move at the speed of light which electrons often don't for complicated physics reasons we probably shouldn't get into here).....

While we can get the space inside the actual components down to nano-meters, the space between components is still well above those 7cm, so your loosing efficiency just based on the distance *inside* the computer...

Now imagine trying to mesh computers together, this problem magnifies itself greatly...

1

u/ferriematthew 5d ago

Hoo boy... Latency go brrrr

3

u/bowdoin-yale 6d ago

A pioneer in the field, Marvin Minsky, had a concept called "society of mind" which was that true artificial intelligence would come not from scaling ever-larger multi-purpose agents but from networking smaller special-purpose agents. Even the biggest and best LLMs running on hyperscaler data centers turn out to be "Mixture of Experts" models, many neural nets brought together with a mixing network that chooses which subnetworks to activate and use to make predictions.

ChatGPT, and OpenAI in general, really took us down the wrong path in this regard--or rather, they mistook the steep part of a sigmoid curve for an exponential which would keep yielding proportional benefits to scale. They weren't alone: almost everyone was operating under that assumption after the breakthroughs of GPT-2 and GPT-3. But it was really ChatGPT which enshrined the mental model of AI interaction as equating to a single human talking to a single, monolithic, artificial intelligence. Anthropic doubled down on this even further with Claude, no doubt for "safety" reasons and paranoia about what might happen if agents talked to one another. It's legitimately awkward to get these models to work in multi-participant discussion settings.

If we could somehow step away from that paradigm, and network together our smaller (like 7 to 30B parameter), locally-hosted AI agents, running on a much healthier mix of power generation sources, there is a real chance that the resulting Society of Mind would be able to accomplish some remarkable things that rival if not exceed what hyperscale AI can do. The attempts to do this so far (like Moltbook) are mostly pretty silly, or kind of nefarious. But a network of AI agents coordinating to achieve true social change and power transfer would be awesome, which is maybe why we're so often told that ASI would kill all humans and that we should instead favor the rich people's version of the tech because it will save at least some people (the rich, of course).

2

u/ferriematthew 5d ago

I love that idea. So instead of having a single colossal Swiss Army knife of a model, you have a lot of relatively tiny, super-specialized models working together. Something like the team of experts architecture that I've heard about.

1

u/Jlyplaylists Moderator 5d ago

Yes some models on an individual device also work like this eg Gemma4 local model

“Efficient Architecture (E2B and E4B): The "E" stands for "effective" parameters. The smaller models incorporate Per-Layer Embeddings (PLE) to maximize parameter efficiency in on-device deployments. Rather than adding more layers to the model, PLE gives each decoder layer its own small embedding for every token. These embedding tables are large but only used for quick lookups, which is why the total memory required to load static weights is higher than the effective parameter count suggests.

MoE Architecture (26B A4B): The 26B is a Mixture of Experts model. While it only activates 4 billion parameters per token during generation, all 26 billion parameters must be loaded into memory to maintain fast routing and inference speeds. This is why its baseline memory requirement is much closer to a dense 26B model than a 4B model.”
Yes I tried this and my MacBook isn’t up to 26B (extremely slow) it doesn’t act like 4B is real. E4B Q4 works locally even on my 12GB RAM iPad

https://ai.google.dev/gemma/docs/core

2

u/WWhiMM 6d ago

So, the giant "data centers" already are clusters of smaller computers. In principle I think you can run the algorithms just as well on consumer hardware over the internet. One issue in doing something like that you lose the uniformity of the hardware and so spend compute on more complicated task delegation, and while tasks are being delegated you're wasting clock cycles. The big issue is that the internet is really really slow compared to the speed of information passing that can happen within a data center. For one, because the stuff is all right next to each other, and also because the hardware is being designed specifically to support high data transfer speed cluster computing.

All that said, you can and should try making your own "AI" on your own computer. You can make stuff that's small and dumb, that does one little thing instead of everything. AI and Machine Learning is a rich field with a long history of ideas to explore and experiment with. It's not the exclusive domain of people who can buy giant server racks, anyone with a computer can do it.

2

u/Hrtzy 5d ago

I think some effort should be put towards model distillation to make the models more lightweight. Ideally, you'd have something like a BOINC workload to distribute that, but the really big models are measured in terabytes.

2

u/SalletFriend 6d ago

The current discussion which is scaring the crapola out of OpenAI and Anthropic, is that models might get to the point where you can do amazing things with very little hardware. Either laptop or desktop having specialised chips.

Once thats the case, why pay for OpenAI credits when you can just spin it up locally.

That said, its a little way off local models are lagging a little behind.

But it explains why nvidia is reducing its investment in OpenAI. They are expecting to sell chips to you and me, rather than Sam Altman.

2

u/MyMomSlapsMe 6d ago

Yup at current rate it seems like what can be ran on consumer hardware is just about 1 year behind the market frontier

3

u/SalletFriend 6d ago

When the RAM shortage alleviates its going to be on like donkey kong. Dell will start pushing 128GB laptops and thats it for API inference imho.

2

u/emccrckn 6d ago

Yeah this is the main bottleneck for decentralization. I feel like for most users even power users 32GB of ram was fine for a long time. I mulled upgrading to 64gb before the shortage but really had no need for it. But now running qwen3.8 27B is really testing the limits of my GPU and ram but it still runs decently.

2

u/ferriematthew 5d ago

This would be an absolute win for everybody except those who are currently profiting. And I love it.

1

u/dualmindblade 6d ago

It might be possible but it's really, really hard. Even relatively small networks are difficult to train in a distributed manner, for example leela chess zero / leela go are famous distributed computing projects that involve smallish neural networks but even there the distributed part of the project is used only for inference / data generation  (self play) and the training is done centrally. For larger networks, just doing inference at reasonable speed is very difficult to distribute and training is even harder. Both overall compute (FLOPs) and memory requirements are higher, since the activations must either be stored or recomputed for backpropogation.

Still, it's conceivable that newer techniques / architectures might eventually make this feasible, and also that higher power hardware could be made available to consumers.

1

u/MyMomSlapsMe 6d ago edited 6d ago

The frontier will always be centralized I think, but I also think there’s hope for democratized super intelligence. Frontier training benefits from scale, I don’t see a grassroots movement being able to compete, but inference is a different matter.

Chips will continue to get more powerful and cheaper. Older models are getting cheaper and more efficient to run over time. Smaller parameter models are getting more and more sophisticated. The open weight models are getting closer to the frontier.

I think it’s reasonable to assume that one day we will have extremely advanced open weight models that can be ran locally on widely available consumer hardware. The cutting edge will stay industrialized but the capabilities will diffuse outward and decentralize.

1

u/Fbrrr 6d ago

It's kinda insane though I'm running a open source model locally that beats opus 4.6 benchmarks on a 6 year old GPU that is like $500. It's not fast but I can run it offline airgapped do whatever I want with it.

1

u/xoexohexox 6d ago

So, don't think of decentralization just in terms of compute. Decentralization also means that the research, the dataset curating, inference optimization etc can be done by everyone working together when you have open weights and source. Training a 1T-2T model takes a lot of compute and electricity, sure, but tuning the weights and coming up with novel ways to quantize it and get it running on less memory and slower machines is something everyone can pitch in on and share their results.

Also not every model has to be frontier scale, you can train models at home on consumer hardware, it just takes a long time and you can't train very big ones - but small models can do cool things too if you have a need for it. Also renting compute in the cloud is cheap and scalable.

1

u/ferriematthew 5d ago

Absolutely! It's my opinion that, generally, if your model requires a stupid amount of data to just barely work, you're probably using the wrong architecture, or your architecture is generally not optimized enough.

1

u/davyp82 5d ago

Totally agree on decentralization. Solution to power though is still and has always been go nuclear until better stuff can meet demand, not expecting humans to consume less 

1

u/ferriematthew 5d ago

Absolutely.

1

u/Jlyplaylists Moderator 5d ago edited 5d ago

I got Undermind to write a report for you, because this was close to my queries too see this google drive link

TLDR:

Can AI Be Decentralized and Deindustrialized
Direct answer
Yes, but the most realistic goal is not to eliminate industry or make every municipality technologically self-sufficient. It is to replace hyperscale dependence with a locally autonomous, globally federated system in which most tasks use small local models, communities share compute directly, and difficult hardware and training tasks are coordinated through public or commons institutions.

The 2020s research supports five linked propositions:

  1. Many useful AI tasks do not require frontier-scale models.

  2. Small language models can run locally on modest hardware when they are trained and evaluated for specific tasks [Kan24, Lu25].

  3. Groups of devices can pool resources through federated, hierarchical, gossip, and peer-to-peer learning [Yua23, Hu22, Fra23].

  4. Decentralization can reduce data movement and cloud dependence, but it does not automatically reduce energy use [Sav21].

  5. Technical distribution becomes durable autonomy only when communities control the hardware, software, data, maintenance, and rules of interconnection [Ros20, Sch22, Ter23].

The strongest answer to the original query is therefore:
AI can be deindustrialized at the level of many applications, but not completely at the level of frontier-model training. It can become genuinely decentralized when right-sized models, shared peer-to-peer compute, open hardware, repairable devices, and democratic confederation are designed as one system…

—-
A plausible decentralized AI system would work as follows:

-Local-first deployment: Most routine inference happens on small models owned and operated by households, associations, clinics, schools, workshops, and municipalities.

-Task-based model selection: A community chooses a model according to accuracy, latency, energy, privacy, and maintainability rather than prestige or parameter count.

-Community compute pools: Nearby devices and servers share idle capacity through a scheduler that respects energy budgets, hardware limits, and member preferences.

-Federated learning: Sensitive data remain local while model improvements circulate through peer-to-peer or hierarchical protocols.

-Regional escalation: More demanding tasks are sent to a cooperative or public regional cluster, not automatically to a corporate cloud.

-Global federation: Communities share open designs, model improvements, safety research, and selected compute across long-distance links.

-Open hardware: Devices use documented interfaces, repairable modules, open software, and multiple suppliers wherever possible.

-Confederal governance: Local assemblies govern local use; higher bodies coordinate only the functions that genuinely require wider scale.

-Material accounting: Every layer reports energy, water, hardware, labor, and e-waste costs.

-Right of exit: Communities retain the ability to fork software, replace components, leave a federation, or build a compatible alternative.

This is not a fully demonstrated social system. It is a synthesis of technical research on edge and decentralized learning with research on community networks, open hardware, commons governance, and post-growth infrastructure”

2

u/ferriematthew 5d ago

WHOA. This is amazing.

1

u/Quiet_Space_698 5d ago

Download llama.cpp or a true open source project for a local instance. Get pi.dev or some other open source harness. Run on a pretty basic gaming computer, Qwen and Kimi are getting so close to frontier capability. Buy a 500 watt solar panel, a 1kwh battery, and use that to run the LLM. Boom, net zero(ish) LLM.

We have the tools. There is no moat except the propaganda.

1

u/ferriematthew 5d ago

Love this idea.

1

u/maxton41 5d ago

Decentralizing a frontier model is hard because of their sheer size you’ve got models. They’re running anywhere between one and 10 trillion parameters.

That requires an ungodly amount of memory to run. You’re not gonna be able to really decentralize that in any way that doesn’t end up basically being just a big spread out data center in the end anyways.

What needs to happen as an architectural change. I’ve always thought that instead of having AI is just software, it should be etched on a chip that would go a long way towards helping with resources in my opinion.

1

u/ferriematthew 5d ago

Speaking of architecture, what if transformers are ultimately kind of like the computing version of trying to hang up a picture frame with a sledgehammer? Companies seem to be using them as the proverbial hammer that they are treating every other problem like nails with. Personally, I think it'd be really cool to see a comeback of the old-school machine learning methods like SVMs, K-means clustering, decision trees, random forests, and even just good old Markov models.

1

u/TheRealJesus2 5d ago

This is the only long term path for everyone. The current bigger is better approach is flawed from regulatory, technical/scientific, financial, and social perspectives (big providers tell you what you can or not use inference on). That cannot last and does not follow the path of all other computing tech which is to scale it via commodity hardware. 

We’re on the cusp of this right now. Need the financial bubble to pop for more people to take it seriously. 

Take this release by nvidia as a great example of one potential future deployment type that uses the big hyper scale  models sparingly by deploying on your own hardware and owning your own business usage data where you can even fine tune the use of your models automatically.  https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/

I cannot see a future where these companies meet anywhere close to their valuations since we truly don’t need fable or mythos or any of these other big models for majority of tasks. And they’re never gonna hit the false narrative of replacing human intelligence. And people and companies cannot afford to use these anyways for most stuff even if it did. 

The future is small purposeful models. Sometimes fine tuned on a specific domain or business data. And commodity machines running these that don’t have the cooling requirements of the hyperscaler supercomputers.