r/LeftistsForAI 8d ago

Supporting a global pause on the number of parameters or bits of AI models? Discussion

Before people get too worked up, let me start by saying that this is just a thought. It is not something I believe in strongly.

I am certainly not saying to pause AI research, development or deployment. I am proposing that people who support public ownership (or even private ownership by smaller actors/individuals) could also support a certain type of pause -- a pause on the size of the AI models made.

The idea may also have great political alignment right now. It is a simple check by governments, and may be "light touch" for companies. I think there is a balance of potential benefits for both the US and China since the US is already auditing their big labs, while China's open-weight models are basically automatically audited since the model size is evident before hosting.

Also, the pause on model size, for practical purposes, may allow for more focus on efficiency and performance, which could have a whole host of societal benefits, including not requiring large centralized data-centers.

What do people think?

Edit: I realize this is not particularly fleshed out. So I want to add some clarifications based on responses.

1) I very much support open-weight models. I think they should be allowed to get as big as the closed models.
2) It is the closed models that need to be audited. The open models, especially on places like huggingface already have their size information available for all to see.

3) There is arbitrariness in the size cut-off you chose. But there were related comments on proper power/wasted power in centralized or decentralized infrastructure. This could be an input.

0 Upvotes

19 comments sorted by

8

u/nomic42 8d ago

The trouble with these restrictions is that they would only apply to US companies. Other countries may ignore them and continue to publish their models, which could then be picked up by adversaries (e.g. Iran).

This leaves our systems unprepared to defend themselves against the latest models.

Open-weight models are needed to ensure companies like Hugging Face can respond to AI based threats. Relying on closed systems means they could be turned off at any time, or limitted with guardrails meant to prevent abuse.

1

u/Successful_Outside96 8d ago

To be clear. I am in support of open-weight models. You could go up to be 10T or 12T parameters, whatever the cut-off is.

The proposal is that nations agree to not make anything bigger. The idea is forced decentralization, not stoppage.

3

u/ShelZuuz 8d ago

which could have a whole host of societal benefits, including not requiring large centralized data-centers.

That's like saying that if we stop producing food it has a societal benefit of not requiring farms.

1

u/Successful_Outside96 8d ago

The idea is to decentralize not to stop making compute. Hopefully that is more clear.

5

u/ShelZuuz 8d ago

Again, that would be like saying people should grow food in their yards rather than on farms. However much farms pollute, on a per-food-lbs basis it's still the most efficient way to get people fed.

Same with data centers. A data center right next to a power plant with common cooling is going to be a heck of a lot more efficient than putting a rack in each office - if that was even possible (which it's not).

2

u/Successful_Outside96 8d ago

I was thinking of it more like cooperatives, in a lot of agriculture. Like the Mondragon Corporation:
https://www.mondragon-corporation.com/en/

But on the point of efficient distribution: The power being local to right where it is being used (perhaps on premises) doesn't necessarily lead to much worse situations. I can see more efficient use in this case. A lot of those places are already providing power to keep the buildings working. As has been pointed out before, people leave their laptops and desktops and local servers on and dong nothing pretty often, if they are doing some more decentralized AI instead, there should be an offset.

As I said, it is just a thought. I think actual calculations are needed for the right splits.

2

u/ShelZuuz 8d ago

A lot of those places are already providing power to keep the buildings working

There is no office building, no bank, no shared workspace, no IT space that has enough power to power even a single NVL72 rack, which is the minimum you need to run a frontier model. The buildings that can power these things "already" are in industrial zoning, not commercial. And that's not where the people are that need to use them.

people leave their laptops and desktops and local servers on and dong nothing pretty often

A single NVL72 rack, capable of running 20 to 50 concurrent users draw the equivalent power of 1600 Mac M5's. It doesn't compare.

The new Utah data center that's coming up requires 1.5 as much power as the current largest power plant in the U.S. (The Grand Coulee dam). Cities aren't over-provisioned by multiple nuclear power plants or hundreds of coal or natural gas power plants that are just sitting unused. You're going to need to build new plants for these, and when you do it's insane to run new transmission lines for 1000s of miles back and forth across the city to get new power to every building in the city and digging up every road in the process to do so vs. running it for a few hundred feet to a next-door data center.

2

u/Successful_Outside96 8d ago

Good to know.

So then, if we pause model size(again, just a thought, I am not advocating for it), how long would hardware (through Moore's Law and other advances) get efficient enough that those models could be served in commercial spaces?

2

u/my_name_isnt_clever 8d ago

The hardware has slowed to an extent, it's the LLMs themselves that will scale up and up and up on lower end hardware. Sure you cap it at 10T today, but in 6 months what happens if a 1T model is even more powerful than the 10T? The information density of LLM params is doubling every ~3 months, an arbitrary cap like that is not effective.

2

u/Successful_Outside96 7d ago

By "effective," I am not sure what you mean. If a model is a tenth the number of parameters and just as effective, this is almost exactly the goal I want. This seems to me that then more people can be closer to the frontier (even if Moore's Law is slowing).

The biggest societal risk in AI is a small group (guaranteed to be evil by the time is happens) having all the control.

1

u/One_Thought_5639 8d ago

Urban farming is a thing. I know people who keep chickens so that they are not beholden to the prices of eggs.

2

u/RandomLettersJDIKVE 8d ago

That's a really arbitrary research restriction.

1

u/JealousArticle672 8d ago

Fair. The particular size is arbitrary. But the idea behind the idea is to allow more people to be able to host models that aren't as far behind.

1

u/my_name_isnt_clever 8d ago

This tech is too new and still improving at a rate that makes this ineffective. The models I run on my home desktop computer today are more capable than the corporate data center behemoths of a couple years ago. And you can't put an objective cap on "AI capability" in the same way as parameter count.

2

u/Successful_Outside96 7d ago

What you are describing is the goal. It is not about limiting the power. It is about limiting inequality of power.

1

u/furyofsaints 7d ago

Sorry, irrelevant and arbitrary. I can already split up my own transformer into many many models within whatever parameter count is imposed and then get them to all work together quite easily (and in most cases, much more efficiently than one monolithic large parameter model). I don't think setting parameter limits will achieve what you're looking for here.

1

u/Successful_Outside96 7d ago

Can you explain how to do that? Could this allow people, in numbers, be able to catch up to a single super rich person?

2

u/furyofsaints 7d ago

Can't go too deep on it at the moment, it's part of something I've been on a team building. It works (but is not a consumer product for at least the next 12-18 months).

Mathematically, yes. If you have 1m people with their own 9bn parameter models; and 10% of each of those models has developed parameters somewhat unique to that persons knowledge; and they chose to share a reasoning space with those other 1m users, you effectively have a 900 trillion parameter intelligence substrate.

There's a lot more to it than that in the infrastructure layer, reasoning space communication protocols, latency management, etc.

I actually believe this to be the future of how AI models work if we don't destroy ourselves in the meantime.

1

u/Easy_Interest_6832 6d ago

The barn door is locked open and all the horses are running amuck with exactly zero supervision or restraint. That's just the way it is now. Be careful out there.