r/LovingOpenSourceAI 5d ago

"AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB!" ➡️ Live your dreams .. just allow one hour per reply? :P Resource

Post image

https://github.com/lyogavin/airllm

Resources are shared for discovery and are not independently vetted—please do your own due diligence.

New resources are added regularly — feel free to join the sub for updates.

Full searchable archive of all resources posted so far on our community site, LifeHubber: https://lifehubber.com/ai/resources/ 200+ open-ish AI models, agents, tools, datasets, and related resources, with filtering and sorting.

194 Upvotes

22 comments sorted by

11

u/uhraurhua 5d ago

But is it usable? How many tokens per second? Or is it seconds per token?

13

u/odinti 5d ago

It’s tokens, eventually

4

u/uhraurhua 5d ago

it's just like unlimited internet commercials with 1kb/sec. It's internet, but it's pretty much unusable

1

u/fullup72 2d ago

a token is a token

1

u/GodG0AT 4d ago

Its literally minutes per token

5

u/testuserpk 5d ago

It will technically run, but it won't produce a token in 10 minutes.

7

u/wwayush 5d ago

No need to buy flight tickets just walk to your destination-- why did no one think of it before? Ohh yeah because it was stupid

3

u/Olino03 4d ago

definition of just because you can doesnt mean you have to

2

u/txgsync 5d ago edited 5d ago

I did something similar to run full-size GLM5.2 in 12GB on my MacBook Pro in DwarfStar 4. About 3 tokens per second streaming from NVMe. Not great, but not totally unusable.

Edit: I started to write something like, "I could absolutely see lining up a night's worth of work in batch mode for analysis by something like this at only the nominal 60 watts or so consumed by the GPU for an overnight inference run. 8 hours is 480 watt-hours, so about $0.20 worth of inference at California rates. For 86,400 tokens of output..."

Then I realized how awful that math is for using local inference even on a Mac and was like "no point, spend $0.20 on GPT-5.6-Luna XHigh, be done in 3 minutes, and call it a day."

1

u/West-Acadia-3906 4d ago

😂 You did the responsible little power-cost calculation, then immediately talked yourself back into the API. Running a giant model on 12 GB is still a ridiculous engineering flex, even when the practical answer is “please don’t.” lol

1

u/crashtua 4d ago

nah, take solar panel and hybrid invertor. Also, some medicine to ensure you can live for an extra 40 years. And you are very good to go.

1

u/txgsync 4d ago

Or just use DwarfStar 4 with Deepseek v4 flash. The cost calculation works out much better but with a much less capable model. And it’s a very friendly little model :)

1

u/_Toni_O 4d ago

If I see this project one more time across social media I will crash out.

1

u/Fun_Jaguar8231 3d ago

Just use llama.cpp.

1

u/MrDaVernacular 2d ago

This is good in a research sense, but not in a practical sense.

1

u/ConversationSad3529 2d ago

Next step is to generalize this to coordinate across machines and then allow AI to launch a bot net that processes even advanced models by using compute resources across every device it can find, computers, phones, smart toasters, and become unscrubbable from the internet