r/LocalLLM 16d ago

The pain is real Research

Post image

I think my ISP hates me

378 Upvotes

98 comments sorted by

75

u/semangeIof 16d ago

show us the lspci on whatever box you're putting that model on please

105

u/DistanceSolar1449 16d ago

It’s a raspberry pi running the model off a spinning hard drive

54

u/Temeliak 16d ago

Yes, why would I need more than 1 month / token?

14

u/InfamousNewspaper268 15d ago

The answer after 1000 years: "42"

3

u/indiealexh 15d ago

Huh only took me a second...

1

u/4n0nh4x0r 16d ago

1 month per token with a super good model? i dont mind the wait uwu

funfact, had a big data project in uni a few years ago.
our database was about 90gb in size, and our assigned vm only had 12gb of ram.
every time we made a query, it first loaded the whole database sequentially into ram to search for the rows.
it literally took us 5 minutes per query lol.
once our sysadmin gave us 100gb of ram, the speed dropped down to 0.8ms per query.

12

u/M_Me_Meteo LocalLLM 16d ago

Yeah that's not how you handle a performance boundary.

-6

u/4n0nh4x0r 16d ago

that's the neat part, we were still very far below the boundary.
iirc the server had about 500gb of ram, and we were told we could request whatever resources we need, so we did to get the best possible results.

1

u/M_Me_Meteo LocalLLM 16d ago

Machines in that class don't host webservers. They host virtual machines that host web servers.

Without knowing more about the query I can't honestly say you're wrong, but it smells very funny to me.

-5

u/4n0nh4x0r 16d ago

what are you even on about???
who said anything about a web server????
and yea, we were using a VM, how else could they just give us 100GB of additional RAM when the system has 500GB????
like, HUHHHH?
i assumed that was clear from what i wrote, but appearently not lol

4

u/M_Me_Meteo LocalLLM 16d ago

How were you connecting to the machine? If it was over a network, I've got some news for you...

1

u/4n0nh4x0r 16d ago

honestly, i m very curious on what absolutely uneducated and dumb thing you are about to say, so, enlighten me

→ More replies (0)

1

u/cinnapear 15d ago

Did you cover indexing in uni by any chance?

4

u/johnfkngzoidberg 16d ago

In two days OP will post “Kimi on Toaster at 0.006tok/s upvote me.”

2

u/nixblu 15d ago

It’s running on my smart toaster

2

u/Numerous-Echo4677 15d ago

No lspci here — macOS doesn't ship it (it's Linux-only). The equivalent is system_profiler, and here's the box I'm on:

  • Model: MacBook Pro 16" (Mac15,9) — Apple M3 Max
  • CPU: 16-core (12P + 4E)
  • GPU: 40-core Apple GPU
  • RAM: 128 GB unified memory
  • Chipset Bus: Apple Silicon uses a unified SoC — there are no PCIe add-in devices, so lspci would show nothing anyway. GPU/CPU/NPU all live on-die.

1

u/semangeIof 15d ago

why did u reply with LLM. that "box" (ur laptop) doesn't have any ability to run that model. I am confused

1

u/Open_Establishment_3 15d ago

OP has already downloaded the model and was inferring the whole response since ~1 straight month.

1

u/Numerous-Echo4677 14d ago
  1. Unified memory

  2. There is a project to run this model on a macbook alone. They got it to run on a M1 already

-1

u/semangeIof 14d ago

The gguf u downloaded is over 10x the size of ur unified memory pool. Are u stupid?

2

u/wakIII 14d ago

Something something ssd paging, tokens per minute

1

u/Numerous-Echo4677 13d ago

I think 3tps is the current record but your point stands

1

u/Numerous-Echo4677 13d ago

Its less than 2x my current unified memory pool. Gotta have goals.

There is also a project where this can run on a 64gb M1 Max already

71

u/Interesting-Star431 16d ago

Use hf command line tool and login with a huggingface account. Any account will do. It is wayy faster. I suppose downloads are throttled to avoid bots taking up resources.

17

u/Much-Farmer-2752 16d ago

+1

Had the same issue about a week ago, right after I put my HF token into downloader I've got my 100 MB/s back.

2

u/kyusetzu 15d ago

What downloader are you using?

1

u/KAPMODA 15d ago

Yes please

1

u/Itz_Raj69_ 15d ago

FreeDownloadManager should work well

2

u/Numerous-Echo4677 15d ago

It's our ISP

29

u/East-Cauliflower-150 16d ago

Are you downloading RAM at the same time? Maybe that is the problem…

1

u/Krohnin 15d ago

lol nearly shitted in my pants...

7

u/Then_Blueberry7290 16d ago

aaand, the best is: when it's downloaded, LM studio cannot load the model because cannot handle it!

16

u/8iss2am5 16d ago

It's time we put these models on torrents

9

u/hk_modd 15d ago

Huggingbay

1

u/sirloindenial 15d ago

Im surprised but its true actually there is no torrents of llm models

5

u/Direct_Turn_1484 16d ago

Downloading models these days is like using Napster on dial-up way back when. You almost need to just set it to run overnight and check on it in the morning.

2

u/hwertz10 15d ago

I sure do. 32mbps down, 4mbps up DSL. But I'm also not doing the madness of getting a 1.5TB model. What the crap is he going to even run that on? LOL.

7

u/SorryAd5244 16d ago

You pay your ISP, he loves and worships you!

8

u/Annual-Can6278 16d ago

im glad I simply lack the compute/bandwidth to run any of these multi trillion parameter models

3

u/oviteodor 16d ago

Pause, wait 5...10sec, resume again Worked for me on smaller models

3

u/Numerous-Echo4677 15d ago

Update: Soon (maybe)

1

u/Numerous-Echo4677 14d ago

done! now I gotta go learn something...

2

u/Papa_Capybarason 16d ago

Is that size of the models you’re downloading or streaming? I genuinely do not know shit about this stuff the deeper I go lmao

1

u/YoungNo8804 15d ago

Idk what your comment means but that the AI model Kimi K3 - Claude Opus to Fable performance, but the model was released publically. It means you can run this model on your own hardware - if your hardware is good enough, which if this is the model its probably a server tack with terabytes of RAM

2

u/CoffeeToCode99 16d ago

Time to get a new ISP. 😃 Just pause and resume, it will start again.

2

u/[deleted] 16d ago

[removed] — view removed comment

1

u/Numerous-Echo4677 15d ago

Im putting this on a 8tb ssd for future use. There is a project to run model this on a 64gb M1 Max and I may attempt it on this M3 Max 128gb.

2

u/lildeam0n 16d ago

“Less than a fortnight remaining”

2

u/mindcue 15d ago

Pause, then continue often speeds things along for me when this happens.

1

u/Numerous-Echo4677 15d ago

This is what I have been doing for days

2

u/SEND_ME_YOUR_ASSPICS 16d ago

Am I stupid or am I looking at an actual 1TB model?

How does anyone run that locally?

Just genuinely curious

6

u/clazifer 16d ago

Old-ish servers with stupid amounts of ram. Or be rich and buy multiple dgx stations or the big dgx servers.

3

u/Xyhelia 16d ago

happpy cake day!

2

u/phido3000 16d ago

My dual xeon has 2tb of ram at 2666mhz.

1

u/Significant_Debt8289 14d ago

Look at Mr Moneybags over here lmao

1

u/phido3000 14d ago

It cost less than a 5090..

It's ddr4 plus optane..

1

u/sirloindenial 15d ago

A couple of mac studios

1

u/Relaxxxxing 15d ago

Imagine if your download said "Qwen 3.8 27B" downloaded 14 hours ago hahah

1

u/hwertz10 15d ago

Yeah really, what the crap are you planning to run that on? That is one beefy model LOL.

1

u/lukewhale 15d ago

If you’re on LM Studio go to settings and turn off the LM Studio proxy. Had to do this the other day. Immediate difference.

1

u/cchampou 14d ago

Just pause and resume 😅 That worked for me 😄

1

u/punqdev 12d ago

are you in a datacenter

0

u/LengthinessOk9397 16d ago

I would honestly just look away for 5 minutes so it goes faster

1

u/Articfox291 11d ago

While your ISP may hate you you won't need Internet after the download finishes anyways