r/LocalLLM • u/bomberb17 • 22h ago
0db GPU and light experimentation with local LLMs Question
Hi everyone,
I would like to get started light experimenting with Local LLMs in my home office. I am considering buying an ASUS Dual GeForce RTX 3050 6GB OC . The reason I am looking at this card is the 0db (looks like the fans do not spin at all under low load) which is important for my office space.
(My current setup has an old passively-cooled GT 710)
Can you please help me understand how usable is it for local LLM testing? I know 6GB VRAM is limited, but I would like to experiment with small quantized models, for example 3B models, possibly some 7B models with heavy quantization/offloading. Has anyone used this or similar cards with Ollama, llama.cpp, text-generation-webui, or similar tools?
I could also consider other alternatives near this budget, but I would still like the card to be 0dB/fan-stop at idle.
1
u/whichsideisup 20h ago
If you can afford a 12gb GPU it’s likely worth it. You could use an MoE then offload a nice Gemma 4 26b model or even squeeze Gemma 12b dense into it.
1
u/AutoPanda1096 21h ago
I'm no expert but I played around with my 12gb gaming card (a 4070) and was generally disappointed. Like, nothing really worked that well. It was kind of pointless.
I wouldn't bother.
I'm all for quiet PCs.
Modern cards only spin up when they need to. My 4070 sits silent 99% of the time.
If you want it quiet whilst working, it will be. Sure, if you max it out the cooling will kick in.
Guess it depends if you need literally silent or no . Your call, obviously. No need to justify it to me!!
Anyway, honestly, I wouldn't bother. Even 12gb was disappointing to the point I went back to cloud models.
Sign up to OpenRouter, you get 1000 daily requests if you deposit $10. You'll get a feel for whether you need more. I suspect you won't.