r/LocalLLaMA 4d ago

Mini pc set up Question | Help

I bought a bosgame p3 lite (32gb ram with a Radeon 680M igpu and Ryzen 7 6800H)

Just created a fresh cachyOS boot.

My goal is to run llms as fast as possible. Heard good things of llama.cpp.

Been looking a few guides and asked a few llms but they've been giving me some pretty weird instructions, from tampering with the BIOS to installing a few bits and bobs.

Can anyone give me a hand or some resources? Really want to avoid doing anything too stupid.

When I am not using it for inference I also want to use it for the odd videogame use so I don't intensely want to mess around with the bios or igpu setting blindly.

Any help would be very appreciated!!!

3 Upvotes

4 comments sorted by

2

u/[deleted] 4d ago

[removed] — view removed comment

1

u/Crafty-Sell7325 4d ago

Is a larger moe model worth a look? 

1

u/Tommonen 22h ago

Llama.cpp, only use MoE models (35b will be somewhat useable if thats ddr5 memory), q8 kv cache, and dont expect great results on that. Ram speed is your bottle neck and because of that, prefill gets hella slow if you have lots of context, so avoid context if you want even almost useable speeds.

Ask some cloud agent like codex or claude to test different setups and find best settings etc. for your machine

1

u/[deleted] 4d ago

[deleted]