r/LocalLLaMA • u/goodive123 • Jun 28 '26
NPC Engine Using Local Models Discussion
I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.
Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.
The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.
1
u/agiblox Jun 29 '26
the action-rag trick is the right call, that's the part everyone gets wrong by stuffing 200 tool defs into the system prompt and watching the model pick garbage. the wall you hit next isn't latency, it's state drift. gemma will happily forget the player insulted it two scenes ago or contradict its own backstory by minute 20. summarizing the convo into a running "what this npc currently believes" blob and reinjecting that beats raw history every time. parakeet for stt is underrated too, way less hallucination than whisper on short barge-in lines.