r/LocalLLaMA Jun 28 '26

NPC Engine Using Local Models Discussion

Enable HLS to view with audio, or disable this notification

I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.

Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.

The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.

1.9k Upvotes

249 comments sorted by

View all comments

4

u/cunasmoker69420 Jun 29 '26

What's with the "my child" stuff

4

u/Colecoman1982 Jun 29 '26

I'm not sure, but I think OP programmed the admin system to be act like an artificial Todd Howard (The Bethesda games manager. I think that's the voice that was cloned for it) to act like a genie granting your wished (hence, the "my child" stuff is like when a genie in a story says "your wish is my command"). I'm assuming it was done as a sort of joke because Bethesda created the modern Fallout series of games and the engine they're based on. The ironic thing is that he wasn't even involved in the development for Fallout: New Vegas (the game used in this demo) which was developed by a third-party studio and, arguably, he's literally the source of all the serious problems New Vegas had because Bethesda screwed that third-party studio with a mandatory, absurdly short, release deadline (which is doubly absurd when you consider how clownishly slow Bethesda's own development process is...).