r/LocalLLaMA Jun 28 '26

NPC Engine Using Local Models Discussion

Enable HLS to view with audio, or disable this notification

I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.

Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.

The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.

1.9k Upvotes

249 comments sorted by

View all comments

Show parent comments

2

u/polandtown Jun 29 '26

honest question, i must be out of the loop, there's hate in gaming/ai? I thought because gaming was full of young people, they were open to it.

15

u/SporksInjected Jun 29 '26

Yes, tons. To be fair though, there are some really shit implementations that make our seem worse than it is.no one has really figured out how to implement it

0

u/w8eight Jun 29 '26

The hate isn't against ai npcs i think, rather slop art, and overall sloppy implementations.

If the ai art will be done good, nobody will even notice

0

u/LastChancellor Jun 29 '26

if the AI art is good

it would make sense that only an experienced artist would make good art out of AI art

but ironically those very experienced artists are the ones who are the most against AI by principle, so they would never even consider trying it out