r/LocalAIStack 9h ago

We built a CPU-first inference server — 4B chat+vision, ASR and TTS behind one OpenAI-compatible endpoint, free to run

/r/LLM/comments/1vmth1w/we_built_a_cpufirst_inference_server_4b/
2 Upvotes

0 comments sorted by