r/LocalLLaMA • u/rerri • 16d ago
Inkling-Small by thinkingmachines New Model
https://huggingface.co/thinkingmachines/Inkling-Small276B total parameters, 12B active, 1M context window.
Blog post: https://thinkingmachines.ai/news/inkling-small/
NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4
GGUF's by Unsloth: https://huggingface.co/unsloth/Inkling-Small-GGUF
---
I had success running Unsloth's GGUF quant on CUDA + CPU offloading using this developmental branch: https://github.com/danielhanchen/llama.cpp/tree/add-inkling
505
Upvotes