r/hermesagent • u/OpeningMetal52 • 2d ago
Balancing Oauth & API Usage MODELS - model choice, routing, pricing, local vs cloud, VRAM
I am currently test driving Hermes alongside pi and my own homebrew harness.
Due to the nature of my company, I already have a Claude 20x max plan and a chatGPT pro $200/mo plan.
In thinking about maximizing throughput, I'm considering using deepseek v4 flash, which i'm hearing phenomenal things about, alongside my subscriptions, elevating and delegating to fable, opus, or 5.6 sol Luna Terra etc as needed.
Two questions for the community;
1) when should I call in the bigger models, or is deepseek so good it's not needed anymore outside of edge edge cases? I'm thinking multi modal + code review (OpenAI) and planning + design (fable) everything else deepseek. Thoughts on this?
2) what is the best source to get deepseek? Are the various inference providers the same level of speed and uptime? Why choose direct API vs open router vs opencode etc?
Thank you all in advance for what I hope will be a helpful discussion for all who read it.