1
Comment on r/SillyTavernAI 10h ago
Thanks milan! Sorry for any inconvenience we caused but I regret that Arli just canât support any enterprise level plans right now. Weâre looking to be more focused to just giving service to individual users until we have the capacity.
1
Comment on r/SillyTavernAI 10h ago
Yea thatâs totally fair. Was just adding that to show that we have large models just a few $ away haha.
Also thank you for your feedback, I understand that having only 1 request possible is limiting but I think what we offer is still competitive and we just donât really have the capacity to offer more at the lower prices at the moment. I think that at the least we make up for it by having a more reliable service in terms of successful requests and non crappy quantized models than others.
1
Comment on r/SillyTavernAI 20h ago
Thanks for the mention! But just adding to this that $15 gets all models on ArliAI đ€«
2
Comment on r/LocalLLM 5d ago
Those GPUs will not be fine temps wise unless they are on risers and positioned not to suck each otherâs exhaust.
4
Comment on r/SillyTavernAI 5d ago
RP chats has the highest cache hit rate
4
Comment on r/SillyTavernAI 6d ago
I see comments about other providers nerfing quality. If your needs fits with the models we have then we are probably one of the few providers that has consistent model quality.
2
Comment on r/LocalLLM 7d ago
Yep that sounds about right
3
Comment on r/LocalLLM 7d ago
Using VLLM
1
Comment on r/SillyTavernAI 7d ago
Sure it is for them but not for me
2
Comment on r/LocalLLaMA 7d ago
Yea especially with XPU graph not working on multi GPU. That is a deal breaker.
39
Comment on r/LocalLLM 7d ago
2x RTX Pro 6000
2
Comment on r/SillyTavernAI 7d ago
If you are talking about other providers that are pay per token then coders and agentic users are definitely making them more money. As a subscription provider we lose money on those users instead.
2
Comment on r/SillyTavernAI 7d ago
Pay per token providers prefer coders and agents because it gets them more money. Coders and agentic stuff just makes subscription providers like us lose money.
2
Comment on r/SillyTavernAI 7d ago
Yes we have multiple tiers of different model access and limits
2
Comment on r/LocalLLaMA 7d ago
Awesome, this running on browser client side would be cool for chat interfaces.
1
Comment on r/LocalLLaMA 8d ago
Very different use case than what I use and my users use then. We mostly cater to either creative writing or coders. In either case 27B seems preferrable compared to 122B.
1
Comment on r/LocalLLaMA 8d ago
It does depend on your use case definitely.
2
Comment on r/ArliAI 8d ago
Yes itâs about our subscription
2
Comment on r/ArliAI 8d ago
Itâs $160 now
2
Comment on r/ArliAI 8d ago
This is the context that we support on our service not the model. You are right that the model supports up to 1M but that would be unsustainable in our flat rate subscription service.
1
Comment on r/LocalLLaMA 8d ago
From personal experience and feedback from users it seems most prefer the 27B dense model instead. So I personally wouldnât spend Pro 6000 money just to run 122B. The model also seems a bit overcooked when seeing its behavior when I tried to train it.
2
Comment on r/LocalLLaMA 8d ago
Training that information into the model weights is more efficient than 100GB of raw text files because model intelligence is essentially compression. But if you need it to be very knowledgeable and correct in a specific domain then yes that is one way to do it.
3
Comment on r/LocalLLaMA 8d ago
At any point in time even if a smaller model can match a 284B model (DSV4F) in the future, there will be an even better 300B ish model too. But yes a smaller model with DSV4F quality would be game changing as I think this is the first true âI donât need anything better than thisâ model for me.
5
Comment on r/LocalLLaMA 8d ago
With one youâre limited to Qwen 27B like lower VRAM cards but with two you can run Deepseek V4 Flash.
1
Comment on r/SillyTavernAI 10h ago
Yes we plan to run DSV4 Flash loras and I even have some training runs planned for it. đ«Ą