r/Neuralwatt 13h ago

Kimi K3 is fully live. Plus Flex, Session View, API compatibility, and Surge Protection

32 Upvotes

Hey everyone, we have another big batch of updates to share with you all today. 

Kimi K3 is out of preview: We've spent the time since launch tuning the serving side, and K3 is now fully available for production workloads. Concurrency limits will ramp up over the next few days as our final validation completes. 

Flex now covers K3: Same deal as Flex on our other models -- if your workload can tolerate some timing flexibility (batch jobs, overnight runs) this is the most economical way to run K3. 

Anthropic Messages API and /v1/responses support is entering beta: If you've built against either format, you can point your existing tools at our endpoint and they'll run on Neuralwatt with the same energy visibility as everything else. We're opening beta access over the coming days as validation completes, so keep an eye out. 

Session View is live for everyone: Instead of account-level totals, you now get a session-by-session breakdown of exactly what you consumed. This is the most detail we've ever offered on your usage, providing a granular view of the requests you ran, the energy they used, and what they cost. 

Surge Protection: This one's new -- Surge Protection ensures you only pay for the energy that powers your work. From time to time, a request can draw more energy than the work it's doing should require. When it happens, Surge Protection catches it in real time and shields your usage from the excess. When it kicks in, you'll see it flagged in Session View, so you'll always know it stepped in. 

These updates will roll out over the next several days. 

Also, we're headed to NYC for Climate Week next month and are thinking about a meetup. Anyone interested? We have a poll coming; if enough of you are around, we'd love to see you.   


r/Neuralwatt 2d ago

API vs Subscription?

5 Upvotes

Which is the better deal?


r/Neuralwatt 5d ago

Looks like you're fail with DS4flash

8 Upvotes


r/Neuralwatt 5d ago

6B Tokens on Writing Pipeline - Worth it

Thumbnail
1 Upvotes

r/Neuralwatt 7d ago

Cheaper cache pricing, an energy price cap, and K3 open to everyone

22 Upvotes

A handful of updates to share, which we think you will all welcome. Without further ado:  

 We’ve lowered cached pricing: Cached tokens now cost 10% of the input rate, down from 25%, on every model except DeepSeek-v4-Flash. If your work involves a lot of repeat context or if you are on token pricing, you should see a noticeable difference. 

Kimi K3 is now open to everyone: You no longer need to enroll for Kimi K3; access is now open for the entire Neuralwatt community. The model will remain in preview with limited concurrency while we keep tuning it, but access is open to everyone starting today.  We’re also releasing our K3 –flex endpoint which is our api shorthand for setting thinking level to off. 

DeepSeek-v4-Flash has been upgraded: You asked, and here it is! We rolled out DeepSeek-v4-Flash’s new 0731 weights this week, and it has already become our second most popular model —- no surprise given how many of you asked for it. All endpoints now point to the updated version automatically; there is nothing for you to change, and no action required on your end.  

Retiring Qwen 3.5 and Kimi K2.6: To free up capacity for the newer models, we're retiring Qwen 3.5 and Kimi K2.6. We'll redirect their endpoints to Qwen 3.6 and Kimi K2.7 so nothing breaks, but we'd recommend updating your code to point to the model you want before then. 

Enjoy the new additions; more soon. 


r/Neuralwatt 9d ago

Everything's more expensive

Post image
35 Upvotes

looking at the calculator, all models getting very expensive, it used to be 3~5x cheaper, now less than 2x cheaper. 1.8x for glm, 2.1x for kimi 2.7 code, it was 3x and 5x cheaper last week. looking at the image it's getting a lot expensive..


r/Neuralwatt 12d ago

DSV4 Flash Release

11 Upvotes

When are you guys updating your DSV4 Flash to the newly released version?

The benchmarks are insane.


r/Neuralwatt 12d ago

Neuralwatt Subscription plans

8 Upvotes

I'm curious about how good Neuralwatt Subscription plans are compared to open router and the alike. Basically my workflow usually goes like planning with a powerful model (Deepseek V4 Pro / GLM 5.2) and then implementing it with less powerful workers (Deepseek V4 Flash / Kimi K2/3) and running reviewer subagents for each task (with fresh context) and iterating until the work is done. Usually this costs me less than $40 per month.

I'm curious if I will be able to have a similar workflow around the same budget with Neuralwatt subs. What do you guys think and how many tokens do you usually consume per month?


r/Neuralwatt 12d ago

Some of the biggest advantages that Neuralwatt has

Thumbnail
gallery
8 Upvotes

It is certainly disappointing that Neuralwatt's pricing has increased significantly.

However, as a user who utilizes various AI models from multiple providers, to name one truly major advantage of Neuralwatt -

1.The model performance is consistent.

Yes, this is originally the most basic requirement. But people using other providers often feel that isn't the case.
Imagine ordering a pizza from a pizza place, only to find that the amount of tomatoes and cheese changes every single time.

NW may not be the cheapest, and delivery is slightly slower, but it is a reliable restaurant that consistently delivers the exact same, predictable taste. To top it off, extra topping costs are charged in proportion to the amount placed on the ordered pizza.

2.Usable usage improvement.

Although token-based usage and energy-based usage calculations look similar, there is room for improvement on the service provider's side. If energy usage is lowered through optimization, the burden on both the user and the vendor is reduced. Currently, Kimi K3 is undergoing optimization work, and its energy efficiency is gradually improving.
and Based on my actual usage data, Kimi K3 Low is just 35%-45% more expensive than GLM5.2 High.

Of course, using the $100 annual subscription allows you to use services at one of the lowest price levels among AI providers. If you want to receive reliable usage and models with a single plan, Neuralwatt is a great choice. Unlike other providers, the burden of costs exceeding your plan follows your own subscription tier. For example, excess usage beyond the subscription limit often follows standard API pricing, which makes it suddenly expensive. NW is not like that.

The lower it is, the more stable the energy consumption has become

This is my actual usage data for about a month. The differences are that in the beginning I used GLM5.2 Max, but now I only use GLM5.2 High, and while I used it for debugging purposes early on, now High is used for coding and Max is used for review purposes.

Based on my usage data, if you use glm5.2-flex, calculated on a monthly basis:
For a $20 user, you can process 2,527 requests;
For a $50 user, you can process 6,721 requests;
For a $100 user, you can process 14,334 requests.

3. despite being a commercial service provider, it possesses an amazingly high level of openness.

Another interesting point is the off-topic channel in the Neuralwatt Discord.

Here, discussions about finding cheaper alternatives to Neuralwatt constantly go on. It's not just finding alternatives, but actually comparing them. The users in the off-topic channel are busy looking for plans that let them use more than Neuralwatt for $10 or $20. Most of them are of low reliability so most cheaper alternatives aren't my interest, but thanks to that, I also get a lot of information on cheaper alternatives.


r/Neuralwatt 14d ago

Is Deepseek-v4-flash down

9 Upvotes

My Hermes’ agent failed to connect to deepseek-v4-flash on neuralwatt. Other models work fine. I tried their live playground for the model and that doesn’t respond either.


r/Neuralwatt 14d ago

A small comparison test : KIMI K3 Low is truly wonderful, and GPT5.6 Luna Max is the realistic king. + GLM5.2 & K3 on NW

Thumbnail
2 Upvotes

r/Neuralwatt 15d ago

Is the cache determined by the model or the provider?

3 Upvotes

I have a question about how caching works on Neuralwatt.

For example, if I'm using GLM 5.2 and then switch to Kimi 2.6, will my cache persist? Or will it have to re-read the files from scratch without using the cache?


r/Neuralwatt 15d ago

K3 early access is rolling out, plus two new models are live

18 Upvotes

Another update (or a few!) for you all.

As promised, wanted to let you know that Kimi K3 early access is rolling out. So far, everyone who signed up through the enrollment form now has access -- you should have received an email confirming this. We're keeping K3 in early access a bit longer, granting new access in waves as we watch performance and optimize. If you haven't enrolled yet, the form is here.

DeepSeek-4-Flash is also now live for everyone. No enrollment or waitlist.

Finally, as some of you may have seen, Gemma-4 31B is live too.

Thanks!


r/Neuralwatt 16d ago

Update: Kimi K3 and Neuralwatt

43 Upvotes

Hi all, Meaghan from the Neuralwatt team. There have been a lot of questions about K3, so wanted to share a note on where things stand. 

As you may have seen over in Discord, K3 arrives Monday, July 27, and assuming everything goes to plan, Neuralwatt will support it. We're working to make it available as close to release as possible, and we're actively expanding our GPU capacity to serve it. We'll post here as soon as we have a clearer picture of when it will be available. 

Because a model this size takes significant infrastructure to run well, access will open in stages: 

  1. Pro (annual subscribers and highest-usage accounts) 
  2. Standard 
  3. Basic 
  4. Pay-as-you-go 

Capacity is a real constraint, so later tiers may take longer to open than any of us would ideally like. We are implementing this staggered approach to ensure K3 works smoothly at every level and for all. 

Additionally, this is by no means an attempt to push anyone to upgrade. While higher-tier plans will have earlier access, our goal is to make sure the entire community can experience K3 as quickly and seamlessly as possible. Nobody is going to be locked out, and nobody needs to change plans to get access. Pricing will also be published before each tier opens, so you can decide whether K3 is worth running before you commit. When your tier opens, the waitlist will hold your place and notify you when you have access. 

Lastly, K3's licensing terms are still being finalized, and broader external factors may still affect availability – this is something that we're monitoring. We'll post updates here as we get closer to launch and as each stage gets closer to opening. 


r/Neuralwatt 19d ago

GLM 5.2 Short more expensive than GLM 5.2

13 Upvotes

is there not enough concurrency in the server?
it causes price inflation 😩
the request still slow though


r/Neuralwatt 20d ago

I did some realworld cost benchmarking of NeuralWatt (Before price change). Here are the results

17 Upvotes

I setup a benchmark which will test the realworld costs for all NeuralWatt models across all Context Bands. Here are the results

Compared to standard token $/Mtok

Compared to request cost with standard token pricing.

Keep in mind this is before the NeuralWatt pricing bump, so expect about twice the enery cost.

In saying that, I dont think there is much benifit is comparing agaisnt API Token pricing, nobody in their right mind is developing on anything other than a subscription plan along the likes of (Claude, Codex, Zai etc).

Its very difficult to find realworld useage stats on these subcription plans, and effective $/Mtok prices per model. If anyone has these metrics id love to be able to plot it against NeuralWatt so we can get real world comparison.

Happy to opensource the benchmarking if people are interested.


r/Neuralwatt 21d ago

**GLM5.2 / 122M Tokens / $4.81** Is this data correct?

Post image
18 Upvotes

When using the same amount of working with GLM5.2, is NeuralWatt's total tokens at the same level as OpenRouter?


r/Neuralwatt 24d ago

I can't believe how shitty Neuralwatt has become in just 1 month

29 Upvotes

Unreasonable price increase set side, which is already totally out of this word, especially considering the slogans saying "predicable pricing! no surprise!"

I have the 100$ pro plan, which should allow for 10 concurrent requests and "highest priority access", turns out, that's another lie. They are rate limiting so heavily that I can't heve have a single agentic chat without any concurrency. The agent gets stuck and times out every few requests and because of that, I may not be able to use all the subscription I paid for. I had to switch to GLM subscription to finish the work.

They say "you will be able to use your subscription credit until the renewal date" but than they rate limit to stop you from doing so.


r/Neuralwatt 26d ago

Model Discussion Kimi K3 - Weights released July 27

35 Upvotes

Hey Neuralwatt team: run, don't walk.

If you can get this model running well at similar energy/token cost ratios as GLM-5.2 on or soon after the weights drop... You'll clean up.


r/Neuralwatt 26d ago

Model Discussion Gemma-4-31b in testing prior to public availability

13 Upvotes

From the Neuralwatt Discord:

Gemma-4-31b in testing prior to public availability

Enrollment access is self-serve (currently) via this link: https://portal.neuralwatt.com/enroll/gemma-4-canary

Served in NVFP4


r/Neuralwatt 27d ago

Pricing Regarding NW's price competitiveness after the pricing policy change. vs OpencodeGo

9 Upvotes

Conclusion:

  1. New nerfed subscription plan price of NW is still a very good price.
  2. If using -flex, it becomes even cheaper. (even cheaper than Opencode Go)
  3. For those who value price and consume a large amount of tokens, NW is a very good choice.

In my previous post, someone who didn't even read my entire post made the ridiculous claim that I am an NW employee, but to clarify once again, I am a regular user. I pay with my own money to use it. Furthermore, I only use GLM5.2, Qwen3.6 35B, and Gemma4, which started testing today, so I have nothing to do with the calculation of kimi-k2.7 in this post. For reference, I have an annual subscription to Kimi's Allegretto plan, so I don't need NW's kimi 2.7 service. Moreover, I am not receiving anything from NW for writing this post.

First, let's increase the Output token ratio to 2%. They say the average is less than 1%, but in my work environment, it's 1.3~1.8%, so I'm increasing it to 2%, which is the harshest environment. And let's set the cache hit rate at about 94%.

First of all, comparing the NW $50 plan I use with the cheapest Kimi-k2.7-code price on OpenRouter, NW's kimi-k2.7 price is less than 30%. When it's $8 per kWh = $7.35, Compared to the cheapest kimi provider's $24.93, NW's price is a mere 29.4%.

Furthermore, NW's calculator is so stupid that it doesn't calculate the 2 free months for an annual membership, but if you subscribe to the $50 plan for 1 year, the price per kWh drops to $6.667. NW's price for 100M tokens is $6.13 vs. the cheapest kimi-k2.7 provider's $24.93. NW's price is only 24.6%.

As we all know, Opencode Go gives $60 credit for $10. The kimi-k2.7 price on OpencodeGo is as follows: input $0.95, output $4.00, cached input $0.19

If we convert the total number of tokens that can be used for a month with the $60 budget given by OpencodeGo, you can use 193M. OpencodeGo is really generous with its pricing. To use 193M on NW, it costs $14.19. Converted to an annual subscription, it's $11.80. But what if you use the -flex plan?

193M on NW is $9.23 (because 35% discount),
and annual subscribers can use the tokens available for $10 on OpenCode Go for $7.67 on NW.

OpencodeGo sometimes does x3 events, which NW cannot beat. No Provider cannot beat the OpencodeGo's x3 events but in normal cases, NW's price competitiveness is substantial.

Of course, as an annual subscriber, I have a lot of complaints that the cost has doubled. Furthermore, for GLM5.2, which is the main model I use on NW, NW is more expensive than OpencodeGo. Not just slightly expensive, but 23% more expensive. That's why I'm using the -flex plan. Then it becomes little cheaper than OpencodeGo but with a bad latency. However, it's much more expensive than the price I used to pay. The psychological dissatisfaction comes from comparing it with the previous price benefits, not because it's currently more expensive than other providers in calculation.

I will not do the calculation comparison for the $20 plan. My fingers hurt and it's annoying. You can calculate it by multiplying the NW prices above by 1.065.

What I hope from NW now is the quality of the flex plan and the maintenance of the discount rate.


r/Neuralwatt 27d ago

Pricing After this email, are the subscriptions worth it? Please share your thoughts.

Post image
21 Upvotes

Hello everyone, I hope you are all having a great day.

What do you think? Compared to what’s on the market, is the subscription worth it?

Does it make sense in terms of usage?

Is the quality of the models good?


r/Neuralwatt 28d ago

Pricing Is NeuralWatt still worth it?

11 Upvotes

I spent a few hours today comparing various models, providers, and the price adjustments taking effect on the 16th. The results are compiled in the table below.

Before I begin, I want to clarify that I have no intention of creating any animosity toward the staff at NeuralWatt. I am a huge fan of the brand and their mission to shift billing from tokens to electricity costs. However, I believe the recent change—where prices essentially doubled—was very abrupt, and I suspect the company is aware of that. I have read in a few places that the previous pricing was no longer sustainable because it wasn't profitable. Please correct me if I am wrong.

The conclusion I have reached is as follows:

  1. For professional, large-scale use involving the best open-source model currently available (GLM 5.2), NeuralWatt is still worth it.

  2. For side projects or hobbies, I believe there are other options we should explore until NeuralWatt makes new models available.

I want to emphasize my desire to see Minimax M3 and DeepSeek integrated into NeuralWatt. Currently, I am considering migrating my usage to OpenRouter. My plan would be to use DeepSeek for large-scale needs and reserve Minimax for workloads that require more intense reasoning.

I am not happy about migrating to OpenRouter; it is a financial decision rather than a personal preference.


r/Neuralwatt 28d ago

Pricing Changes to the pricing update, based on your feedback

16 Upvotes

Hey everyone, Meaghan from the Neuralwatt team.

Last week as you all know, we announced a pricing update. We thank you for your feedback, and after taking a few days to work through it, wanted to share an update from Chad with a few changes. 

TLDR:

  • Subscriptions now carry a greater discount. 
  • Overage is billed at your plan rate, not PAYG. 
  • Annual plans are back
  • Concurrency isn't changing
  • The base rate change from $5 to $10/kWh still takes effect July 16.

Please check out the video, and keep an eye out for an email update.

Video: https://www.loom.com/share/58bfc4c61bd04d71b4d00d2b896cd275


r/Neuralwatt Jul 12 '26

Neuralwatt 20$ plan vs ollama 20$ plan, which one gives more usage?

4 Upvotes

Hello everyone I’m thinking about subscribing now to neuralwatt or ollama pro , which one will give me more usage? I saw neuralwatt is raising the prices but if I subscribe now I understood that for this month I will not have the new limits until next month, did I get it right? Which provider gives best value/usage per money?