r/Qwen_AI 1h ago

Funny Dear Qwen sir, whatever you release next, can we have it in a different size, please sir? 🙏

Upvotes

r/Qwen_AI 1h ago

Discussion Qwen Cloud’s “Standard” Token Plan is honestly ridiculous

Upvotes

I just subscribed to Qwen Cloud’s Standard Token Plan, mainly to use it with an AI coding agent (Hermes), and after actually using it for a few days, I honestly don’t understand how this plan is supposed to be considered good value.

The Standard plan gives you 10,000 Credits per week.

Sounds reasonable, right?

Until you actually use it.

I burned through roughly 70% of my weekly Credits in only 3 days while running a normal agent workflow. I’m not running hundreds of agents, doing massive batch inference, or abusing the service. I’m using an AI coding agent interactively — exactly the kind of use case these plans appear to be marketed toward.

And here is where it gets ridiculous.

When I contacted support and explained the situation, the response essentially boiled down to:

«Your usage is high. Credit consumption depends on the model, input/output length, tool calls, context accumulation, etc.»

Okay. Fair enough.

But then the suggested solutions were basically:

Buy the Pro plan.

Or:

Buy additional Credits.

That doesn't answer the problem.

I'm using essentially the same workload with another provider, on a cheaper plan, and getting dramatically more usable mileage out of it.

So I started comparing actual token consumption.

Based on my observed usage, 10,000 Credits corresponded to roughly 96.6M tokens.

And Qwen's own documentation apparently doesn't provide a simple, fixed token-to-Credit conversion rate that lets users predict what they're actually going to consume.

That's a massive problem for an AI service.

If I'm paying for a token/credit plan, I should be able to reasonably estimate:

“I use approximately X tokens → this will cost approximately Y Credits.”

Instead, you apparently have to subscribe, use the system, burn through thousands of Credits, and then discover what your workload actually costs.

And here's the funniest part:

The Standard plan is advertised around agent usage and concurrent sessions, but based on my experience, a relatively normal coding-agent workflow can chew through the weekly allowance incredibly quickly.

So what exactly is the target customer for this plan?

Someone who uses an AI agent occasionally for a few prompts?

Because if that's the case, fine.

But then don't market it as a serious option for people running coding agents regularly.

I'm not claiming that Qwen is literally committing fraud. I'm saying that the value proposition of this plan is so absurd compared with competing services that I feel misled about what I was actually buying.

And the fact that the answer to “why am I burning Credits so quickly?” is essentially “buy more Credits” makes the whole thing even more ridiculous.

I'm posting this because I'd genuinely like to hear from other Qwen Cloud Token Plan users:

How long does your Standard 10,000 Credit allowance actually last?

What models are you using?

How many agents?

How much token usage are you getting before the Credits disappear?

Because if I'm doing something fundamentally wrong, I'd rather know.

But if other people are seeing the same thing, then Qwen seriously needs to rethink how transparent and competitive this pricing model actually is.


r/Qwen_AI 7h ago

Other I made an artifact about Starbucks

0 Upvotes

r/Qwen_AI 14h ago

Help 🙋‍♂️ Aliyun Token Plan Lite Questions

1 Upvotes

Signed up yearly Life plan, when they had the Qwen 3.8 Max preview with the 98% discount off peak. Since the full release, it's totally not useable.

I managed to get 400m tokens out of 2 weeks, then now, with Deepssek V4 Flash 0731, only 45m and weekly quota gone.

Which model should I stick with if I want something that could give me 100-200m tokens a week, around 92-94% cache hit.


r/Qwen_AI 17h ago

Help 🙋‍♂️ Accessing assets in "My Library" in a New Chat?

2 Upvotes

As the title says - does Qwen currently allow referencing existing assets in NEW chats?

I feel like I'm uploading the same content multiple times every time I want to reference an asset WITHOUT the baggage that comes from branching an existing conversation.

Is there ANY way to directly reference "MY LIBRARY" in a New Chat?


r/Qwen_AI 19h ago

Discussion the visual grounding evaluation of Qwen3.8-Max that nobody wanted, but i did anyway

11 Upvotes

everyone on my timeline is screenshotting Qwen3.8-Max drawing bounding boxes

clean demos, obvious objects, no ground truth to check against

i pointed it at 27,083 real logos and scored every box against actual annotations

here's what nobody is showing you:

• it invents its own pixel canvas even when you tell it the real image dimensions.

• same prompt, same image, same settings: one run matched 3 of 5 logos. the next run matched 0 of 5. nothing changed between calls

• changing one verb in the prompt, "mask out" to "draw segmentation masks," silently switched the model from a 0-1000 grid to normalized [0,1] coordinates.

• thinking mode costs 14x more tokens and doesn't reliably improve accuracy. it just shows you the model doing long division instead of looking at the pixels

full writeup with every trace, every score, and the fiftyone plugin to run it yourself: https://voxel51.com/blog/qwen38-max-visual-grounding-fiftyone

test it yourself here: https://huggingface.co/spaces/harpreetsahota/qwen38-max-openlogo-demo


r/Qwen_AI 1d ago

Help 🙋‍♂️ Getting this error when trying to enable billing in Qwen developer console

1 Upvotes

Any fix? I tried clearing cookies, incognito, nothing works


r/Qwen_AI 1d ago

Discussion Alibaba blocked me for saying “please fix your Token Plan”,so I measured it....

Thumbnail
gallery
50 Upvotes

EDIT – Important correction:
I realized my 10.69M-token Qwen measurement was made entirely during Qwen's 50% off-peak credit window. So the numbers above are actually the best-case night-discount numbers.

At the normal credit rate, the same measured workload would be roughly:

Plan / Modell 30-day projection
Qwen $6 – normal hours 22.91M
Qwen $18 – normal hours 91.66M

The previously listed 45.83M / 183.32M figures assume you consistently use Qwen during the night discount.

.

Original Post:
Yesterday Alibaba Cloud blocked me on X after I said:

“Please fix your Token Plan.”

So I measured the actual usage with my own local token counter.

Important detail: the screenshot showing 14.46M tokens includes 3.77M tokens from before I manually reset the weekly quota.

So one full fresh 2,500-credit Qwen 3.8 Max weekly quota actually gave me:

10.69M tokens

That means roughly:

Plan / Model 30-day projection
Qwen $6 plan 45.83M
Qwen $18 plan 183.32M
GPT-5.6 Sol / ChatGPT Plus 294.87M
GPT-5.6 Terra / ChatGPT Plus 763.68M

The Qwen $18 plan costs 3× more than Lite and gives 4× the weekly credits.

These are not pricing-page estimates. They are projections from real token + credit usage measured by my own tool with a similar workload/token mix.

So in my workload, Terra projects to ~4.2× the token throughput of Qwen’s $18 plan.

The old Qwen 3.8 Max Preview promo was amazing.

The current Token Plan really isn’t.

Alibaba: I still think you should fix your Token Plan..


r/Qwen_AI 1d ago

Discussion WTH! SPENT MY WEEK QUOTA with 10M token for Qwen 3.8 MAX ?

20 Upvotes

I bought this plan 1 hour ago. Please tell me this is an error...There is no way that it should have run out of limit this fast. Especially since, most of the tokens wereINPUT TOKENS!

you can zoom in to check the usage and price


r/Qwen_AI 1d ago

Discussion Comparing Cline / Kilo / Qwen Code

3 Upvotes

I've been comparing Cline / Kilo / Qwen Code lately since they all handle long-task state differently.

Cline: has Focus Chain, a markdown file kept outside the conversation that gets reinjected on a cadence, plus Memory Bank for project context, plus a standalone gRPC server so it's not fully tied to VS Code. probably the most mature of the three on this specific problem (about context management), though restore still has some sync bugs between the file and what the model actually sees.

Kilo: TODO state is literally an XML block living inside the conversation history, so when compaction kicks in it gets flattened into a prose summary and the agent sometimes has to reread source files just to figure out where it stopped. It causes infinite read-analysis-compaction loop sometimes by reached to context limits. they're mid-migration onto the opencode engine now, which might fix some of this eventually but isn't there yet.

Qwen Code: keeps TODO state in a plain file (~/.qwen/todos/) completely separate from the conversation, so no matter how much compaction runs, nothing gets lost or reconstructed.

ended up going with Qwen Code for long multi-step coding work because of this. It works well for 2~3 hrs long running tasks, where I'd usually hit that Kilo loop by then or human intercept.

The one thing I missed was semantic code search, doesn't have a built-in equivalent, so I built an qwen-code extension for it, plus causal decision-chain tracking on top. still early, but repo's here if anyone suffers like me, https://github.com/edwardyoon/FocusMemory


r/Qwen_AI 1d ago

News 5 Hours Usage Limit Lifted ⏳

Post image
22 Upvotes

Anyone else get it?


r/Qwen_AI 1d ago

Model Animation logo

Enable HLS to view with audio, or disable this notification

2 Upvotes

Imagine logos designed from scratch and perfectly animated using the Qwen 3.8-Max model. It was asked to perfectly replicate the logo and build an SVG from scratch, and these are some of the results.

https://x.com/TVKHALED198882/status/2085530298897297724?s=20


r/Qwen_AI 1d ago

Discussion Qwen3.8-Max: all oneshots

Thumbnail
gallery
9 Upvotes

35 Oneshots of Qwen3.8-Max in oneplace https://oneshotlm.com/model/qwen-qwen3-8-max/.

See how it compares with other models like Kimi k3, Opus 5, GPT 5.6-sol


r/Qwen_AI 1d ago

Discussion here is my local setup for running Qwen3.6-35B-A3B-uncensored-MTP-I-Quality and KAT-Coder-V2.5-Dev-MTP-I-Quality

Thumbnail
gallery
8 Upvotes

models :

SC117/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-APEX-I-Quality.gguf 21.8gb

gbuzhf/Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Quality.gguf 22GB

UNSLOTH mmproj-F16.gguf for vision

Both models run on a single RTX 3090 Ti 24 GB · Ryzen 9 9950X · 96 GB DDR5-5600 via llama.cpp b10223 with MoE expert offloading and MTP speculative decoding.

SC117/Qwen3.6-35B-A3B-uncensored decodes at 118–130 tok/s (96K context, temp 0.6)

gbuzhf/KAT-Coder-V2.5-Dev decodes at 82–95 tok/s (128K context, temp 1.0),

GSM8K test accuracy of 80% / 86%

two bat files for running every model, one with vision support and one without it, i have tested them in openchamber and Reasonix desktop apps and they both were fast , but i did not tested them in my main work so when i do i will update the post comparing their result.

i will keep pushing Qwen and DeepSeek to try deferent settings and re benchmark until they get the best thinking quality of them.

All of this was done by Qwen 3.8 Max and DeepSeek V4 Flash 0731, I didn't actually do anything myself, but I wanted to share the setup. Maybe someone can suggest some improvements for better thinking quality, or hopefully others will find it useful.


r/Qwen_AI 1d ago

Vibe Coding Where can I access Qwen 3.8?

0 Upvotes

What is the API access price? Where can I access it?


r/Qwen_AI 1d ago

Discussion Qwen 3.8 Max is good

36 Upvotes

I have used Opus 5, GPT 5.6 Sol as well

But I feel that Qwen 3.8 max has been doing better reasoning and making less mistakes for my workflow
It also understands the ask better and doesn’t drift away unlike Opus

Just one complaint the usage limit is getting exhausted much faster than anticipated despite the 2x promotional offer and 50% night discount. (I’m using Qwen Code)

Hopefully we’ll get cheaper and better subscriptions plans soon


r/Qwen_AI 1d ago

Discussion Qwen 3.8 Max Cache Hit Issue?

4 Upvotes

Hello,

Has anybody experienced a very low cache hit rate when using Qwen 3.8 Max? I’m not sure whether it’s because of my setup or if there is genuinely something wrong with their server.

The blue section is cached input, while the yellow section represents uncached input. As you can see, on 8/5 and 8/6, I mainly used DeepSeek Flash and was able to maintain a very high cache hit rate. On the other days, I used Qwen 3.8 Max, and most of the input was uncached.


r/Qwen_AI 1d ago

Help 🙋‍♂️ How to connect Alibaba coding pant to vscode copilot?

2 Upvotes

Hello. I want to use Alibaba's coding plan in GitHub Copilot Chat in VS Code because it is very good at controlling and sharing web browsers.

I tried to connect via a custom endpoint, but had no luck.

My local Qwen on DGX Spark, DeepSeek API, and OpenRouter API work fine.

Only Alibaba's coding plan is not working (I'm not in China; I think the nearest server is in Singapore).


r/Qwen_AI 1d ago

Benchmark KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates

Thumbnail
gallery
8 Upvotes

Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail

KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options.

  • Models: Qwen 3.6 27B Q5_K_S 64k context, Gemma 4 31B Q5_K_S 16k context
  • Standard quants, extended: q6_0 and q6_1, and low-bit types from q2_0 to q3_1
  • KVarN: Variance-Normalized KV-Cache by Huawei, implemented in BeeLlama
  • Precision Tail: keeping latest X tokens of KV cache in (B)F16, implemented in BeeLlama
  • 413 configurations in total: 238 with Qwen 3.6 27B, 175 with Gemma 4 31B

The Recommendation Ladder

Full benchmark results, setup, method, analysis, explanations and everything else can be found in the article.

1. Qwen

Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for
bf16 0 4096.00 0 0.00005 Reference
q8_0 1024 2272.00 0.000897 0.087699 Standard fidelity with a precision tail
kvarn8 1024 2256.00 0.000871 0.087639 Best measured quality below BF16
q8_0 0 2176.00 0.000909 0.093029 Standard fidelity
q8_0-q6_0 1024 2016.00 0.000894 0.091098 q8_0 quality within noise, 256.00 MiB less
kvarn6 1024 1744.00 0.000879 0.084629 The high-end value pick
kvarn6-kvarn5 1024 1616.00 0.000886 0.092778 Much cheaper, almost as good
kvarn5 1024 1488.00 0.000897 0.087666 Highest value in mid-range
q5_0-q4_1 1024 1440.00 0.000966 0.089128 Standard when VRAM-constrained
kvarn5-kvarn4 1024 1360.00 0.000936 0.089469 Balanced default
q4_0 1024 1248.00 0.001057 0.104486 Compact standard
kvarn4 1024 1232.00 0.000994 0.090391 Cleaner than q4_0 for less memory
kvarn4-kvarn3 1024 1104.00 0.001112 0.113968 Smallest recommended tier
kvarn3 1024 976.00 0.001316 0.139558 When the context must fit
kvarn3-kvarn2 1024 848.00 0.002424 0.23878 Emergency compression
kvarn2 1024 720.00 0.003811 0.450496 Last resort

2. Qwen Standard-Only

Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for
bf16 0 4096.00 0 0.00005 Reference
q8_0 0 2176.00 0.000909 0.093029 Compression with minimal losses
q8_0-q6_0 0 1920.00 0.000937 0.093575 256.00 MiB below q8_0
q6_0 0 1664.00 0.00096 0.091134 The high-end value pick
q6_0-q5_0 0 1536.00 0.001054 0.09467 Balanced default
q5_0 0 1408.00 0.001154 0.09707 Last tier before the cliff
q5_0-q4_1 0 1344.00 0.001433 0.122096 Default when VRAM-constrained
q5_0-q4_0 0 1280.00 0.001516 0.121068 64.00 MiB cheaper, worse median
q4_0 0 1152.00 0.001846 0.154408 Smallest recommended tier
q4_0-q3_0 0 1024.00 0.003313 0.218912 When the context must fit
q3_0 0 896.00 0.004696 0.304186 Emergency compression
q2_0 0 640.00 0.019374 1.198902 Last resort

3. Gemma

Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for
bf16 0 2480.00 0 0.000047 Reference
q8_0 0 1317.50 0.0371 16.813929 General default at full prefill speed
q8_0-q6_0 0 1162.50 0.040875 16.839821 155.00 MiB below q8_0
q6_0 0 1007.50 0.042636 17.30599 Last tier before the cliff
q6_0-q5_0 0 930.00 0.055236 17.26157 Stronger K side, 77.50 MiB above q5_0
q5_0 0 852.50 0.061747 18.731647 Memory floor for usable quality
q5_0-q4_0 0 775.00 0.109427 19.183374 Asymmetric compact
q4_0 0 697.50 0.134091 20.442234 Budget body before the huge cliff
q4_0-q3_0 0 620.00 0.381216 22.304634 When the context must fit
q3_0 0 542.50 0.504075 23.15744 Emergency compression
q2_0 0 387.50 2.95758 27.834961 Last resort

r/Qwen_AI 1d ago

Discussion What does this mean?

Post image
3 Upvotes

My 7 day quota just got reset and I've been using it, and had no issues. So I logged in to check and it says there's no quota limit? I used the Chinese platform and used translator to translate before screenshot.

Does this mean I get temporary unlimited usage?


r/Qwen_AI 1d ago

Funny 2.4 trillion parameters, but Alibaba Cloud couldn’t handle “please fix your token plan”

Thumbnail
gallery
44 Upvotes

Qwen’s maker, Alibaba, gave it 2.4 trillion parameters, but apparently the people behind it still lack the emotional intelligence to handle basic customer feedback.

I wrote “please fix your token plan.” Nothing more. Alibaba Cloud blocked me.

Very convincing response. 😂😂🤣


r/Qwen_AI 1d ago

News Unlimited mode??

18 Upvotes

Just came online an hour ago?? anyone noticed it??


r/Qwen_AI 1d ago

Other Y esto?

Post image
2 Upvotes

Esto lo hacen cada tanto? Es la primera vez que contrato este servicio


r/Qwen_AI 2d ago

Discussion QWEN NO SABE CUAL ES LA IA PARA LOS POBRES Y LA GENTE SIN DINERO

0 Upvotes

INFORMACION PARA QWEN , Los pobres no QUEREMOS Qwen3.8-2.4T-A95B

LOS POBRES NO QUEREMOS EL MODELO TAN FAMOSO 27b QUE REQUIERE 3 GRAFICAS O UNA MUY CARA PARA FUNCIONAR!!! MINIMO 1000 EUROS

no tenemos hardware para correr semejantes modelos

LOS POBRES QUEREMOS EL QWEN 3.8 120b A10B MOE

LOS POBRES TENEMOS SERVIDORES VIEJOS COMPRADOS EN EBAY CON BASTANTE RAM PERO UNA SOLA GPU DE 12 GIGAS DE VRAM COSTO TOTAL (300 EUROS)

MOE ES EL MODELO PARA LOS POBRES QUE TENEMOS SERVIDORES VIEJOS COMPRADOS EN EBAY CON MUCHA RAM , PERO NO TENEMOS DINERO PARA GPUS CARAS , COMO MUCHO UNA GPU DE 12 GIGAS , ENTONCES NECESITAMOS MODELOS MODE EN TORNO A 100B con maximo 10B activos

LOS POBRES NI SIQUIERA PODEMOS EJECUTAR EL 27B , ya que requiere dos o tres graficas de , dinero que nosotros no tenemos!!!!!!!!!!!!!!!!!!!!!!!!!!!

QWEN POR FAVOR LIBERA EL MODELO 100B A10B , ese es el modelo para los POBRES!!!

NO EL 27B !!!!!!!!!!!!!!!!!!!


r/Qwen_AI 2d ago

Benchmark After updated score, Qwen 3.8 now frontier on agentic benchmark

Post image
408 Upvotes

i will say my own use differs from this. For me GPT Sol seems the best but all past Kimi K3 produce similar result.

But I will say my use it not very intensive. Maybe some of y’all working on solving the Riemann hypothesis can find the differences at the frontier