The AI arms race reached a fever pitch on Monday after Chinese language e-commerce and cloud supplier Alibaba known as into query America’s technological lead with the launch of Qwen 3.8-Max, a 2.4 trillion-parameter mannequin that goes toe-to-toe with the perfect fashions from Anthropic and OpenAI.
The brand new mannequin comes simply days after the launch of DeepSeek V4 Flash 0731, which, in response to unbiased benchmarks by Synthetic Evaluation, performs inside a single level of OpenAI’s budget-friendly GPT-5.6 Luna whereas costing 40 p.c much less per activity. What’s extra, at simply 284 billion parameters, it’s sufficiently small to run on comparatively modest enterprise servers and workstations.
Chinese language mannequin devs like Moonshot, Alibaba, and DeepSeek at the moment are attacking their American counterparts on each value and efficiency. The pincer motion comes as US mannequin devs like Anthropic and OpenAI stoke fears over the origins and security of China-made AI fashions.
In a current weblog submit, Anthropic CEO Dario Amodei insisted he’s not against open fashions, simply ones made in China, ones distilled from proprietary fashions, and ones that don’t meet rigorous security metrics. In different phrases, something that really competes with Anthropic’s personal fashions.
The protection bit is especially disingenuous, as the corporate’s fearmonger-in-chief has gone out of his option to stoke fears amongst US authorities officers. Proprietary fashions could be managed, however open weights, as soon as launched within the wild, are not possible to claw again.
However neither Amodei’s feedback nor commitments from main American and European tech giants change the truth that China is offering the one significant competitors within the open weights area.
China dominates right here. Talking on CNBC Monday, Clément Delangue, CEO of Hugging Face, the most important and most influential mannequin repo on this planet, mentioned as a lot.
“They’re clearly dominating on open fashions proper now, and I wouldn’t be stunned if they begin dominating on the frontier both by the top of this yr or subsequent yr on the fee of progress,” he advised the monetary TV information community.
America’s most succesful open weights mannequin, Inkling, is simply shy of a billion parameters, and nonetheless can’t sustain with DeepSeek’s best. Fashions like Moonshot’s Kimi K3 and Z.ai’s GLM 5.2 are in a completely completely different orbit.
In reality, an enterprise’s solely credible options to proprietary fashions and their doubtful safety insurance policies are Chinese language fashions. DeepSeek, Alibaba, Moonshot, MiniMax, and Z.ai aren’t going to move that chance up.
Alibaba lets its strongest mannequin free on the world
Of the Chinese language mannequin devs, Alibaba is arguably essentially the most like its American competitors. Its open-weight fashions are properly regarded, and their vary of sizes and permissive licensing have made them enticing for fine-tuning application-specific techniques. Nonetheless, like Google and OpenAI, its prime fashions have been locked behind an API – till now.
Seizing the second, Alibaba is making its most succesful mannequin weights accessible for obtain for the primary time with the launch of Qwen 3.8-Max.
The blog post incorporates the standard array of vaguely intelligible bar charts showcasing how the mannequin compares to OpenAI and Anthropic’s personal fashions, in addition to a slew of demos that might have made Billy Mays proud.

If you would like specifics, we suggest testing the launch weblog here. You don’t need to look that carefully to get what Alibaba is promoting: something OpenAI and Anthropic can do, we are able to do as properly, if not higher, cheaper, and on the {hardware} you personal.
With that mentioned, 2.4 trillion parameters is a slightly huge carry for many enterprises, doubtless requiring 48-64 Nvidia B200-class GPUs if deploying in a customer-facing capability. For inner workloads, 8-16 B300 or AMD MI355X GPUs can be satisfactory.
If that’s somewhat wealthy on your blood, Alibaba hasn’t forgotten its roots and can be releasing a 27-billion parameter model of the mannequin alongside the Max variant.
With benchmarks, mannequin devs can often discover a assortment of checks that paint their mannequin in a constructive mild. And unsurprisingly, Synthetic Evaluation’ personal Intelligence leaderboard tells a barely completely different story than Alibaba’s, with Qwen 3.8-Max matching Anthropic’s less-capable, however nonetheless formidable, Claude Sonnet 5.

Past the advertising and pitch demos, Qwen’s weblog submit is surprisingly brief on element. What we do know is it’s a multi-modal combination of consultants (MoE) mannequin. Meaning of the two.4 trillion complete parameters, solely 95 billion are literally used to generate tokens for anybody request. We will additionally assume that the mannequin makes use of the identical hybrid Transformer+Mamba structure as earlier Qwen fashions to be able to keep efficiency throughout massive contexts. Talking of which, the mannequin will assist context home windows as much as 1 million tokens, although it’s nonetheless not clear whether or not that depends on strategies like rope scaling to increase it or not.
The mannequin is at present accessible by way of Alibaba’s API service QwenCloud for $2 per million enter tokens, and $6 per million output tokens. Cached tokens are charged on a sliding scale with $0.25 charged per million implicit cached tokens, $0.17 per million express cache token reads, and $2.5 per million express cache tokens created.
For comparability, Anthropic’s Claude Sonnet 5 will set you again $2/M enter tokens, $0.20/M cache hits, and $10/M output tokens. And Sonnet 5 pricing is about to extend 50 p.c beginning September 1.
In the meantime OpenAI’s GPT 5.6 Luna, which falls simply behind Qwen 3.8-Max and Sonnet 5 on the Synthetic Evaluation leaderboard, will run you $0.20/M enter, $0.02/M cached enter, $0.25/M cached write, and $1.20/M output tokens for brief context lengths underneath 272,000 tokens and double that for jobs exceeding that context.
The mannequin weights are slated for launch on well-liked mannequin repos, together with Hugging Face, beginning subsequent week.
DeepSeek undercuts OpenAI with flashy new V4 refresh
Whereas Alibaba joins Kimi K3-maker Moonshot.AI’s assault on frontier fashions, DeepSeek has taken a really completely different tack with its newest open weights mannequin: squeeze each ounce of efficiency from as few weights as potential.
The result’s DeepSeek V4-Flash-0731, a 284-billion parameter mannequin that may match into round 142 GB of GPU reminiscence (at FP4). Meaning enterprises can simply run this mannequin at scale on a single system.
And regardless of its smaller stature, the Flash mannequin truly outperforms the 1.6 trillion-parameter DeepSeek V4 Professional by almost 14 p.c on Synthetic Evaluation’ Intelligence leaderboard. With that mentioned, the refinements made to DeepSeek V4 Flash will little question discover their method into the Professional mannequin earlier than lengthy.
Like Qwen 3.8-Max, DeepSeek V4 Flash undercuts OpenAI and Anthropic on pricing – this time by a substantial margin. For API entry, DeepSeek is at present asking $0.14/M enter tokens, $0.0028/M cached tokens, and $0.28/M output tokens.
Nonetheless, in an agentic world stuffed with reasoning fashions, API pricing doesn’t paint an entire image. A mannequin might seem cheaper, but when it consumes twice as many tokens as one other increased priced mannequin, it could not truly be inexpensive. Due to this, it is vital to take a look at how effectively the mannequin solves actual world issues.

Alarmingly for the US LLM makers, in response to Synthetic Evaluation, DeepSeek V4-Flash isn’t simply cheaper per token; additionally it is extremely environment friendly at its job. In comparison with OpenAI’s GPT 5.6 Luna, which is among the many best and least costly fashions in Sam Altman’s present lineup, DeepSeek’s newest mannequin is a full 40 p.c inexpensive, with a value to resolve of simply three cents versus 5 cents.
That distinction might really feel small, but it surely’s value remembering that for vibe coders consuming tens or a whole lot of tens of millions of tokens a day, that distinction provides up fairly rapidly.
One of many secrets and techniques to DeepSeek’s effectivity is the mixing of DSpark speculative decoding instantly into the mannequin weights. We’ve explored speculative decoding prior to now, however in a nutshell, it entails utilizing a smaller draft mannequin to foretell the outputs of a bigger, extra succesful one. When it really works, inference efficiency will increase. When it guesses mistaken, it falls again to the bottom mannequin.
Extra importantly, as a result of the smaller mannequin is simply predicting the output of the bigger one, it’s fully lossless and subsequently requires no compromise by way of efficiency. Alibaba and others have applied comparable speculative decoding mechanisms, like multi-token-prediction (MTP), because of this. With DSpark, DeepSeek claims it may extract 57–85% extra per-user velocity on the very same {hardware}, which is a formidable declare, and one which a minimum of in our testing rings true.
Your individual private frontier mannequin?
DeepSeek V4 Flash is simply sufficiently small that working it at house is fully potential in case you’ve received some deep pockets, or a heck of plenty of reminiscence mendacity round.
Testing on a 128 GB DGX Spark, we had been capable of get the mannequin working at a good 128,000 token context window utilizing Unsloth’s IQ3-XXS quant in Llama.cpp.
Three bits per weight provides us simply sufficient room to pack the DSPARK draft mannequin into reminiscence, but it surely’s additionally a bit extra compression than we usually suggest for homelab use. Due to this, Unsloth warns that the mannequin’s outputs might present some indicators of high quality loss, but it surely does run.
Unsloth has a full information on get the mannequin up and working, assuming you’ve received beefy sufficient {hardware}.
We admit, the DGX Spark will not be an affordable field. At $4,699, it’s squarely in workstation territory, as is the $3,999 Ryzen AI Halo we checked out early final month. However the truth that you’ll be able to even ponder working a mannequin like DeepSeek V4-Flash at house is spectacular in its personal proper.
The unique IBM PC kitted out with a monitor and diskette drive value about $3,735 in 1981. Adjusting for inflation, that works out to about $13,700 in immediately’s cash. If historical past repeats itself, the {hardware} essential to run fashions like DeepSeek V4-Flash ought to change into far more accessible inside the subsequent decade.®
Source link

