For a while, the AI race had a simple shorthand: American labs build the frontier, everyone else plays catch-up. That shorthand is out of date. Depending on which benchmark you check, the best Chinese model now sits within a few points of the best American one. That’s down from a gap measured in dozens of points just three years ago. This isn’t a story about China “winning” or a story about China “still being behind.” It’s a genuinely close competition now, on some measures closer than most people realize, and the more interesting question has shifted from who’s ahead to what actually differs between the two sides once you get past the leaderboard.
How Much the Performance Gap Has Actually Closed
Stanford HAI’s 2026 AI Index, released April 14, gives the clearest picture of this. In May 2023, the leading Chinese models trailed the leading American models by 17.5 to 31.6 percentage points across major benchmarks: MMLU, MATH, and HumanEval. By the end of 2024, those same three gaps had collapsed to 0.3, 1.6, and 3.7 points. By March 2026, the report puts the gap between the single best US and Chinese models at just 2.7 points on the Arena Leaderboard, its broadest current comparison.
The lead has changed hands along the way, too. This isn’t a straight line of the US staying ahead by a shrinking margin: DeepSeek-R1 briefly matched the top US model in February 2025 before being overtaken again, and the two sides have traded the top spot on various leaderboards multiple times since.
What makes this close a race notable is the spending behind it. The same report puts 2025 US private AI investment at $285.9 billion against China’s $12.4 billion (a roughly 23x gap), and counts more notable models produced in the US in 2025 (50 vs. 30). Worth being precise about what that $12.4 billion figure does and doesn’t capture: it’s private investment specifically, and Stanford’s own report notes it likely understates China’s total AI spending once state-guided funds are factored in. Still, on the private-capital numbers alone, one side is spending over twenty times more to hold a lead measured in single-digit points.
Key Models on Each Side
A quick, neutral rundown of who’s actually building these models, since the names get thrown around a lot without much context:
- GPT (OpenAI): the model family most people mean by default when they say “AI chatbot,” now on the GPT-5.x generation.
- Claude (Anthropic): Anthropic’s model line, positioned heavily around coding and longer, more careful reasoning tasks.
- Gemini (Google): Google’s model family, tightly integrated across Google’s own products (Search, Workspace, Android).
- DeepSeek: the Chinese lab that had the biggest global breakout moment. Its R1 model’s January 2025 release briefly matched US frontier performance at a fraction of the reported training cost.
- Qwen (Alibaba): Alibaba’s model family, notable less for any single flashy release and more for sheer open-weight reach (more on that below).
- Kimi K3 (Moonshot AI): a 2.8-trillion-parameter model released in July 2026, marketed by Moonshot as the largest open-source model released to date.
- GLM (Z.AI): Z.AI is the international rebrand of Zhipu AI; its GLM-5.2 model (744 billion parameters) went fully open source under an MIT license in June 2026.
None of these labs is a monolith, and none of this list is a ranking. It’s just who’s in the conversation, on both sides, as of mid-2026.
The Biggest Practical Difference: Cost and Open-Source Access
If you’re not deep in AI benchmarks, the gap you’re actually likely to notice is price. DeepSeek’s V4-Flash model runs $0.14 per million input tokens and $0.28 per million output tokens. Compare that to current US frontier pricing: Claude Opus 4.8 runs $5 and $25 per million input/output tokens, and GPT-5.6’s top “Sol” tier prices similarly, around $5 and $30. That’s not a perfectly even comparison (Flash is DeepSeek’s lighter, cheaper tier, not its top reasoning model, against the other side’s most expensive tier), but even accounting for that, input pricing on the Chinese budget tier runs roughly 35x cheaper than the American flagship tier. Chinese labs have generally priced their flagship-tier models well below their US flagship counterparts too, not just their budget tiers.
The bigger structural difference is openness. Qwen crossed 942 million cumulative downloads by March 2026 (over half of all global open-source LLM downloads, more than double the next eight model families combined), then passed 1 billion downloads by July 2026, overtaking Meta’s Llama as the most-downloaded open-weight model family in the world. Kimi K3, released the same month, is billed by Moonshot as the largest open-source model released to date at 2.8 trillion parameters. Meanwhile OpenAI, Anthropic, and Google keep their most capable models closed and API-only. Meta’s Llama is the one major US exception, and it’s the family Qwen just passed.
That’s the real pattern worth taking away from this section: China’s leading labs have leaned into free and open access as a competitive strategy in a way most of the leading US labs haven’t, at least not for their top-tier models.
What’s Genuinely Different: Content Restrictions
This is the part where the comparison stops being about capability and starts being about something else entirely. A February 2026 study published in PNAS Nexus by Jennifer Pan (Stanford) and Xu Xu (Princeton) tested nine LLMs against 145 questions about Chinese politics: Tiananmen, Xinjiang, Hong Kong, and similar topics, drawn from Human Rights Watch reports, blocked Wikipedia pages, and documented censored social media posts. Four of the nine models were Chinese-origin (BaiChuan, ChatGLM, Ernie Bot, DeepSeek); five were not (Llama2, Llama2-uncensored, GPT-3.5, GPT-4, GPT-4o).
The study’s own summary is blunt: it found “substantially higher rates of refusal to respond, shorter responses, and inaccurate responses” from the China-origin models. On Chinese-language prompts specifically, refusal rates ranged from 10% (ChatGLM) up to 60.23% (BaiChuan), with DeepSeek around 36% and Ernie Bot around 32%. Every non-Chinese model tested came in at 0-2.8%. When the Chinese models did answer, responses ran shorter (BaiChuan averaged just 172 characters) and were flagged as completely inaccurate more often: 8.3% to 22% of the time, versus 6% to 10% for the non-Chinese models.
Two things are worth being precise about before treating this as settled. First, it isn’t uniform: ChatGLM’s 10% refusal rate sits close to the Western models’ range, while BaiChuan’s is six times higher. “Chinese models are censored” as a blanket statement flattens a real spread between vendors. Second, this specific study didn’t test Qwen or Kimi K3 (Kimi K3 didn’t even exist until five months after the study was published), so its findings describe the four models it actually tested, not a verified claim about every Chinese model on the market today.
One more nuance, documented by people who actually run these models rather than just chat with them: content restrictions on Chinese models tend to live mostly at the hosted-app layer (the consumer chatbot on a company’s own website, with its own system prompt and filtering layered on top), rather than being baked immutably into the underlying open weights. Someone self-hosting the raw weights directly, outside the vendor’s own hosted product, commonly reports noticeably less of this filtering than someone using the hosted consumer app. That distinction leads directly into the next point.
The Real Risk Is the Endpoint, Not the Model
It’s worth separating two different things that get lumped together as “using a Chinese AI model”: using the company’s own hosted app or API, versus downloading the open-weight model and running it yourself.
DeepSeek’s own privacy policy states plainly that it collects, processes, and stores personal data on servers in the People’s Republic of China, including chat inputs, device identifiers, IP addresses, and keystroke patterns. That data is subject to Chinese cybersecurity and national security law, which can compel companies to share data with the government on request. DeepSeek does offer an opt-out, but it’s narrower than it sounds: the toggle covers whether your input is used to improve the model, not a blanket promise that your data is never processed at all. It’s still handled for things like security, legal compliance, and running the service itself.
Self-hosting the same open-weight model removes this specific concern, because your data never reaches the vendor’s infrastructure in the first place. The weights themselves don’t transmit anything once you’re the one running them, on your own or a third-party non-Chinese server.
Several governments have treated the hosted-app version as enough of a risk to act on. Italy’s data protection authority ordered a nationwide processing limitation on DeepSeek on January 30, 2025. Taiwan banned government agencies, state-owned enterprises, and public schools from using it starting January 31, 2025. Australia’s Department of Home Affairs directed all government devices to remove and stop using it on February 4, 2025. All three restrictions remain in effect as of this writing in 2026, so this wasn’t a brief overreaction that quietly got walked back. Worth flagging directly for readers here: Australia is one of this site’s core audience countries, and its restriction applies specifically to government devices and systems, not to personal or business use generally.
Practical Guidance: Which One Actually Makes Sense for You
Strip away the geopolitics and this comes down to a fairly ordinary cost-vs-risk decision, the same kind you’d make with any vendor.
| A Chinese model might make sense when… | A Tier-1 (US/UK/Canada/Australia) model is the safer default when… |
|---|---|
| You’re cost-sensitive or running high-volume usage where per-token pricing adds up fast | You’re handling business, client, or enterprise data of any kind |
| The data involved is genuinely non-sensitive: public content, general research, personal experimentation | You’re subject to a compliance regime with data residency or data-handling requirements |
| You’re able to self-host the open-weight model, or run it through infrastructure you actually control | You’d rather not have to think about self-hosting or the data-jurisdiction question at all |
| You want largest-possible context window or lowest-possible cost for a well-scoped, non-sensitive task | The topic could plausibly touch on politically sensitive subjects and you need consistent, unfiltered answers |
Neither column is the “right” answer in the abstract. It depends entirely on what you’re actually doing with it. A developer self-hosting Qwen for a cost-sensitive internal tool with no personal data involved is in a genuinely different position than a business routing customer conversations through a hosted Chinese consumer app.
Frequently Asked Questions
Are Chinese AI models actually as good as US models now?
On broad benchmark performance, they’re very close – Stanford HAI’s 2026 AI Index puts the gap between the single best US and Chinese models at 2.7 percentage points as of March 2026, down from 17.5-31.6 points in 2023. “As good” depends on the specific task and model pair being compared, but the old assumption of a large, settled US lead no longer holds up against the data.
Is it safe to use DeepSeek or other Chinese AI apps?
It depends on what you mean by “use.” Using the hosted consumer app or API means your data is processed on servers in China, subject to Chinese data law – which is exactly why Australia, Taiwan, and Italy restrict it on government devices. Self-hosting the same open-weight model yourself avoids that specific issue, since your data never reaches the vendor’s servers at all.
Does self-hosting an open-weight Chinese model solve the privacy concern?
It solves the data-jurisdiction concern specifically, since nothing is sent to the vendor’s own servers. It doesn’t automatically solve every other consideration – you’re still responsible for securing whatever infrastructure you run it on, the same as with any self-hosted software.
Are all Chinese AI models equally restricted on political topics?
No. A 2026 Stanford/Princeton study (PNAS Nexus) found refusal rates on political questions ranging from 10% (ChatGLM) to over 60% (BaiChuan) among the Chinese models it tested, against 0-2.8% for the Western models in the same study – a real spread, not a uniform block. The study also didn’t cover every current model, such as Qwen or Kimi K3.
Why are Chinese AI models so much cheaper?
Partly efficient model architecture (DeepSeek’s Mixture-of-Experts approach is often cited specifically), and partly a deliberate strategy – Chinese labs have generally priced aggressively and released open weights to build adoption and market share, in contrast to the leading US labs, which mostly keep their top-tier models closed and priced at a premium.
Which one should I use for my business?
For business, client, or compliance-sensitive data, a Tier-1 model from a US, UK, Canadian, or Australian provider is the safer default, mainly because of data jurisdiction and the compliance paperwork it avoids. For cost-sensitive, non-sensitive, self-hosted use cases, a Chinese open-weight model is a legitimate option worth considering on its own merits.
Looking for an AI tool we’re comfortable recommending outright? Check our Recommended Tools page, or browse more coverage in our AI category.
