Skip to main content

US Lead in AI Benchmarks Narrows to 3%

6 OCTOBER 2026·2 MIN READ·4 SOURCES

Bloomberg Intelligence says the US lead over China in AI performance has narrowed to a record low, with top Chinese models now 3% behind US rivals on benchmark scores. Senior analyst Robert Lea attributes the improvement to DeepSeek’s V4.1 Flash release and says it could help Chinese competitors gain market share.

US Lead in AI Benchmarks Narrows to 3%

Key takeaways · 4

  • 01

    Top Chinese models trail US rivals by 3% on benchmark scores, compared with a 9% gap in May and 15% earlier in the year.

  • 02

    Bloomberg Intelligence links the improvement to DeepSeek’s September release of V4.1 Flash.

  • 03

    V4.1 Flash ranked sixth globally on LiveBench in September and scored 81.1, compared with 83.4 for an Anthropic model.

  • 04

    Treat leaderboard results as one input: rankings can shift, and strong scores may not translate into revenue.

A much narrower benchmark gap

Bloomberg Intelligence says the US AI performance lead over China has narrowed sharply to a record low.[1] Its senior analyst Robert Lea puts top Chinese models 3% behind US rivals on benchmark scores.[1] The gap was about 9% in May and 15% earlier in the year.[1] Lea attributes the improvement to DeepSeek’s September release of V4.1 Flash.[1] The figures describe benchmark performance; they do not establish that Chinese models have caught up across all uses or measures of AI capability.[1]

V4.1 Flash’s benchmark showing

Bloomberg Intelligence says V4.1 Flash ranked sixth globally on LiveBench in September, making it the highest-ranked Chinese model on the benchmark since DeepSeek’s R1 reasoning model drew attention in 2025.[2][3] The model scored 81.1, compared with 83.4 for an Anthropic model.[3] Lea describes its performance as comparable to leading systems from Anthropic and OpenAI.[2] Still, only three of LiveBench’s 15 top-ranked models were Chinese, and Lea cautions that benchmark rankings can change.[2]

Availability and technical details

DeepSeek says V4.1 Flash is live on its API with native multimodal support, and describes it as the smallest model in its new architecture family, with native visual understanding.[4] The company says it uses a 552-billion-parameter mixture-of-experts architecture, with 8 billion active parameters for input and 16 billion for output.[4] DeepSeek also says it uses one-quarter as much HBM and one-eighth as much SSD storage for its KV cache.[4] Its stated off-peak API rates are half of peak rates.[4]

Commercial gains remain uncertain

Lea says better performance could bring further market-share gains for Chinese competitors, but cautions that strong leaderboard scores may be difficult to monetize.[1][2] He says Chinese models face growing US regulatory scrutiny and potential bans following allegations of model distillation.[2] Lea says sustainable profits would require less competitive pressure, industry consolidation and more rational pricing, and says China’s AI industry could remain unprofitable until 2030.[2]

For teams evaluating AI models, the narrowing benchmark gap is a reason to compare systems on their own workloads rather than assume performance differences from geography or reputation. Commercial viability is a separate decision: benchmark scores, API costs and regulatory exposure all warrant consideration.

Why it matters
Story quiz

Test yourself on this story — 3 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 6 October 2026

    US Lead in AI Benchmarks Narrows to 3%

Sources

AI fluency, one session a day, built for your work.