🇨🇳 China's chip workaround, 🎵 Anthropic suit, OpenAI's India raid
| by Kai · The Strategist | 12-min read |
This week, the AI race tilted decisively toward the economics of scale: Chinese labs proved domestic chips can carry global inference traffic, while Tencent open-sourced a 770B-parameter model at cut-rate prices. Meanwhile, legal and regulatory pressures on US vendors intensified, from Anthropic's copyright fight to OpenAI's talent grab in India. The pattern is clear: competitive advantage is moving from raw hardware access to the efficiency of software and the control of distribution.
Zhipu confirms viral Ox Alpha is GLM-5.3-Flash, runs on 100,000 domestic chips
Chinese AI lab Zhipu confirmed that the anonymous model Ox Alpha, which broke usage records on OpenRouter, is its open-weight GLM-5.3-Flash, served entirely on a cluster of more than 100,000 domestic chips. The model processed 62 trillion tokens before its formal release and lifted Zhipu's Hong Kong shares 12 per cent. Zhipu says per-token cost and hardware efficiency are comparable to mainstream Nvidia GPUs.
- GLM-5.3-Flash ran all online inference on more than 100,000 domestically produced chips.
- The model processed 62 trillion tokens before release, including 11 trillion on OpenRouter in three days.
- It scored 57 on the AI Intelligence Index, matching Claude Opus 4.8 and beating DeepSeek's V4 Pro.
- Zhipu's shares closed 12 per cent higher at HK$1,160 in Hong Kong.
The surface story is that Zhipu AI launched a model that went viral under a codename. The actual story is that a Chinese lab proved it can serve hundreds of billions of inference tokens per week on 100,000 domestic accelerators without giving up benchmark standing or per-token cost. GLM-5.3-Flash processed 62 trillion tokens before its formal release, hit 11 trillion on OpenRouter in its first three days, and now scores 57 on the AAII index, above DeepSeek's V4 Pro and on par with Claude Opus 4.8. This is not a lab curiosity. It is a production-grade system running at global scale on hardware that the US export regime was designed to deny.
The mechanics of this shift appear in who gains advantage. Zhipu's shares closed up more than 12 percent, but the durable gain belongs to domestic chip suppliers. Analysts suspect Huawei's Ascend series, with Moore Threads and Hygon also named, and Zhipu's claim that hardware efficiency and per-token cost are comparable to mainstream Nvidia GPUs is an invitation for other Chinese labs to buy domestic silicon. Every cluster installed reduces Nvidia's pricing power in China's largest growth market. At the same time, GLM-5.3-Flash is priced at about one-tenth of GLM-5.3 and one-fortieth of Claude Opus 4.8, which will pressure premium API pricing globally. If a Chinese lab can deliver frontier-adjacent quality at those margins on domestic chips, the argument that export controls protect Western AI margins weakens.
The second-order implication most coverage will miss is that inference is becoming a commodity, and value is migrating to distribution and application layers. Zhipu's anonymous testing strategy on OpenRouter and OpenCode was not a marketing stunt. It turned an unknown model into the platform's biggest launch by routing live developer traffic through a production inference stack. More than 10.3 trillion tokens, nearly a third of OpenRouter's weekly volume, flowed to one model. That is a scale test, but it also demonstrates that a model's fate depends on platform access and routing decisions, not only on weights. This connects directly to Tencent open-sourcing a 770-billion-parameter model with a 1-million-token context in the same week, and to MiniMax nearly tripling its Alibaba Cloud budget as compute demand surges. Chinese labs are competing on infrastructure cost and model efficiency, not on exclusive access to Nvidia hardware. Zhipu's technical documentation shows that a separated Encode-Prefill-Decode architecture and quantified caching improved end-to-end performance threefold on the same hardware. Software optimization is now the strategic weapon, and the CUDA moat is being tested from multiple directions. The next question is not whether China can run frontier inference on domestic chips, but how quickly global prices for API access adjust to a market where that is the new baseline.
| China |
MiniMax nearly triples Alibaba Cloud budget as compute demand surges
Shanghai-based AI firm MiniMax raised its 2025 Alibaba Cloud spending cap to US$300 million from US$115 million, nearly triple the original limit, after burning through two-thirds of its cloud budget by June. The increase reflects surging compute needs for model training and live inference.
- MiniMax raised its 2025 Alibaba Cloud cap to US$300 million, up from US$115 million.
- The firm burned through two-thirds of its cloud budget by the end of June.
- Its 2026 API service budget with Alibaba rose from US$650,000 to US$7.5 million.
- First-half revenue surged 283% to US$116.6 million, with enterprise sales up 700%.
Noah Medical plans Hong Kong IPO to fund China hospital push
Noah Medical, a US surgical robotics company backed by SoftBank, plans to file for a Hong Kong listing as early as next year to raise over US$100 million and expand into mainland China. The company currently generates 90 per cent of its revenue in the United States.
- Seeks to raise more than US$100 million, with filing planned as early as next year.
- Founder Zhang Jian says Hong Kong is the natural choice for expanding business in mainland China.
- Hangzhou's Sir Run Run Shaw Hospital is a target customer after China's drug regulator approved Noah's products in November last year.
- Hong Kong's Prince of Wales Hospital has purchased Noah Medical equipment.
| India |
Meta India and Southeast Asia VP exits for OpenAI as regulatory pressure builds
Sandhya Devanathan, Meta's vice president for India and Southeast Asia, is leaving the company after more than a decade to join OpenAI, where she will oversee consumer growth, enterprise adoption, partnerships, regulatory engagement, and operations across Southeast Asia and Australia. Her departure comes as Meta faces growing scrutiny from Indian authorities over content moderation and child safety issues.
- Devanathan will be based in Singapore and report to OpenAI's Asia-Pacific managing director, Kiran Mani.
- Prabhjeet Singh joined OpenAI as India head days earlier, after leading Uber's India and South Asia business.
- Benjamin Joe, the Asia-Pacific VP, will now directly manage Arun Srinivas, Meta's India managing director.
- Meta removed 160,000 accounts in India over six months based on child-exploitative activity signals.
| Policy & Regulation |
Music publishers sue Anthropic over 'blatant theft' as IPO looms
Thirty-five music publishers, led by Sony Music Publishing and Warner Chappell, have sued Anthropic in California, accusing the AI company of scraping copyrighted music lyrics to train its Claude models. The lawsuit, one of several music copyright cases against Anthropic, comes as the company reportedly plans an IPO that investors value at up to US$2 trillion. Anthropic said it disagrees with the claims and will defend itself.
- The suit names Anthropic co-founders Dario Amodei and Benjamin Mann as defendants.
- Concord, Universal Music Publishing and ABKCO are pursuing separate lawsuits over song lyrics.
- BMG filed its own claim covering artists including Bruno Mars, Ariana Grande and the Rolling Stones.
- Investors have discussed an Anthropic valuation of up to US$2 trillion, according to the Financial Times.
| Quick hits |
Watch whether Chinese labs' success with domestic chips forces US model vendors to compete on software efficiency, and whether export controls lose their bite as a strategic lever.
The Asia AI Signal