As artificial intelligence steps out of the digital realm and into the real world, the race to build the embodied “brains” powering next-generation robots has become the newest battleground in tech competition between China and the United States. Two days after Nvidia launched its Cosmos 3 model – designed to help physical AI “think before it acts” – a Chinese start-up stole the spotlight. Hangzhou, Zhejiang province-based Spirit AI said its foundation model for embodied intelligence, Spirit v1.6, had become the first from China to top the RoboArena global leaderboard. It scored 1,924 on the benchmark, edging out Nvidia’s Cosmos3-Nano-Policy, which took second place with a score of 1,881. Coming in third was DreamZero with a score of 1,763 – another Nvidia project unveiled in February.
The RoboArena benchmark, which evaluates how effectively generalist robot policies translate to real-world actions, was co-developed by Nvidia alongside elite institutions including Stanford University and the University of California, Berkeley. The fierce competition underscores a broader shift: robotics is AI’s next frontier. Nvidia’s partnerships with China’s Unitree Robotics and Singaporean robotic hand pioneer Sharpa also highlight this trend.
Unlike large language models (LLMs), which are built to process and generate text and code, a physical AI model allows machines – such as humanoids, robotic arms or autonomous vehicles – to perceive, understand and interact with the physical world. Physical AI relies on two core capabilities. Policy capabilities are the model’s ability to take actions based on what it observes. This is the main metric measured by the RoboArena leaderboard. The second one is world capabilities, which is the model’s ability to simulate and predict what will happen next if a specific action is taken. While these functions are often developed separately, the industry is moving towards consolidation. Last September, Chinese researchers introduced a unified “Policy World Model” that integrates world modeling and trajectory planning into one architecture.
China’s dominance is not limited to policy models. The WorldArena benchmark, which evaluates embodied world models, is currently topped by WorldScape-0.2, developed by Chinese start-up Manifold AI. It beat out Nvidia’s Cosmos-Predict 2.5 in the policy evaluator track. Chinese firms are currently also leading other tracks within the WorldArena and broader evaluation ecosystems. The perception track is led by AgiBot with its GenieEnvisioner-Sim2.0-2B model, a video world simulator for robotic manipulation. The data engine track is led by Chinese start-up DexForce’s DSCFuncWorld, which focuses on optimizing the pipeline of training data. Meanwhile, Manifold AI’s WorldScape-0.2 has claimed the top spot on the WorldScore benchmark, designed to evaluate a model’s ability to generate worlds from text prompts, outperforming WonderJourney, a joint project by Stanford and Google, the South China Morning Post reports.