Baidu and Zhipu's LLMs lead Chinese generative AI rankings

Baidu’s Ernie Bot 4.0 and start-up Zhipu AI’s GLM-4 top the ranking of Chinese large language models (LLMs), but their foreign rivals still lead in overall capabilities, according to a new test by Tsinghua University in Beijing. The SuperBench assessment report examined 14 representative LLMs – the technology underpinning generative artificial intelligence (AI) chatbots – and found that overseas models, such as OpenAI’s GPT-4 and Anthropic’s Claude-3, came out on top in multiple capabilities, including semantic comprehension, coding abilities and alignment with human commands.

Researchers found “obvious gaps” in the code-writing and operative abilities in the real-world environment between domestic and first-class foreign models. The report aims to “provide objective and scientific evaluation criteria” to examine a growing number of LLMs that have emerged recently, according to a WeChat post published by Tsinghua’s Basic Model Research Center, which conducted the assessment with the Zhongguancun Laboratory. Chinese tech giants and start-ups have been racing to improve their LLMs since OpenAI, a U.S. start-up backed by Microsoft, launched a series of innovative tools powered by generative AI, including ChatGPT and text-to-video service Sora. Around 200 LLMs have been introduced in China, where OpenAI’s services are officially unavailable.

The Tsinghua report echoes a recent comment by Alibaba Group Holding Co-founder and Chairman Joe Tsai, who said China is about two years behind U.S. companies in the global AI race, citing how OpenAI has leapfrogged the rest of the tech industry in AI innovation. Revisions to existing U.S. export controls, which took effect earlier this month, have made it harder for China to access advanced AI processors and semiconductor manufacturing equipment.

Despite the challenges faced by Chinese LLM developers, Tsinghua’s report showed that Ernie Bot 4.0, the latest version of the generative AI chatbot launched by Baidu, and GLM-4 from Zhipu AI, have gradually narrowed the gaps with the world’s best models in overall performance. One area where China’s LLMs performed better is Chinese text-language tasks. Start-up Moonshot AI’s Kimi chatbot, Alibaba’s Tongyi Qianwen 2.1, GLM-4 and Ernie Bot 4.0 ranked in the top four in that category, although GPT-4 still came first in Chinese text-language reasoning.

Moonshot AI and Zhipu AI, along with Baichuan and MiniMax, are locally known as the “four new AI tigers” of China for being some of the country’s most promising generative AI start-ups. Established in 2019, Zhipu AI has raised CNY2.5 billion since last year from investors, venture capitalists and Big Tech companies such as Alibaba, Tencent Holdings and Meituan. Moonshot AI, also based in Beijing, raised USD1 billion in a funding round in February, the South China Morning Post reports.