AI Author:EqualOcean News Updated 57 mins ago (GMT+8)

Zhipu AI(智谱)on August 14 released GLM-5.3, a large language model the Beijing-based company describes as the strongest open-weight model for coding. The company said it plans to publish the model’s weights two weeks after launch, following additional security testing and hardening.

zhipu

GLM-5.3 retains the same base model as GLM-5.2, with the latest gains coming entirely from post-training. Zhipu said it expanded task environments by tens of times, added more environment types and extended the duration of post-training, seeking to improve the model’s ability to complete long-horizon coding and agent tasks.

According to the company, GLM-5.3 delivers a 50% improvement in coding performance over GLM-5.2 in its internal evaluations. It also said the model ranks first among open-weight systems on public coding and agent benchmarks, including Terminal Bench 3.0 and Agents’ Last Exam (CLI). Such rankings can depend on the agent harness, settings and evaluation date, rather than the underlying model alone.

Zhipu attributed the improvement to its IndexShare, SAO and next-generation Slime reinforcement-learning frameworks, rather than to a new model architecture. The decision to leave the base model unchanged highlights a broader shift in the open-model race: as base architectures converge, labs are increasingly competing through post-training data, reinforcement learning and more demanding task environments.

The company also reported an unexpected improvement in cybersecurity-related tasks. It said GLM-5.3 performed on par with Mythos 5 in white-box code review and vulnerability discovery, although it did not provide independent results in its announcement. Zhipu said the capability emerged as training environments became longer and more complex, describing security work as a form of tightly constrained coding.

Zhipu began investing in the area in September 2025, it said, working with domestic security laboratories to develop longer task environments and specialised evaluation frameworks. The company’s decision to delay the release of model weights is intended to give it time to assess potential misuse risks and strengthen safeguards before the model is made broadly available.

The release arrives as coding has become a central battleground for open-weight models. Strong coding performance can drive adoption among developers and support agentic workflows, both of which help models gain practical use beyond benchmark tests. Chinese AI companies have increasingly paired commercial API offerings with public-weight releases, while many leading US labs keep their most capable systems behind paid, hosted services.

For developers outside China, an openly downloadable coding model could lower barriers to local deployment and customisation. That may be particularly relevant in markets where access to US-hosted AI services is limited by pricing, infrastructure or commercial restrictions.

Zhipu’s claim underscores the strategic importance of open-weight releases to China’s AI industry. As Chinese models improve in coding and security-adjacent tasks, the question will be not only how they compare on benchmarks, but whether developers and enterprises adopt them at scale.