AI Policy Wiki
Dashboard

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models (Zhipu AI, 2025)

high confidence · updated 2026-07-26

Technical report for GLM-4.5, an open-source Mixture-of-Experts model with 355B total and 32B activated parameters, using a hybrid reasoning method supporting both thinking and direct response modes. Trained on 23T tokens; reports 70.1% TAU-Bench, 91.0% AIME24, 64.2% SWE-bench Verified, ranking 3rd overall and 2nd on agentic benchmarks among evaluated models. Released with a 106B GLM-4.5-Air variant.

The technical report for GLM-4.5, from Zhipu AI and Tsinghua University.

Architecture and training

An "open-source Mixture-of-Experts large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes."

Training: "multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning."

Reported results

The report organizes capability around ARC — agentic, reasoning, and coding:

BenchmarkScore
TAU-Bench70.1%
AIME2491.0%
SWE-bench Verified64.2%

The efficiency claim is the report's emphasis: "with much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks." The agentic placement above the aggregate is the notable asymmetry — this is a model whose relative strength is in tool use and multi-step tasks rather than raw capability.

Two models were released: GLM-4.5 at 355B and GLM-4.5-Air at 106B.

Relationships