View
KV TecHub
0%

GLM-5.2: The Most Powerful Open Model Can Now Run Locally

Z.ai has released GLM-5.2, a new open-weight model with 744 billion parameters that sets out to compete with the most capable closed-source systems. Thanks to a collaboration with Unsloth, it can now run locally — even on a Mac with 256 GB of unified memory.

GLM-5.2: The Most Powerful Open Model Can Now Run Locally

What Is GLM-5.2

GLM-5.2 is the most ambitious open-weight model released to date. With 744 billion total parameters and 40 billion active parameters (Mixture-of-Experts architecture), it comes with a 1 million token context window — enough to fit entire codebases or lengthy documents in a single call.

On benchmarks, GLM-5.2 performs at or near the level of models like Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro across reasoning, coding, and agentic tasks.


The Challenge: 1.51 TB on Disk

The biggest barrier to running this model locally has always been its size. The full model requires 1.51 TB of storage — well out of reach for most setups.

The solution came from Unsloth with their Dynamic GGUFs technology:

  • Dynamic 2-bit (UD-IQ2_M): 239 GB (-84% size) — fits on a 256 GB unified memory Mac

  • Dynamic 1-bit (UD-IQ1_S): 217 GB (-86% size) — for systems with 223 GB RAM

The idea behind "Dynamic" is smart: rather than applying uniform quantization everywhere, the most important layers are automatically upcast to 8-bit or 16-bit, so accuracy is preserved where it matters most.


Three Thinking Modes

A standout feature of GLM-5.2 is its support for three reasoning modes:

  • Non-thinking: Fast responses without extended reasoning

  • Thinking (High): Moderate reasoning depth for most tasks

  • Thinking (Max): Intensive reasoning for complex problems

For most use cases, the recommended settings are temperature = 1.0 and top_p = 0.95. For harder coding tasks (e.g. SWE-Bench), top_p should be raised to 1.0.


Hardware Requirements

Quantization

Memory (RAM + VRAM)

1-bit

223 GB

2-bit

245 GB

3-bit

290–360 GB

4-bit

372–475 GB

5-bit

570 GB

8-bit

810 GB

The 2-bit quant is the most accessible option: it fits on a Mac with 256 GB unified memory or a system with 1×24 GB GPU + 256 GB RAM using MoE offloading.


How to Run It

With Unsloth Studio (Easy)

Unsloth Studio is an open-source web UI that automatically detects multi-GPU setups and handles RAM offloading. It runs on macOS, Windows, and Linux.

Install:

bash

# macOS / Linux / WSL
curl -fsSL https://unsloth.ai/install.sh | sh

# Windows PowerShell
irm https://unsloth.ai/install.ps1 | iex

Launch:

bash

unsloth studio -H 0.0.0.0 -p 8888

Then open http://127.0.0.1:8888, search for GLM-5.2, pick your quant, and you're running.

With llama.cpp (Advanced)

If you prefer llama.cpp, download the model from Hugging Face and run it with:

bash

./llama.cpp/llama-cli \
    --model unsloth/GLM-5.2-GGUF/UD-IQ2_M/GLM-5.2-UD-IQ2_M-00001-of-00006.gguf \
    --temp 1.0 \
    --top-p 0.95 \
    --min-p 0.01

For extended context, KV cache quantization can stretch the usable context window by up to 3.5×:

bash

--cache-type-k q4_1 --cache-type-v q4_1

How Accurate Is the 2-bit Version?

The concern with quantized models is accuracy loss. Unsloth ran KL Divergence analysis to measure exactly that:

  • Dynamic 4-bit and 5-bit: Essentially lossless

  • Dynamic 2-bit: ~82% top-1% accuracy, 84% smaller file

  • Dynamic 1-bit: ~76% top-1% accuracy, 86% smaller file

For most tasks, the 2-bit version performs reliably. For complex out-of-distribution tasks, 4-bit is recommended.


Benchmarks: How Does GLM-5.2 Stack Up

A selection of comparisons against leading models:

Benchmark

GLM-5.2

Claude Opus 4.8

GPT-5.5

AIME 2026

99.2

95.7

98.3

GPQA-Diamond

91.2

93.6

93.6

SWE-bench Pro

62.1

69.2

58.6

Terminal Bench 2.1

82.7

78.9

83.4

MCP-Atlas

76.8

77.8

75.3

In mathematical reasoning (AIME 2026) and some coding benchmarks it matches or beats closed-source models. In others, such as SWE-bench, it trails Claude Opus 4.8 slightly.


What This Means in Practice

Until now, models in this performance class were only available through APIs — local execution was out of reach for most people. GLM-5.2 changes that equation.

The 245+ GB memory requirement still puts it beyond the typical laptop. But for developers, researchers, or companies that want frontier-class AI with full control over their data — without depending on external APIs — this is a meaningful step forward.

The combination of open weights, a 1M token context window, and smart quantization makes GLM-5.2 one of the most compelling open-weight models to have emerged so far.