GLM-5.2-MLX-4bit

Runtime — updated 2026-08-28: load with --trust-remote-code

This repository now bundles glm_moe_dsa.py (declared via model_file in config.json), a fixed runtime for this architecture, and needs it:

mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-4bit --trust-remote-code --prompt "..." --max-tokens 300

mlx-lm's own glm_moe_dsa builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights on 21 (indexer_types: the other 57 "shared" layers reuse the previous full layer's top-k selection). mlx_lm.load loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048 tokens were unaffected (the indexer is bypassed below index_topk); beyond that, 57 of 78 layers attended to keys chosen by random projections. The bundled runtime implements the schedule as the reference does (plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against transformers 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it: github.com/PipeNetwork/glm53-mlx. The weights are unchanged.

MLX (Apple Silicon) conversion of zai-org/GLM-5.2 — a glm_moe_dsa MoE (256 experts, DeepSeek-V3.2-style sparse attention) — quantized to 4-bit.

Quantizations

Part of the GLM-5.2 MLX collection.

Variant Notes
8-bit 8-bit · ~800GB · needs ~1TB RAM · integrity-checked
6-bit 6-bit · ~625GB · needs ~768GB RAM · integrity-checked
5-bit 5-bit · ~530GB · needs ~640GB RAM · integrity-checked
4-bit (this repo) 4-bit · ~430GB · tight on 512GB · smoke-tested
mixed mixed · experts@3-bit / non-expert@6-bit · ~360GB · 512GB-fit · smoke-tested

Use with mlx-lm

pip install mlx-lm
python -m mlx_lm generate --model pipenetwork/GLM-5.2-MLX-4bit --prompt "Hello" -m 256

Validation

Smoke-tested locally (loads + generates coherent text).

License

MIT (inherited from base). Quantization config (excerpt): {"group_size": 64, "bits": 4, "mode": "affine"}.

Downloads last month
566
Safetensors
Model size
116B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pipenetwork/GLM-5.2-MLX-4bit

Base model

zai-org/GLM-5.2
Quantized
(145)
this model

Collection including pipenetwork/GLM-5.2-MLX-4bit