Instructions to use canada-quant/GLM-5.3-Flash-W4A16-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use canada-quant/GLM-5.3-Flash-W4A16-MTP with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="canada-quant/GLM-5.3-Flash-W4A16-MTP") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("canada-quant/GLM-5.3-Flash-W4A16-MTP") model = AutoModelForMultimodalLM.from_pretrained("canada-quant/GLM-5.3-Flash-W4A16-MTP", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use canada-quant/GLM-5.3-Flash-W4A16-MTP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "canada-quant/GLM-5.3-Flash-W4A16-MTP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "canada-quant/GLM-5.3-Flash-W4A16-MTP", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/canada-quant/GLM-5.3-Flash-W4A16-MTP
- SGLang
How to use canada-quant/GLM-5.3-Flash-W4A16-MTP with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "canada-quant/GLM-5.3-Flash-W4A16-MTP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "canada-quant/GLM-5.3-Flash-W4A16-MTP", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "canada-quant/GLM-5.3-Flash-W4A16-MTP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "canada-quant/GLM-5.3-Flash-W4A16-MTP", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use canada-quant/GLM-5.3-Flash-W4A16-MTP with Docker Model Runner:
docker model run hf.co/canada-quant/GLM-5.3-Flash-W4A16-MTP
GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP
Model description
INT4 weight-only quantization of zai-org/GLM-5.3-Flash with the BF16 MTP draft head kept for speculative decoding. Only the 36,288 routed-expert GEMMs are INT4 (GPTQ, symmetric, group-size 128); attention, router, shared experts, embeddings, the vision tower and the MTP head stay in BF16. Not a new model — all capability comes from the base model.
| Size | 177.7 GiB (BF16 ≈ 599 GiB, −70%) |
| Runs on | NVIDIA H100, H200, RTX PRO 6000 and DGX Spark (GB10) — validated GPU counts and context per configuration in Hardware & context limits |
| Context | full 1,048,576 tokens on 2× H200, and on 2× DGX Spark with the incoai reference drafter (with our DFlash2-G drafter: 262K default, 800K validated); 262K–512K on the 4/8-GPU x86 configurations (KV-memory-bound) |
| Quality | AIME 2025 0.8833 (n=120, H100) · AIME 2026 85.0% (n=120, 2× DGX Spark) · GPQA-Diamond 0.8586 · GSM8K 0.97 · vision: MMMU 0.747 · OCRBench 888 |
| Throughput | 8× H100: 181.2 tok/s single-stream (TP8 + DFlash2-G K=7) and 2,418 tok/s aggregate @c256 (TP8, no speculation) · 4× H200 TP=4: 195 → 1,954 tok/s from c1 to c256 · 2× DGX Spark: ≈31–33 tok/s single-stream at short context |
Full benchmark grids, protocols and research notes: BENCHMARKS.md.
Uses & recommended recipes
Quick start
# 1. Download (~178 GiB)
huggingface-cli download canada-quant/GLM-5.3-Flash-W4A16-MTP --local-dir /models/glm53-flash-w4a16-mtp
# 2. Serve on 4× H100 / H200 (other hardware: see Serving recipes)
docker run --gpus '"device=0,1,2,3"' --ipc=host --network=host --rm \
-v /models:/models vllm/vllm-openai:glm53-flash-x86_64-cu130 \
vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53-w4 \
--tensor-parallel-size 4 --enable-expert-parallel \
--max-model-len 262144 --max-num-seqs 512 \
--gpu-memory-utilization 0.92 --no-enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--reasoning-parser glm45 --tool-call-parser glm47 --enable-auto-tool-choice \
--trust-remote-code --port 8000
# 3. Call it (OpenAI-compatible)
curl http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "glm53-w4",
"messages": [{"role": "user", "content": "Prove there are infinitely many primes."}]
}'
Two things that bite. If your
config.jsonpredates 2026-09-08, re-download it — older copies fail in vLLM withKeyError: 'layers.0.mlp.gate_up_proj.weight'(weights are unchanged). And always pass--max-num-seqs ≤ 512— the vLLM default of 1024 exceeds this hybrid linear-attention model's 512 Mamba/KDA-state cache blocks.
Serving recipes
| Architecture | Image |
|---|---|
| SM90 (H100 / H200) | vllm/vllm-openai:glm53-flash-x86_64-cu130 (validated). Upstream vllm/vllm-openai:nightly-x86_64 ≥ 2026-09-08 also boots this checkpoint (vllm-project/vllm#53906); pass --attention-backend FLASH_ATTN_MLA_SPARSE there, its default backend faults on ≥131K prompts. |
| SM90, 8× H100 + DFlash2-G drafter | ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2 — vLLM nightly pin + baked drafter patches (EAGLE3 aux taps, drafter-aware KV partitioning); image of record for the 2026-09-28 8×H100 sweep |
| SM120 (RTX PRO 6000) | cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1 |
| SM121 (DGX Spark) | ghcr.io/canada-quant/vllm-glm53-flash-sm121:v2-w4a16-dflash2e, built from canada-quant/vllm-glm53-flash-sm121 with the two SM121 serving patches baked in |
All configurations use expert parallelism. fp8 KV is not available on Hopper for this NoPE model.
H100 / H200, TP=4 (pinned image) — the Quick start command is the benchmarked recipe. num_speculative_tokens: 2 is the sweet spot on Hopper (52–55% acceptance; N=5 collapses acceptance to ~30%). Keep prefix caching off — it measured −2…−5% on H200. On 8× H200, two independent TP=4 replicas behind a load balancer beat one TP=8 endpoint by +28–38% aggregate at c128–c512; TP=8 wins single-stream and holds one 7.8M-token pool. On 2× H200 (89.5 GiB weights per GPU) add --tensor-parallel-size 2 --max-num-seqs 128 --max-cudagraph-capture-size 128 for 262K, or --max-model-len 1048576 --max-num-seqs 16 --gpu-memory-utilization 0.95 --max-cudagraph-capture-size 64 --max-num-batched-tokens 4096 for 1M.
8× H100, TP=8 (v2 image, bench-validated 2026-09-28) — 8192/1024 prompts. The latency recipe is TP=8 + DFlash2-G K=7: 181.2 tok/s c1, +34% over TP8 without speculation (135.4). The aggregate recipe is TP=8 without speculation: 2,418 tok/s @c256; MTP N=2 (2,330 @c256) is the balanced middle and has the best grid c1 (215.0). K=4 beats K=7 on aggregate at every concurrency. Avoid TP=4 with the DFlash2 drafter past c32: its 8 full-attention KV tensors cut the KV pool ~6× vs TP8 (311,999 vs 1,811,949 tokens) and force a structural preemption/recompute TTFT cliff. Uncontended TP=4 with MTP N=2 is the best single-stream shape measured on this rig (c1 193.18, 1,292 @c256). With any drafter, cap --max-model-len ≤ 131071: the speculative path faults on 131K-token prefills (without speculation they run). Image ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, cold boot ≈ 12 min. Full grids: BENCHMARKS.md.
RTX PRO 6000, TP=4 — same command with the SM120 image and --max-num-seqs 64 --max-num-batched-tokens 8192 --kv-cache-dtype fp8 --enable-prefix-caching. fp8 KV is required at 262K on 96 GB cards. Keep MTP on at every concurrency here: it adds +70% at c1 and +48% at c32.
2× DGX Spark, TP=2 — prebuilt image and one-command launcher in canada-quant/vllm-glm53-flash-sm121; the launcher also ships in the drafter repo. Drafter: canada-quant/GLM-5.3-Flash-DFlash2-G, our self-trained DFlash2 drafter (Apache-2.0; 3.676 mean acceptance at K=7 on our 500-prompt holdout vs 3.632 for the incoai reference on the same hardware; drop-in successor of -F and -E). Start the worker rank first, wait 25 s, then the head rank.
# on both nodes, rank1 (worker) first, then rank0 (head) 25 s later
# 262K context (production config): 8 GiB fp8 KV -> 366,749-token pool (the launcher default since 2026-10-04;
# older launcher copies defaulted to 3 GiB, which is too small for DFlash2-E/-F/-G, so passing it explicitly is safe)
MAX_MODEL_LEN=262144 KV_CACHE_MEM=8053063680 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>
# long context with DFlash2-G (validated 2026-09-26/27): 800K, 16 GiB KV pin -> 888,729-token pool
MAX_MODEL_LEN=800000 KV_CACHE_MEM=17179869184 GMU=0.90 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>
- Drafter memory (corrected 2026-10-04). DFlash2-E/-F/-G have 8 full-attention layers that each cache the whole context, about 3× the KV per token of the incoai reference drafter (2,048-token sliding window). Measured pools on this pair: 366,749 tokens at 8 GiB (262K) and 888,729 at 16 GiB (800K) with G, vs 1,360,420 at 9 GiB (1M) with the incoai drafter. 1M on 2× Spark is validated with the incoai drafter only; with G, 9 GiB holds ≈450K tokens. Earlier versions of this card paired G with 1M at 9 GiB.
- Drafter or built-in MTP head? DFlash2-G suits short-context English and code. For long contexts, non-English text or memory headroom, use the checkpoint's own BF16 MTP head (
{"method":"mtp","num_speculative_tokens":3}). Our G recipe slows at depth: 31.4 / 30.7 / 15.3 / 10.0 tok/s at 0 / 4K / 65K / 100K (llama-benchy pp2048/tg128). A community user reported 31.6 / 29.8 / 25.4 / 31.0 tok/s at the same depths with this checkpoint on eugr's b12x build + MTP (1700 MHz GPU clock, different harness; not measured by us), and G 20–30% slower than MTP on non-English text. MTP on our SM121 image is not yet validated by us; a same-pair comparison started 2026-10-04. - Notes: 7 speculative tokens is the drafter's trained and measured setting (k=5 also boots, not benchmarked). Confirm the boot log shows the mask-embedding load (
mask_token_id 154856); keep single prompts ≤ ~310K tokens; stop withdocker stop -t 30, neverrm -f. Cold boot is 6–10 minutes.
2× DGX Spark on eugr's b12x build (community-validated, not one of our validated recipes) — this checkpoint needs a fused-attention ignore-list fix there (eugr/spark-vllm-docker#403): bind-mount a config.json copy whose quantization_config.ignore adds re:.*self_attn\..*, re:.*\.in_proj_qkvgfab.* and re:.*in_proj.* (safe: nothing under self_attn is quantized), and pass --block-size 256 --moe-backend marlin --attention-backend B12X --linear-backend b12x --kv-cache-dtype fp8 with --speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"B12X"}'. We have not tested our DFlash2 drafters on that build (its b12x attention lacks a full-window non-causal path, and the drafter cache stays BF16).
Quality
| Benchmark | Hardware | W4A16 (this) | Notes |
|---|---|---|---|
| AIME 2025 — n=120, max thinking, 131,072-token budget | H100 | 0.8833 | — |
| AIME 2025 — same protocol | RTX PRO 6000 | 0.8083 raw · 0.8833 with a budget-commit fix | see below |
| AIME 2026 — n=120, max thinking, 131,072-token budget | 2× DGX Spark | 85.0% (102/120) | EXL3 (MiaAI-Lab) 80.0% (96/120), matched protocol |
| GSM8K | H100, RTX PRO 6000 | 0.970–0.975 | — |
| GPQA-Diamond — n=198 @131K | H100 · RTX PRO 6000 | 0.8586 · 0.8586 | — |
| MMMU (validation) — n=900, lmms-eval 0.7.3 task spec | 8× H100 TP=8, greedy | 0.74667 (672/900) | MC 0.766 · open 0.434 |
| OCRBench — n=1000, lmms-eval 0.7.3 task spec | 8× H100 TP=8, greedy | 888 /1000 | weakest: handwritten math 56/100; strongest: doc-VQA 192/200 |
On RTX PRO 6000 the raw AIME 2025 score trails the H100 result: ≈63% of the gap is a budget wall (11.7–14.2% empty answers) and ≈37% is SM120 kernel numerics. A zero-cost commit hook closes it but is not part of the published recipes. The vision tower is BF16 passthrough and was not covered by the text-only calibration, so MMMU and OCRBench are this artifact's own vision baseline (greedy, seed 42, measured 2026-09-28). Details: BENCHMARKS.md.
Benchmarks summary
Output tok/s, thinking ON, measured by the authors. Competitor comparisons against NVFP4 checkpoints were removed on 2026-10-04: earlier tables mixed two different NVFP4 checkpoints under one label and compared against NVIDIA's without its MTP head. A clean same-rig comparison will be published when it's done.
| Hardware & configuration | Output tok/s |
|---|---|
| 8× H100, TP8 + DFlash2-G K=7 (latency recipe), 8192/1024 | c1 181.2, +34% over TP8 without speculation (135.4) |
| 8× H100, TP8 without speculation (aggregate recipe) | c256 2,418 |
| 8× H100, TP8, MTP N=2 | c1 215.0 · c256 2,330 |
| 8× H100, TP4 uncontended, MTP N=2 | c1 193.2, the best single-stream shape on this rig · c256 1,292 |
| 4× H200, TP=4, MTP N=2 | c1 195 · c8 698 · c32 1,258 · c64 1,529 · c128 1,789 · c256 1,954 (8× H200 TP=8: 217 → 2,681; two TP=4 replicas: 391 → 3,911 aggregate) |
| 4× RTX PRO 6000, TP=4, MTP N=2 | c1 109.9 · c8 318.5 · c32 534.4 |
| 2× DGX Spark, incoai DFlash2 drafter, 1M serve, 8K/256 (2026-09-03) | c1 33.0 · c2 35.6 · c4 59.6 · c6 67.1. Against EXL3 under a matched protocol: +10.4% c1 and +31–39% long-prefill; EXL3 leads at c2 (+67%) and c4 (+89%) |
| 2× DGX Spark, DFlash2-G K=7, 800K context, llama-benchy pp2048/tg128 (2026-09-27/28) | single-stream 31.4 / 30.7 / 15.3 / 10.0 at depth 0 / 4K / 65K / 100K |
Single-stream decode is insensitive to KV length up to ≥486K on RTX PRO 6000; at batch, long-KV decode plateaus at ~2–4 tok/s per stream and long prefills serialize at a ~6–8.5K tok/s aggregate ceiling. All grids and protocols: BENCHMARKS.md.
Hardware & context limits
Each row is the largest context serving-validated on that configuration.
| Configuration | GPUs | Validated context | Stack |
|---|---|---|---|
| DGX Spark GB10 (SM121), TP=2 | 2× 128 GB UMA | 1M — KV pool 1,360,420 tokens (1.30× a full 1M request), 2026-08-31 | incoai DFlash2 drafter, fp8 KV, 9 GiB KV |
| DGX Spark GB10 (SM121), TP=2 + DFlash2-G | 2× 128 GB UMA | 800K — KV pool 888,729 tokens at a 16 GiB KV pin (2026-09-26); default 262K = 366,749 at 8 GiB. The drafter's full-attention cache costs ~3× the KV per token, so 1M does not fit | DFlash2-G K=7, fp8 KV |
| RTX PRO 6000 (SM120), TP=4 | 4× 96 GB | 512K (486K prompts measured) | MTP N=2, fp8 KV |
| H100 (SM90), TP=4 | 4× 80 GB | 262K (256K prompts measured) | MTP N=2, bf16 KV |
| H100 (SM90), TP=8 + DFlash2-G | 8× 80 GB | 131,071 with any drafter — 128K-token prefills die on the speculative path (spec-off survives 131,072); with MTP-N2 the 262K row above applies | v2 image, K=7 latency / spec-off aggregate |
| H200 (SM90), TP=4 | 4× 141 GB | 262K — KV pool 6.19M tokens (≈23 concurrent 262K requests) | MTP N=2, bf16 KV |
| H200 (SM90), TP=8 | 8× 141 GB | 262K — KV pool 7.79M tokens | MTP N=2, bf16 KV |
| H200 (SM90), TP=2 | 2× 141 GB | 1M — KV pool 2.84M tokens (929K-token prompt measured) | MTP N=2, bf16 KV |
Known issues
config.json(2026-09-08): vLLM matchesquantization_config.ignoreagainst its own fused module names, so the ignore list now carries both the HF and vLLM spellings plusre:.*\.layers\.45\..*for the MTP head. Older 765-entry copies fail at load. Weights unchanged.- DFlash2 admission wedge (SM90 research stack only, MTP recipes unaffected): with the DFlash2 drafter at block size 2304, prompts above ~15.5K tokens are never admitted. A fix was validated to 256K prompts; block size 1536 avoids it. Filed as vllm-project/vllm#55800.
- Marlin no-split-K path on SM121: deterministic illegal memory access at M=256 when forcing
split_k=1; the stock heuristic used in serving is clean. Filed as vllm-project/vllm#56064.
Quantization details
| Field | Value |
|---|---|
| Architecture | Glm5NextForConditionalGeneration (glm5_next) — 45 decoder layers + MTP layer 45, 288 routed experts (top-8) + 1 shared, KDA + DSA attention, 24-block vision tower |
| Quantized | 36,288 tensors = 42 MoE layers × 288 experts × 3 GEMMs — W4A16, INT4, symmetric, group 128, GPTQ, compressed-tensors pack-quantized |
| Kept in BF16 | attention (incl. DSA indexer), dense prefix layers 0–2, shared experts, router, mHC tensors, embeddings, lm_head, norms, vision tower (348 keys), MTP layer 45 (889 keys) |
| Kept in FP32 | A_log, dt_bias, e_score_correction_bias, hc_* — verbatim from source |
| Calibration | 256 samples × 4096 tokens, in-distribution chat/code mix, sequential per-layer GPTQ |
| Built on | 8× NVIDIA B300, 2026-08-27 |
Build gates, all passing: exactly 36,288 packed tensors and nothing quantized outside routed experts; vision key set 348/348 identical to source; MTP layer present; zero dtype drift vs source; no collapsed expert scales. Loads with transformers ≥ 5.16; text generation and image captioning smoke tests pass.
Citation & license
@misc{canada_quant_glm53_flash_w4a16_mtp,
title = {GLM-5.3-Flash W4A16 (INT4) + BF16 MTP},
author = {canada-quant},
year = {2026},
url = {https://huggingface.co/canada-quant/GLM-5.3-Flash-W4A16-MTP}
}
MIT, inherited from the base model. Follow the base model's usage terms.
Built, benchmarked and documented with the Digby.ai coding harness, developed by CQL.ca.
- Downloads last month
- 11,328
Model tree for canada-quant/GLM-5.3-Flash-W4A16-MTP
Base model
zai-org/GLM-5.3-Flash