AI & ML interests
Canada Quant Labs — Canada's open-weight model lab. We train, quantize, and ship sovereign reference models for regulated industries (legal, medical, defence, finance) on a DGX B300 at Equinix Vancouver. Upstream contributors to vLLM and llm-compressor. Recipes: W4A16, NVFP4, MXFP4. Built in Victoria, BC. partnerships@cql.ca · cql.ca
Recent Activity
Canada Quant Labs
Canada's open-weight model lab.
We train, quantize, and deploy sovereign AI models on Canadian Blackwell silicon — for the regulated industries that can't run on someone else's API.
What we do
- Post-training on open base models (SFT, DPO, GRPO, RLAIF)
- Production quantization recipes (W4A16, NVFP4, MXFP4)
- Self-trained speculative-decoding drafters (DFlash2) for the models we serve
- Audited, air-gapped deployment with eval evidence and MRM docs
Where we work
- Legal · Medical · Defence · Finance
- Headquarters: Victoria, BC
- Compute: NVIDIA DGX B300 at Equinix Vancouver · 2× DGX Spark (GB10, SM121) serving cluster
Upstream
- Contributors to vLLM, llm-compressor, compressed-tensors
Open artifacts — 8 checkpoints on this org, every one shipped with its recipe, eval methodology, and engineering log
Partnerships · partnerships@cql.ca Press · press@cql.ca Web · cql.ca
GLM-5.3-Flash DFlash2-G — our own speculative drafter · Sep 2026
Our self-trained drafter, not a quant: a DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash, trained against our own W4A16-MTP quant on 736,675 self-generated samples — every completion regenerated by the target model itself; prompts from public instruction sets incl. real user chat, tool calling and hard math. No third-party drafter weights or traces anywhere in the training path. Apache-2.0.
→ canada-quant/GLM-5.3-Flash-DFlash2-G · serving stack: canada-quant/vllm-glm53-flash-sm121 (Apache-2.0 — one-command serving of GLM-5.3-Flash W4A16 + DFlash2 drafter on 2× DGX Spark: TP=2 + expert-parallel, 262K context, fp8 KV, K=7; its drafter/ folder holds the eval kit and every raw measurement)
Measured — ~1.84B parameters, block size 8 → K=7 speculative tokens, 9 hidden-state taps on the target. Same hardware, same protocol (500 held-out prompts, thinking ON, 8× B300): 3.676 mean acceptance at K=7 vs 3.632 for the incoai reference, 3.697 vs 3.667 single-stream — both outside the ±0.015 run-to-run noise; K=4 3.088 vs 3.136; on 4,096-token traces 3.478 vs 3.473 (parity — a ceiling of the target's long thinking for every drafter). Previous releases -F (3.626) and -E (3.561) stay up with their records. -E was adopted 15 Sep 2026 on the 2× DGX Spark stack (writeup: cql.ca/news/glm-5-3-flash-dflash2-e.html); -G is its drop-in successor.
GLM-5.3-Flash W4A16-MTP · Aug 2026
INT4 weight-only quantization of GLM-5.3-Flash with the BF16 MTP draft head preserved for speculative decoding — 177.7 GiB (−70%), serving the full 1M-token context on 4× H100/H200, 4× RTX PRO 6000, or a 2× DGX Spark desktop pair. MIT.
→ canada-quant/GLM-5.3-Flash-W4A16-MTP
Measured — only the 36,288 routed-expert GEMMs are INT4 (group-128, GPTQ); attention, router, shared experts, embeddings, the vision tower, and the MTP head stay BF16. On 2× DGX Spark the strongest community NVFP4 quant OOM'd 9/9 boots at any context while W4A16 serves 1,048,576 tokens; at 262K it is +53% seq1 / +101% seq6. Matched-protocol quality: AIME 2026 85.0% (102/120), +5.0 pt over the EXL3 arm; GSM8K 0.97. Published 4× H200 recipe: 195 → 1,954 output tok/s from c1 to c256 (GSM8K-200 on the recipe: 0.99); 2× H200 admits a 929K-token prompt on two GPUs.
Writeup: cql.ca/news/glm-5-3-w4a16-mtp.html.
Hy3 W4A16-MTP · Jul 2026
A 4-bit weight quantization of Tencent Hy3 (295B-parameter MoE, 21B active) that keeps the multi-token-prediction (MTP) draft layer in BF16. It matches the FP8 release on quality at 57% of the footprint (≈598 GB BF16 → 172 GB), runs on 8×H100-80GB with headroom (4×H100 at TP=4), and its preserved MTP layer is worth +40% single-stream throughput over the same-scheme quant that shipped without it. Apache-2.0.
→ canada-quant/hy3-w4a16-mtp · calibration data: canada-quant/hy3-w4a16-mtp-calibration
Quality is statistically indistinguishable from tencent/Hy3-FP8 on the same harness (8×H100): AIME24/25 70.0/83.3, MATH-500 94.0, GPQA-diamond 88.4. Speed: 144 output tok/s at concurrency 1 with MTP k=2; k=1 tops the whole 12-config matrix at c=8 (822 tok/s) and c=32 (2,146 tok/s, +21% over FP8). Serving: stock vLLM ≥ 0.25, no patches; native 256K context validated.
Writeup: cql.ca/news/hy3-w4a16-mtp.html.
GLM-5.2 W4A16-MTP · Jun 2026
GLM-5.2 (744B MoE, ~40B active, MLA + DeepSeek Sparse Attention, native 1M context, MIT). We shipped one artifact: W4A16 INT4 on the routed experts with the BF16 MTP draft head preserved for speculative decoding (attention, dense layers 0–2, shared experts, router, embeddings, lm_head, and MTP layer 78 stay BF16).
| Repo | Routed experts | MTP | On-disk | Hardware | Notes |
|---|---|---|---|---|---|
| GLM-5.2-W4A16-MTP | W4A16 INT4 g=128 (GPTQ) | yes (BF16, layer 78) | ~405 GB | 4×H200 (≤128K) · 8×H200 (1M) | matches FP8 quality; fastest 4-bit GLM-5.2 quant at interactive concurrency |
Validated against the official FP8 release (same harness, 8×H200): matches on GSM8K (0.960), IFEval (0.909/0.911), MATH-500 (0.954), RULER@32K/64K, and SWE-bench Verified (82.0% vs 82.2%) — serves at 1M context on 8×H200, and on 4×H200 up to ~128K. Throughput vs the most-downloaded community 4-bit quants: leads the interactive regime — +69–79% at concurrency 1 vs AWQ-INT4 / NVFP4 — from the MTP draft head (neither competitor ships one); the no-MTP quants edge ahead only at full saturation.
Writeup: cql.ca/news/glm-5-2-w4a16-mtp.html.
DeepSeek-V4 quantization family · May 2026
Four artifacts in the same lineage. One base model in two sizes (V4-Flash, V4-Pro); two routed-expert formats (W4A16, NVFP4); MTP draft head retained on three of four. Attention is FP8 block 128×128 across all four. Upstream reference recipes: RedHatAI/DeepSeek-V4-Flash-NVFP4-FP8 (Flash NVFP4) and nvidia/DeepSeek-V3.2-NVFP4 (Pro NVFP4, MTP-exclusion topology).
| Model | Base | Routed experts | MTP | On-disk | Min hardware (TP=2) | When to pick |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash-W4A16-FP8 | V4-Flash | W4A16 INT4 g=128 | no | ≈143 GB | H200 / DGX Spark / RTX PRO 6000 | maximum compatibility, no MTP needed |
| DeepSeek-V4-Flash-W4A16-FP8-MTP | V4-Flash | W4A16 INT4 g=128 | yes (BF16) | 159 GB | H200 / RTX PRO 6000 | best $/token interactive on V4-Flash |
| DeepSeek-V4-Flash-NVFP4-FP8-MTP | V4-Flash | NVFP4 g=16 | yes (BF16) | 172 GB | RTX PRO 6000 / B300 | best Blackwell-native interactive on V4-Flash |
| DeepSeek-V4-Pro-NVFP4-FP8-MTP | V4-Pro | NVFP4 g=16 | yes (byte-identical) | 913 GiB | 8× B300 (TP=8 + EP) | only choice for V4-Pro deployment; +25–37% throughput vs upstream MXFP4 |
Hardware shorthand
- H200 — 8× NVIDIA H200 SXM5 (Hopper SM 9.0a, 141 GB HBM3e/GPU)
- DGX Spark — 2× NVIDIA DGX Spark (GB10, Blackwell SM 12.1a)
- RTX PRO 6000 — NVIDIA RTX PRO 6000 Blackwell Server Edition (SM 12.0, sm_120, 96 GB HBM)
- B300 — NVIDIA B300 SXM6 AC (Blackwell SM 10.3, sm_103a, 288 GB HBM3e/GPU)
Reproduction repos
Every artifact has a public reproduction repo with calibration scripts, vLLM patches, bench harnesses, and findings docs:
canada-quant/dsv4-flash-w4a16-fp8canada-quant/dsv4-flash-w4a16-fp8-mtpcanada-quant/dsv4-flash-nvfp4-fp8-mtpcanada-quant/dsv4-pro-nvfp4-fp8-mtp(recipe repo not yet public)