Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts Paper • 2610.02241 • Published 12 days ago • 3
MoESQ Collection Paired-4:8 + NVFP4 W4A4 expert compressed MoE models for NVIDIA Blackwell Sparse Tensor Cores. • 6 items • Updated 5 days ago • 3
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published Sep 3 • 86
All QUASAR Models Collection All QUASAR checkpoints in one place: 4-bit QAT for Qwen, Gemma and Muse — NVFP4 W4A16/W4A4 for vLLM, Q4_0 GGUF for llama.cpp / Ollama. • 12 items • Updated 24 days ago • 2
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction Paper • 2608.13966 • Published Aug 14 • 6
Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 19 days ago • 93
Swift1.5 Flash Next Collection Swift Flash Next: the next Swift generation. Weights and quantized builds are added here as they are released. • 6 items • Updated 2 days ago • 10
ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 Image-Text-to-Text • 120B • Updated 16 days ago • 2.09k • 14
Swift 1.5 27B Collection Swift 1.5 on Qwen3.8-27B: stronger than Swift 1.0 on agentic and coding tasks, with fewer thinking tokens. BF16 weights and every quant. • 11 items • Updated 2 days ago • 23