Clef-Flash — NVFP4A16 (weights-only 4-bit)

An unofficial, community NVFP4A16 (weights-only 4-bit) quantization of Cloudflare/clef-flash, the 9B sibling of Clef: a multimodal decision model that turns a state and a schema of typed questions into a probability for every allowed option in one forward pass.

NVFP4A16 — 4-bit weights, bf16 activations. More accurate than W4A4 but ≈7x slower; same VRAM. Runs on a small GPU — ≈4 GiB of weights in VRAM.

Not affiliated with or endorsed by Cloudflare. The model, architecture and joint_schema_model.py are theirs; this repo only re-encodes the weights. Apache-2.0, following the base model.

Usage

The quantized weights are not loaded by from_pretrained / load_release_model (those build the full-precision graph). Use the included clef_rt.py:

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("myroslavtryhubets/clef-flash-NVFP4A16")
sys.path.insert(0, path)
from clef_rt import Clef

model = Clef(path, "nvfp4-w4a16")                   # ≈4 GiB VRAM, Blackwell (sm_120)

out = model.probs([model.encode({
    "state": "Our checkout is down and orders are blocked.",
    "questions": {
        "outage": {"type": "noul", "instructions": "Is a service down?"},
        "team": {"type": "choice", "instructions": "Which team should handle it?",
                 "criteria": {"billing": "payments", "tech": "outages"}},
    },
})])[0]
print(out)

Quality — Decision Index 0.2.1

11 benchmarks (20,337 requests) of the Decision Index 0.2.1 suite, rebuilt byte-identical with the official kit and scored with its own scorer. Clef-flash bf16 (CF) is Cloudflare's published result; This is this checkpoint on one RTX 5060 Ti 16 GB. Coverage-adjusted percentages.

Benchmark Metric Clef-flash bf16 (CF) This Δ
BFCL case exact accuracy 98.8 98.6 -0.2
BANKING77 macro-F1 90.9 92.0 +1.1
CLINC150+OOS macro-F1 66.8 61.2 -5.6
ContractNLI macro-F1 84.3 83.7 -0.6
ANLI macro-F1 59.1 58.8 -0.3
ARC-Challenge accuracy 98.3 98.1 -0.2
WinoGrande accuracy 97.5 97.1 -0.4
MuSR accuracy 86.0 85.4 -0.6
FinEntity macro-F1 97.1 97.2 +0.1
CRUXEval accuracy 86.1 85.3 -0.8
PhishNChips accuracy 75.0 76.5 +1.5

Mean absolute gap to Cloudflare: 1.04 points. Median latency ≈1681 ms/request (batch 1).

Variants

Repo Scheme CLINC150 Mean gap Median latency
clef-flash-NVFP4 W4A4 49.9 2.36 ≈240 ms
clef-flash-NVFP4A16 W4 only 61.2 1.04 ≈1680 ms
clef-NVFP4 W4A4, 27B 97.3 1.08 ≈700 ms

Caveats

  • CLINC150 still -5.6 vs Cloudflare (66.8 -> 61.2) even weights-only: the 9B model is sensitive to weight quantization on 150 near-tied classes. Much better than W4A4 (49.9) but not lossless.
  • Slow for a "Flash" model (median ≈1.7s). On a 16 GB card Clef 27B NVFP4 is both faster (≈0.7s) and more accurate; pick this one only when 13 GB of VRAM for the 27B is not available. For speed use the W4A4 variant.
  • Quality verified only through this runtime (clef_rt.py), not a third-party serving stack.
  • vLLM: registers the Qwen3_5 backbone but has no Clef joint-head / decision path, so it cannot produce typed decisions out of the box.

Credits

Model, architecture and decision API by Cloudflare (blog). Base model Qwen/Qwen3.5-9B. Quantization by @myroslavtryhubets.

Downloads last month
26
Safetensors
Model size
9B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for myroslavtryhubets/clef-flash-NVFP4

Finetuned
Qwen/Qwen3.5-9B
Quantized
(35)
this model