Instructions to use myroslavtryhubets/clef-flash-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use myroslavtryhubets/clef-flash-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="myroslavtryhubets/clef-flash-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("myroslavtryhubets/clef-flash-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("myroslavtryhubets/clef-flash-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use myroslavtryhubets/clef-flash-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "myroslavtryhubets/clef-flash-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "myroslavtryhubets/clef-flash-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/myroslavtryhubets/clef-flash-NVFP4
- SGLang
How to use myroslavtryhubets/clef-flash-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "myroslavtryhubets/clef-flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "myroslavtryhubets/clef-flash-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "myroslavtryhubets/clef-flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "myroslavtryhubets/clef-flash-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use myroslavtryhubets/clef-flash-NVFP4 with Docker Model Runner:
docker model run hf.co/myroslavtryhubets/clef-flash-NVFP4
Clef-Flash — NVFP4A16 (weights-only 4-bit)
An unofficial, community NVFP4A16 (weights-only 4-bit) quantization of
Cloudflare/clef-flash, the 9B sibling of Clef:
a multimodal decision model that turns a state and a schema of typed questions into a probability
for every allowed option in one forward pass.
NVFP4A16 — 4-bit weights, bf16 activations. More accurate than W4A4 but ≈7x slower; same VRAM. Runs on a small GPU — ≈4 GiB of weights in VRAM.
Not affiliated with or endorsed by Cloudflare. The model, architecture and joint_schema_model.py
are theirs; this repo only re-encodes the weights. Apache-2.0, following the base model.
Usage
The quantized weights are not loaded by from_pretrained / load_release_model (those build
the full-precision graph). Use the included clef_rt.py:
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("myroslavtryhubets/clef-flash-NVFP4A16")
sys.path.insert(0, path)
from clef_rt import Clef
model = Clef(path, "nvfp4-w4a16") # ≈4 GiB VRAM, Blackwell (sm_120)
out = model.probs([model.encode({
"state": "Our checkout is down and orders are blocked.",
"questions": {
"outage": {"type": "noul", "instructions": "Is a service down?"},
"team": {"type": "choice", "instructions": "Which team should handle it?",
"criteria": {"billing": "payments", "tech": "outages"}},
},
})])[0]
print(out)
Quality — Decision Index 0.2.1
11 benchmarks (20,337 requests) of the Decision Index 0.2.1 suite, rebuilt byte-identical with the official kit and scored with its own scorer. Clef-flash bf16 (CF) is Cloudflare's published result; This is this checkpoint on one RTX 5060 Ti 16 GB. Coverage-adjusted percentages.
| Benchmark | Metric | Clef-flash bf16 (CF) | This | Δ |
|---|---|---|---|---|
| BFCL | case exact accuracy | 98.8 | 98.6 | -0.2 |
| BANKING77 | macro-F1 | 90.9 | 92.0 | +1.1 |
| CLINC150+OOS | macro-F1 | 66.8 | 61.2 | -5.6 |
| ContractNLI | macro-F1 | 84.3 | 83.7 | -0.6 |
| ANLI | macro-F1 | 59.1 | 58.8 | -0.3 |
| ARC-Challenge | accuracy | 98.3 | 98.1 | -0.2 |
| WinoGrande | accuracy | 97.5 | 97.1 | -0.4 |
| MuSR | accuracy | 86.0 | 85.4 | -0.6 |
| FinEntity | macro-F1 | 97.1 | 97.2 | +0.1 |
| CRUXEval | accuracy | 86.1 | 85.3 | -0.8 |
| PhishNChips | accuracy | 75.0 | 76.5 | +1.5 |
Mean absolute gap to Cloudflare: 1.04 points. Median latency ≈1681 ms/request (batch 1).
Variants
| Repo | Scheme | CLINC150 | Mean gap | Median latency |
|---|---|---|---|---|
| clef-flash-NVFP4 | W4A4 | 49.9 | 2.36 | ≈240 ms |
| clef-flash-NVFP4A16 | W4 only | 61.2 | 1.04 | ≈1680 ms |
| clef-NVFP4 | W4A4, 27B | 97.3 | 1.08 | ≈700 ms |
Caveats
- CLINC150 still -5.6 vs Cloudflare (66.8 -> 61.2) even weights-only: the 9B model is sensitive to weight quantization on 150 near-tied classes. Much better than W4A4 (49.9) but not lossless.
- Slow for a "Flash" model (median ≈1.7s). On a 16 GB card Clef 27B NVFP4 is both faster (≈0.7s) and more accurate; pick this one only when 13 GB of VRAM for the 27B is not available. For speed use the W4A4 variant.
- Quality verified only through this runtime (
clef_rt.py), not a third-party serving stack. - vLLM: registers the
Qwen3_5backbone but has no Clef joint-head / decision path, so it cannot produce typed decisions out of the box.
Credits
Model, architecture and decision API by Cloudflare (blog). Base model Qwen/Qwen3.5-9B. Quantization by @myroslavtryhubets.
- Downloads last month
- 26