Instructions to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 # Run inference directly in the terminal: llama cli -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 # Run inference directly in the terminal: llama cli -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 # Run inference directly in the terminal: ./llama-cli -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Use Docker
docker model run hf.co/kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
- LM Studio
- Jan
- vLLM
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
- Ollama
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with Ollama:
ollama run hf.co/kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
- Unsloth Studio
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 to start chatting
- Pi
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with Docker Model Runner:
docker model run hf.co/kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
- Lemonade
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Run and chat with the model
lemonade run user.DeepSeek-V4-Flash-0731-ROCmFP4-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Run Hermes
hermes
- Atomic Chat
DeepSeek-V4-Flash-0731 — ROCmFP4 (Strix Halo) GGUF
This is a ROCmFP4 quant of deepseek-ai/DeepSeek-V4-Flash-0731, built to fit a single AMD Strix Halo box (128 GB unified memory) with full GPU offload. As far as I can tell it's the first ROCmFP4 quant of this model. I made it with the ROCmFPX fork of llama.cpp for the gfx1151 (Radeon 8060S / Ryzen AI MAX+ 395) Vulkan/ROCm stack.
| Base model | deepseek-ai/DeepSeek-V4-Flash-0731 |
| Quant | Q3 — mixed ROCmFP4, experts ~3.14 bpw, 2.92 BPW overall |
| Size | ~101 GB (fits 128 GB unified memory with headroom) |
| Arch | deepseek4 (sparse MoE, 256 experts, indexer/DSA attention) |
| Target HW | AMD Strix Halo gfx1151 iGPU (Ryzen AI MAX+ 395), Vulkan RADV |
| Loader | ROCmFPX fork — stock llama.cpp cannot load ROCmFP4 tensors |
Why I made it
A standard 4-bit GGUF of this model comes out around 141 GB, which overflows a 128 GB Strix Halo's shared pool and spills to CPU. I wanted the largest-quality quant that still fully offloads on a single box and stays coherent, so I mixed the expert tensors down to land it at ~101 GB.
Recipe (quantized from the F16 with the fork's llama-quantize):
- base type
Q2_0_ROCMFPX ffn_down_exps→q3_0_rocmfpx(3.5 bpw)ffn_gate_exps,ffn_up_exps→q2_0_rocmfpx(2.5 bpw)- attention / embeddings → ROCmFPX; norms kept in fp32
The ROCmFP4 (_ROCMFPX) types hold quality better than equivalent-bit k-quants on this hardware while using the FP4 paths on gfx1151.
Running it
Build the ROCmFPX fork (llama-server / llama-cli) for gfx1151, then:
export HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1
export AMD_VULKAN_ICD=RADV VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/radeon_icd.json
./llama-server \
-m DeepSeek-V4-Flash-0731-Q3-ROCmFP4-00001-of-00004.gguf \
-dev Vulkan0 -ngl 999 -fa on -fit off --no-mmap \
-c 8192 -n 2048 -np 1 -b 1024 -ub 512 -t 16 --poll 50 --jinja \
--reasoning-format deepseek \
--chat-template-kwargs '{"enable_thinking":false}' \
--host 0.0.0.0 --port 8084
Notes from getting it stable on my box:
-fit off— the fork's auto-fit step crashed on this arch for me; pin-ngl 999and turn it off.--no-mmap— important for MoE speed. With mmap, experts page-fault per token and throughput roughly halves.-c 8192with-b 1024 -ub 512keeps the graph pool under its limit; larger context can overflow it.-n 2048caps runaway generations so one request can't hold the single slot forever.--chat-template-kwargs '{"enable_thinking":false}'gives fast, direct answers. Drop it (or passenable_thinking:trueper request) for the model's reasoning mode.- Expect roughly 5–8 tok/s — it's a 101 GB model on one iGPU. Use streaming for a usable feel.
A note on MTP
This checkpoint ships a multi-token-prediction (nextn) head, and I kept those tensors in this quant. I got a working MTP inference path running on this arch and tested it thoroughly, but on this hardware/loader combination MTP nets out slightly slower than plain decoding — the draft head's acceptance is low and the sparse-MoE verify step can't amortize its weight reads across draft tokens. I ran it against draft depth, the probability threshold, and draft-head precision; none of them turned it into a win here. So I ship it with MTP off. If you want the model's advertised MTP speedup, run it on a CUDA/vLLM stack instead of this one.
License
Derived from deepseek-ai/DeepSeek-V4-Flash-0731; the original model's license applies (see license_link). This upload is only a quantization — all capabilities and limitations are the base model's.
Other public builds of this model
Compiled from Hugging Face repository metadata — file sizes, shipped files, quant variant as named by each repo. No third-party build was run or benchmarked here, so this table makes no speed or quality claim about any of them. It is here so you can see the size and format options at a glance and pick what fits your hardware.
Base model: deepseek-ai/DeepSeek-V4-Flash-0731. Generated from Hub metadata; download counts move over time.
Acknowledgements
This build would not exist without the work below. Please star and follow these projects — the quantisation format used here is their engineering, not mine.
ROCmFPX — maintained by
charlie12345 / caf
The ROCmFP4 / ROCmFPX tensor formats (ggml types 100–106) exist only in this fork.
Every ROCmFP4 file in this repository was produced with its llama-quantize, and
runs on its runtime. The fork also credits collaborators ciru-ai, Tom Turney,
PlunderStruck and Aydan S., and acknowledges AMD for hardware support.
Licensed MIT, based on upstream llama.cpp.
llama.cpp — ggml-org and contributors The inference engine, GGUF format and conversion tooling everything here is built on.
AMD ROCm The compute platform these builds target — ROCm 7.2.4 on gfx1151 / Radeon 8060S.
Base model authors — see base_model in the metadata above; all model weights,
licences and capabilities are theirs. This repository contributes quantisation and
measurement only.
If you use these files, please credit ROCmFPX alongside this repository.
- Downloads last month
- 217
We're not able to determine the quantization variants.
Model tree for kingjones777/DeepSeek-V4-Flash-0731-ROCmFP4
Base model
deepseek-ai/DeepSeek-V4-Flash-0731