Instructions to use hbin0701/sft-model-1008 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hbin0701/sft-model-1008 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="hbin0701/sft-model-1008") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("hbin0701/sft-model-1008") model = AutoModelForCausalLM.from_pretrained("hbin0701/sft-model-1008", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use hbin0701/sft-model-1008 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "hbin0701/sft-model-1008" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hbin0701/sft-model-1008", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/hbin0701/sft-model-1008
- SGLang
How to use hbin0701/sft-model-1008 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "hbin0701/sft-model-1008" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hbin0701/sft-model-1008", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "hbin0701/sft-model-1008" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hbin0701/sft-model-1008", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use hbin0701/sft-model-1008 with Docker Model Runner:
docker model run hf.co/hbin0701/sft-model-1008
Rewind run4 SFT model ([rewound] marker, single-token <rewind>)
Qwen3-1.7B fine-tuned to write a <rewind> action while thinking. When the tag is written, an inference
controller deletes the last 4 reasoning sections, appends a [rewound] line, and generation continues from
there. This is the SFT stage of run4 (before any RL), trained with LOSS=v2 of the code below.
What changed vs. Sangsang/rewind-run1-sft-model
| run1 SFT | this model | |
|---|---|---|
<rewind> |
4 ordinary text tokens | one added token, id 151669 (a spare embedding row, initialised from <) |
| after a rewind | nothing marks it | [rewound] is appended where the erased text was |
| SFT loss | likelihood | likelihood + "don't fire here": -log(1 - p(<rewind>)) at paragraph starts of non-trigger rows (weight 1) |
| epochs | 2 | 1 |
Training data: the same 584 rows as Sangsang/rewind-run1-sft-data
(146 trigger, 146 post-rewind, 146 negative, 146 clean); post-rewind prompts get [rewound] appended.
System prompt (used in training; use it at inference)
You may write while reasoning. It erases the last part of your reasoning (about the last four steps) so you can continue from the earlier point where things were still on track. Use it only when you realize your reasoning has gone wrong and continuing will not lead to a correct answer; do not use it for routine double-checking. You can use it at most once per problem, and never after . After a rewind, [rewound] marks where the erased part was.
Thinking mode on: the assistant turn starts with an open <think>\n.
Running it
A plain generate() does not perform the rewind. A controller must:
- stop generation on token id 151669 (
<rewind>); - delete the last 4 sections of the reasoning (paragraph runs about one topic; see
section_startsin the code) and append[rewound]\n\n; - resume from that text; after one rewind, ban token 151669 (e.g.
logit_bias={151669: -100}).
The 16,384-token budget counts the tokens on the page, so the erased tokens are given back.
rewind/eval_rewind.py in the code is a complete controller (vLLM).
Known issue
This checkpoint fires far too often. Measured on 6 correct solutions, P(<rewind>) is 0.15-0.49 at almost every
paragraph break after the first, so in RL nearly every solution rewound ~200 tokens in. Cause: the "don't fire here"
term is averaged over each row's paragraph starts, which balances at about 146 / (146 + 438) = 0.25 per break.
Summing it instead is the planned fix. End-of-SFT log: P(tag | trigger point) ~0.19, highest P(tag) at a
"keep going" paragraph start ~0.47.
Training
| setting | value |
|---|---|
| base | Qwen/Qwen3-1.7B, full fine-tuning, fp32 master weights, bf16 compute |
| epochs / steps | 1 / 73 (global batch 8, 584 examples) |
| lr | 2e-6, 8 warm-up steps then cosine, AdamW (wd 0.01), grad clip 1.0 |
| loss | completion-only, mean per example; trigger target = the single <rewind> token |
| hardware | 2x RTX PRO 6000 Blackwell, 19 min |
training_log.jsonl has the per-step loss, P(tag | trigger) and the highest P(tag) at non-trigger paragraph starts.
Code
github.com/hbin0701/sclm, branch rl-resolving, commit c427255:
LOSS=v2 TAG_TOKEN=1 SFT_EPOCHS=1 bash run.sh reproduces this stage. The private repo needs an account with access.
- Downloads last month
- 225