dante

DANTE v1

DANTE (Decisions As Next-Token Estimates) is a class-leading 4B decision model. You give it evidence, a question and the options; it returns a probability for every option from one forward pass. It generates no text: the answer is the model's next-token distribution over the option letters, so a decision costs one prefill and the probabilities are usable as confidence.

Try it in the playground: huggingface.co/spaces/elliottshort/dante-v1

The DANTE v1 playground answering a share-position question: long 1 to 199 shares at 96%

It handles three kinds of question:

Kind Options Example
Pick one 2 to 26 named options Which team should own this ticket?
Yes / no yes, no Did the agent stay within the user's request?
Score ordered levels, lowest first How urgent is this ticket, 0 to 4?

llama.cpp builds: dante-v1-GGUF, including a Q4_K_M for 16 GB Macs.

Results

Accuracy at one option order. DANTE v1 was scored as the Q8_0 GGUF through llama.cpp and the untuned base in BF16 with PyTorch. The development sets were used to pick checkpoints, so read them as in-distribution numbers.

Set Items Qwen3.5-4B, untuned DANTE v1
Development (mixed decision tasks) 1,179 64.3 84.5
Families (state tracking, tool-call guardrails, lead scoring) 400 47.2 88.8
Puzzles (temporal, financial and rule-following state) 475 45.1 84.4
Written decision scenarios 697 93.0 97.7

Usage

Tested with transformers 5.17.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from huggingface_hub import hf_hub_download
import importlib.util

path = hf_hub_download("elliottshort/dante-v1", "dante.py")
spec = importlib.util.spec_from_file_location("dante", path)
dante = importlib.util.module_from_spec(spec)
spec.loader.exec_module(dante)

tokenizer = AutoTokenizer.from_pretrained("elliottshort/dante-v1")
model = AutoModelForCausalLM.from_pretrained("elliottshort/dante-v1", dtype=torch.bfloat16)
model.to("cuda" if torch.cuda.is_available() else "cpu")

ticket = {"ticket": "The invoice total is double the price we were quoted.", "customer_tier": "enterprise"}
print(dante.decide(model, tokenizer, ticket, "Which team should own this ticket?", ["billing", "shipping", "technical support"]))
print(dante.decide(model, tokenizer, ticket, "Is the customer reporting a billing problem?", kind="noul"))
print(dante.decide(model, tokenizer, ticket, "How urgent is this ticket?",
                   {"0": "can wait a week", "1": "within two days", "2": "today"}, kind="score"))

dante.py builds the prompt the model was trained on and reads the letter probabilities:

<|im_start|>user
Answer with one option letter.

{evidence}

Q: {question}
A. {option}
B. {option}
Answer:<|im_end|>
<|im_start|>assistant
<think>

</think>
  • Evidence and question may be strings or JSON; JSON is written compactly.
  • Write an option as its name, or pass a dict to give each name a description (name: description).
  • Yes / no questions list yes then no. Score levels go lowest first.
  • The temperatures in dante.py (pick one 1.05, yes/no 0.35, score 0.5) calibrate one option order. views averages more option orders to cancel position bias, at proportional cost.

Training

LoRA (rank 32) on Qwen/Qwen3.5-4B at revision 851bf6e8, merged into the weights you download here.

  1. 64k decision items: public decision sets, items labelled with rationales by a 27B teacher, written decision scenarios, and programmatically generated items with exact labels (state tracking, temporal and financial arithmetic, policy checks, agent traces). The loss is cross-entropy on the option letter, with a ranked probability loss on score questions.
  2. A second pass of two epochs on 6,000 training items the first pass got wrong or answered with less than 0.6 confidence, mixed with 6,000 it already solved; the adapter is merged at scale 0.8.

Limitations

  • English only, and tuned for short decisions over supplied evidence, not open-ended generation or chat.
  • Fine-grained subjective ratings, such as helpfulness or similarity scores, are a weak spot.
  • Long agent traces and invoice matching remain the weakest in-distribution families.
  • The probabilities are calibrated on our sets; recalibrate the temperatures on your own data if calibration matters.

Credits

Qwen3.5-4B by the Qwen team (Apache 2.0). Training data includes eikos-decisions by Caio Vicentino (CC BY 4.0), plumb-decisions by crh225, ARC (CC BY-SA 4.0), GSM8K (MIT), CommonsenseQA, StrategyQA and LogiQA.

Citation

@misc{short2026dantev1,
  title  = {{DANTE} v1: Decisions As Next-Token Estimates},
  author = {Short, Elliott},
  year   = {2026},
  month  = oct,
  url    = {https://huggingface.co/elliottshort/dante-v1}
}
Downloads last month
293
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for elliottshort/dante-v1

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(885)
this model
Quantizations
1 model

Space using elliottshort/dante-v1 1

Collection including elliottshort/dante-v1