Decider Ettin 68M
A 68M-parameter decision engine fine-tuned from Ettin Encoder 68M on the BEV Decision 150K dataset. Exported to ONNX (opset 18) and quantized to INT8 for browser inference via ONNX Runtime Web (WASM).
Everything runs in the browser. The model is downloaded once, cached in IndexedDB, and never requires a server for inference.
Decision Heads
The model has three output heads activated by the input sequence structure:
| Head | Output | Description |
|---|---|---|
| Choice | choice_logits |
Pick one of up to 10 options |
| Noul | noul_logits |
Binary yes / no verdict |
| Score | score_logits |
Rate on a scale of up to 7 levels |
Training Metrics
| Metric | Epoch 1 | Epoch 2 |
|---|---|---|
| Choice accuracy | 62.5% | 65.8% |
| Noul accuracy | 81.3% | 83.2% |
| Score accuracy | 57.5% | 61.6% |
| Score expected-value MAE | 0.727 | 0.645 |
Hyperparameters: encoder LR 1e-5, head LR 2.5e-4, max sequence length 2048.
Input Sequence Format
The model expects a structured token sequence with trained special markers:
[DECISION] [STATE] <state text> [QUESTION] <question text> [OPT_0] <option 0> [OPT_1] <option 1> ...
Special tokens are pre-trained in the vocabulary at fixed IDs — they must not be added or recreated at runtime:
| Token | ID | Token | ID |
|---|---|---|---|
[DECISION] |
50368 | [OPT_0] – [OPT_9] |
50371 – 50380 |
[STATE] |
50369 | ||
[QUESTION] |
50370 |
ONNX Graph
Inputs (batch size 1, dynamic sequence length):
| Name | Type | Shape |
|---|---|---|
input_ids |
int64 | [1, sequence] |
attention_mask |
int64 | [1, sequence] |
decision_pos |
int64 | [1] |
option_pos |
int64 | [1, 10] |
option_mask |
bool | [1, 10] |
score_n |
int64 | [1] |
Outputs:
| Name | Type | Description |
|---|---|---|
choice_logits |
float32 | Logits over option slots |
noul_logits |
float32 | [no, yes] logits |
score_logits |
float32 | Logits over score levels |
Unused heads output a masked sentinel (float32 minimum).
Files
| File | Size | Description |
|---|---|---|
model_int8.onnx |
67.3 MB | INT8 quantized ONNX model |
tokenizer.json |
3.4 MB | BPE vocabulary (13 trained special tokens) |
tokenizer_config.json |
613 B | Tokenizer settings |
config.json |
1.1 KB | Training config, metrics, graph spec |
manifest.json |
293 B | Version, size, SHA-256 for the web app |
SHA-256 (model_int8.onnx): cfd0184cf450c7ceeaab10ef1a7069bd8924ba46b0d9f73fd13a6a51d9616bd1
Usage
Browser (Decider web app)
The companion web app downloads the model from this repo, caches it in IndexedDB, and runs inference entirely client-side:
// Programmatic API exposed on window
await downloadModel(); // streams from HF, verifies SHA-256, saves to IndexedDB
await loadModel(); // builds ONNX Runtime WASM session
const result = await predict({
type: "choice",
state: "Alice has two projects and limited time.",
question: "Which project should Alice choose?",
options: ["Project A", "Project B"],
});
console.log(result.choice); // logits → softmax for probabilities
Python (reference)
import onnxruntime as ort
session = ort.InferenceSession("model_int8.onnx")
print([inp.name for inp in session.get_inputs()])
# ['input_ids', 'attention_mask', 'decision_pos', 'option_pos', 'option_mask', 'score_n']
Limitations
- Single-threaded WASM inference is slow on sequences above ~512 tokens.
- INT8 quantization trades some accuracy for a 4× size reduction over FP32.
- The model was trained on synthetic decision scenarios; real-world calibration has not been evaluated.
License
Apache 2.0
- Downloads last month
- 33
Model tree for aravind1706/decider-ettin
Base model
jhu-clsp/ettin-encoder-68m