Decider Ettin 68M

A 68M-parameter decision engine fine-tuned from Ettin Encoder 68M on the BEV Decision 150K dataset. Exported to ONNX (opset 18) and quantized to INT8 for browser inference via ONNX Runtime Web (WASM).

Everything runs in the browser. The model is downloaded once, cached in IndexedDB, and never requires a server for inference.

Decision Heads

The model has three output heads activated by the input sequence structure:

Head Output Description
Choice choice_logits Pick one of up to 10 options
Noul noul_logits Binary yes / no verdict
Score score_logits Rate on a scale of up to 7 levels

Training Metrics

Metric Epoch 1 Epoch 2
Choice accuracy 62.5% 65.8%
Noul accuracy 81.3% 83.2%
Score accuracy 57.5% 61.6%
Score expected-value MAE 0.727 0.645

Hyperparameters: encoder LR 1e-5, head LR 2.5e-4, max sequence length 2048.

Input Sequence Format

The model expects a structured token sequence with trained special markers:

[DECISION] [STATE] <state text> [QUESTION] <question text> [OPT_0] <option 0> [OPT_1] <option 1> ...

Special tokens are pre-trained in the vocabulary at fixed IDs — they must not be added or recreated at runtime:

Token ID Token ID
[DECISION] 50368 [OPT_0] – [OPT_9] 50371 – 50380
[STATE] 50369
[QUESTION] 50370

ONNX Graph

Inputs (batch size 1, dynamic sequence length):

Name Type Shape
input_ids int64 [1, sequence]
attention_mask int64 [1, sequence]
decision_pos int64 [1]
option_pos int64 [1, 10]
option_mask bool [1, 10]
score_n int64 [1]

Outputs:

Name Type Description
choice_logits float32 Logits over option slots
noul_logits float32 [no, yes] logits
score_logits float32 Logits over score levels

Unused heads output a masked sentinel (float32 minimum).

Files

File Size Description
model_int8.onnx 67.3 MB INT8 quantized ONNX model
tokenizer.json 3.4 MB BPE vocabulary (13 trained special tokens)
tokenizer_config.json 613 B Tokenizer settings
config.json 1.1 KB Training config, metrics, graph spec
manifest.json 293 B Version, size, SHA-256 for the web app

SHA-256 (model_int8.onnx): cfd0184cf450c7ceeaab10ef1a7069bd8924ba46b0d9f73fd13a6a51d9616bd1

Usage

Browser (Decider web app)

The companion web app downloads the model from this repo, caches it in IndexedDB, and runs inference entirely client-side:

// Programmatic API exposed on window
await downloadModel();       // streams from HF, verifies SHA-256, saves to IndexedDB
await loadModel();           // builds ONNX Runtime WASM session

const result = await predict({
  type: "choice",
  state: "Alice has two projects and limited time.",
  question: "Which project should Alice choose?",
  options: ["Project A", "Project B"],
});

console.log(result.choice);  // logits → softmax for probabilities

Python (reference)

import onnxruntime as ort

session = ort.InferenceSession("model_int8.onnx")
print([inp.name for inp in session.get_inputs()])
# ['input_ids', 'attention_mask', 'decision_pos', 'option_pos', 'option_mask', 'score_n']

Limitations

  • Single-threaded WASM inference is slow on sequences above ~512 tokens.
  • INT8 quantization trades some accuracy for a 4× size reduction over FP32.
  • The model was trained on synthetic decision scenarios; real-world calibration has not been evaluated.

License

Apache 2.0

Downloads last month
33
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aravind1706/decider-ettin

Quantized
(8)
this model

Dataset used to train aravind1706/decider-ettin