TinyBrainBot-100M-v3-Math

A 100M-parameter math/reasoning specialist, SFT'd from tinybrainbot-100m-v3-base on a math-heavy diet (order-of-operations, word problems, chain-of-thought). Beats Supra2-100M-Instruct on 5/7 benchmarks on the official EleutherAI LM-Eval Harness.

  • Architecture: Llama-compatible, 100.1M params (768/12L/12hยท4kv, ctx 1024, vocab 32k).
  • Chat template: <|user|>\n{msg}\n<|end|>\n<|assistant|>\n

Benchmarks (EleutherAI lm-eval, 0-shot, acc_norm; WinoGrande/MMLU = acc)

Benchmark This model Supra2-100M-Instruct
ARC-Easy 53.7 44.4
ARC-Challenge 29.0 24.7
OpenBookQA 32.0 30.4
PIQA 65.7 64.4
WinoGrande 50.9 50.5
HellaSwag 33.3 35.9
MMLU 25.7 25.8

Reproduce these numbers

EleutherAI lm-eval-harness v0.4.12, 0-shot, on the HF repo (not the GGUF โ€” llama.cpp's --multiple-choice path under-reports these tasks):

lm_eval --model hf \
  --model_args pretrained=nkthebass/tinybrainbot-100m-v3-math,dtype=float32 \
  --tasks hellaswag,arc_easy,arc_challenge,openbookqa,winogrande,piqa,mmlu \
  --num_fewshot 0 --batch_size 32

Metrics: acc_norm for HellaSwag / ARC / OpenBookQA / PIQA; acc for WinoGrande & MMLU.

What it does

  • โœ… Order-of-operations: correct โ€” e.g. What is 12 + 7 * 3? โ†’ "First: 7ร—3=21, then 12+21 = 33." (the base and general instruct both get this wrong.)
  • โœ… Reasoning: strongest ARC-Challenge of the family.
  • โš ๏ธ Word problems: attempts chain-of-thought but multi-step arithmetic is unreliable (100M ceiling).
  • โš ๏ธ Not conversational โ€” it specialized toward math; use tinybrainbot-100m-v3-instruct to chat.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("nkthebass/tinybrainbot-100m-v3-math")
model = AutoModelForCausalLM.from_pretrained("nkthebass/tinybrainbot-100m-v3-math")
prompt = "<|user|>\nWhat is 12 + 7 * 3?\n<|end|>\n<|assistant|>\n"
ids = tok(prompt, return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=60)[0][ids.shape[1]:], skip_special_tokens=True))

GGUF

An F16 GGUF is included (tinybrainbot-100m-v3-math-f16.gguf) for llama.cpp / Ollama / LM Studio, with the add_space_prefix=false + leading-space chat template baked in โ€” order-of-operations works faithfully from the GGUF.

Not safety-tuned.

Downloads last month
1,861
Safetensors
Model size
0.1B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support