A newer version of this model is available: mrvictoru/energydecision-dt-v2-sdp

EnergyDecision-DT (Legacy)

⚠️ OUTDATED — Superseded by both the modern v2 pretrained model AND the Stage C standalone DT.

This is the legacy 8×384 pretrained model. It was superseded by mrvictoru/energydecision-dt-v2 (8×768, GQA, RMSNorm, weight-tying), which in turn was superseded by the Stage C standalone DT — a transformer distilled from an honest SDP-planning teacher that beats PPO on all 4 identity surfaces and passes the impact gate:

Surface Stage C DT v2 pretrained (8×768) This model (8×384) PPO
Standard Oct $11,573 $4,991 $2,678 $2,353
Dispatch-matched $35,320 $10,138 $8,242* $22,530
Expanded broad-2024 $34,761 $4,596 $2,459 $19,504

* Phase 1 GRPO result (overfit — collapses to $1,533 on standard surface).

This checkpoint is retained for reproducing the legacy study only. The shipped model is models/aemo/dt/aemo_dt_sdp_jtsoc_fullcorpus.pt.


Model Description

EnergyDecision-DT is a Decision Transformer model trained on simulated battery dispatch data from the Australian Energy Market Operator (AEMO) Frequency Control Ancillary Services (FCAS) market. It models optimal battery dispatch as a sequence prediction problem, conditioning on returns-to-go, observed states, and past actions to predict the next action.

Source code: https://github.com/mrvictoru/energydecision

Key Features

  • Action Space (9-dim):
    • Dim 0: Energy dispatch in [-1, 1] (charge/discharge)
    • Dims 1-8: FCAS contingency bids in [0, 1]
  • State Space (18-dim): Normalized market observations
  • Context Length: 180 timesteps (looks back ~15 hours of history)

Intended Use

This model is intended for:

  • Research into offline RL for energy markets
  • Simulation of battery trading strategies in the AEMO FCAS market
  • Baseline for comparing decision transformer approaches against traditional RL

It is not intended for live trading without further validation, risk management, and regulatory compliance.

Training Data

  • Source: AEMO simulated trade dataset
  • Size: 75,945,600 rows across 2,405 episodes
  • Source policies: A2C-generated trajectories (legacy FCAS-poor corpus)

Model Architecture

DecisionTransformer(
  (embed_return): Linear(1 -> 384)
  (embed_state): Linear(18 -> 384)
  (embed_action): Linear(9 -> 384)
  (embed_timestep): Embedding(100000 -> 384)
  (blocks): 8x TransformerBlock(
      (ln1): RMSNorm(384)
      (attn): MultiheadAttention(384, 8 heads)
      (ln2): RMSNorm(384)
      (ffn): SwiGLU(384 -> 1536 -> 384)
      (dropout): Dropout(p=0.15)
    )
  (ln_f): RMSNorm(384)
  (predict_action): Linear(384 -> 384) -> GELU -> Linear(384 -> 9) -> Tanh
  (predict_state): Linear(384 -> 18)
  (predict_return): Linear(384 -> 1)
)

Hyperparameters

Parameter Value
Blocks 8
Hidden dim 384
Attention heads 8
Context length 180
Dropout 0.15
State dim 18
Action dim 9
Discount factor 0.95
Return scale 2.0

Recommended RTG (legacy model)

The legacy model peaks at rtg_value=0.5. This is the inverse of the modern v2 model (which peaks at 0.0).

Usage

import torch
from huggingface_hub import hf_hub_download
from decision_transformer import DecisionTransformer

model_kwargs = {
    "state_dim": 18, "act_dim": 9, "n_block": 8,
    "h_dim": 384, "context_len": 180, "n_heads": 8,
    "drop_p": 0.15, "max_timestep": 100000,
}

model_path = hf_hub_download("mrvictoru/energydecision-dt", "aemo_dt_fcas_model.pt")
model = DecisionTransformer(**model_kwargs)
model.load_from_checkpoint(model_path)
model.eval()

Citation

@misc{energydecision-dt,
  author = {Victor U},
  title = {EnergyDecision-DT: Decision Transformer for AEMO FCAS Battery Trading},
  year = {2026},
  publisher = {HuggingFace},
  howpublished = {\url{https://huggingface.co/mrvictoru/energydecision-dt}},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train mrvictoru/energydecision-dt