EnergyDecision-DT (Legacy)
⚠️ OUTDATED — Superseded by both the modern v2 pretrained model AND the Stage C standalone DT.
This is the legacy 8×384 pretrained model. It was superseded by
mrvictoru/energydecision-dt-v2(8×768, GQA, RMSNorm, weight-tying), which in turn was superseded by the Stage C standalone DT — a transformer distilled from an honest SDP-planning teacher that beats PPO on all 4 identity surfaces and passes the impact gate:
Surface Stage C DT v2 pretrained (8×768) This model (8×384) PPO Standard Oct $11,573 $4,991 $2,678 $2,353 Dispatch-matched $35,320 $10,138 $8,242* $22,530 Expanded broad-2024 $34,761 $4,596 $2,459 $19,504 * Phase 1 GRPO result (overfit — collapses to $1,533 on standard surface).
This checkpoint is retained for reproducing the legacy study only. The shipped model is
models/aemo/dt/aemo_dt_sdp_jtsoc_fullcorpus.pt.
Model Description
EnergyDecision-DT is a Decision Transformer model trained on simulated battery dispatch data from the Australian Energy Market Operator (AEMO) Frequency Control Ancillary Services (FCAS) market. It models optimal battery dispatch as a sequence prediction problem, conditioning on returns-to-go, observed states, and past actions to predict the next action.
Source code: https://github.com/mrvictoru/energydecision
Key Features
- Action Space (9-dim):
- Dim 0: Energy dispatch in [-1, 1] (charge/discharge)
- Dims 1-8: FCAS contingency bids in [0, 1]
- State Space (18-dim): Normalized market observations
- Context Length: 180 timesteps (looks back ~15 hours of history)
Intended Use
This model is intended for:
- Research into offline RL for energy markets
- Simulation of battery trading strategies in the AEMO FCAS market
- Baseline for comparing decision transformer approaches against traditional RL
It is not intended for live trading without further validation, risk management, and regulatory compliance.
Training Data
- Source: AEMO simulated trade dataset
- Size: 75,945,600 rows across 2,405 episodes
- Source policies: A2C-generated trajectories (legacy FCAS-poor corpus)
Model Architecture
DecisionTransformer(
(embed_return): Linear(1 -> 384)
(embed_state): Linear(18 -> 384)
(embed_action): Linear(9 -> 384)
(embed_timestep): Embedding(100000 -> 384)
(blocks): 8x TransformerBlock(
(ln1): RMSNorm(384)
(attn): MultiheadAttention(384, 8 heads)
(ln2): RMSNorm(384)
(ffn): SwiGLU(384 -> 1536 -> 384)
(dropout): Dropout(p=0.15)
)
(ln_f): RMSNorm(384)
(predict_action): Linear(384 -> 384) -> GELU -> Linear(384 -> 9) -> Tanh
(predict_state): Linear(384 -> 18)
(predict_return): Linear(384 -> 1)
)
Hyperparameters
| Parameter | Value |
|---|---|
| Blocks | 8 |
| Hidden dim | 384 |
| Attention heads | 8 |
| Context length | 180 |
| Dropout | 0.15 |
| State dim | 18 |
| Action dim | 9 |
| Discount factor | 0.95 |
| Return scale | 2.0 |
Recommended RTG (legacy model)
The legacy model peaks at rtg_value=0.5. This is the inverse of the modern v2 model (which peaks at 0.0).
Usage
import torch
from huggingface_hub import hf_hub_download
from decision_transformer import DecisionTransformer
model_kwargs = {
"state_dim": 18, "act_dim": 9, "n_block": 8,
"h_dim": 384, "context_len": 180, "n_heads": 8,
"drop_p": 0.15, "max_timestep": 100000,
}
model_path = hf_hub_download("mrvictoru/energydecision-dt", "aemo_dt_fcas_model.pt")
model = DecisionTransformer(**model_kwargs)
model.load_from_checkpoint(model_path)
model.eval()
Citation
@misc{energydecision-dt,
author = {Victor U},
title = {EnergyDecision-DT: Decision Transformer for AEMO FCAS Battery Trading},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/mrvictoru/energydecision-dt}},
}