ylecun/mnist
Viewer β’ Updated β’ 70k β’ 84k β’ 277
A small CNN trained from random initialization (no pretrained weights) on
ylecun/mnist for ~1 minute of GPU time.
| Test accuracy | 0.9919 (9,919 / 10,000) |
| Training steps | 22,814 (β24 passes over the 60k train set) |
| Training time | 56.1 s |
| Throughput | ~414 steps/s |
| Hardware | 1Γ Nvidia T4 (small), $0.40/hr |
| Final train loss | ~0.0002 (converged) |
| Setting | Value |
|---|---|
| Dataset | ylecun/mnist, train split (60,000) β test split (10,000) |
| Model | 2 conv layers (16, 32 channels) + FC(1568β128β10) = 206,922 params |
| Optimizer | Adam, lr 1e-3 |
| Loss | Cross-entropy |
| Batch size | 64 |
| Seed | 0 |
| Pixels | float32, scaled to [0, 1] |
The training loop was wall-clock bounded (TRAIN_SECONDS=55) with the shuffled
DataLoader cycled, so the run ends on the clock rather than after a fixed number
of epochs. An earlier version of the script stopped after one epoch (938 steps,
10.7 s, 0.9801 test accuracy); the weights here are from the full-55-second run.
import torch
import torch.nn as nn
from huggingface_hub import hf_hub_download
import numpy as np
class Net(nn.Module):
def __init__(self):
super().__init__()
self.conv = nn.Sequential(
nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
)
self.head = nn.Sequential(
nn.Flatten(), nn.Linear(32 * 7 * 7, 128), nn.ReLU(), nn.Linear(128, 10),
)
def forward(self, x):
return self.head(self.conv(x))
model = Net()
path = hf_hub_download("abidlabs/mnist-from-scratch-1min", "pytorch_model.bin")
model.load_state_dict(torch.load(path, map_location="cpu"))
model.eval()
# x: float32 tensor of shape (N, 1, 28, 28), pixel values in [0, 1]
# logits = model(x)
Run metadata (steps, timing, accuracy, seed) is in training_meta.json.