Spaces:
Running
Download CLAUDE.md from autogluon/fev-bench: direct link, hf CLI and curl.
- Browser
- Download file 4.76 kB
-
https://huggingface.co/spaces/autogluon/fev-bench/resolve/main/CLAUDE.md
- Command line
-
hf download hf://spaces/autogluon/fev-bench/CLAUDE.md
-
curl -L -o CLAUDE.md https://huggingface.co/spaces/autogluon/fev-bench/resolve/main/CLAUDE.md
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
fev-bench Leaderboard is a Streamlit web application displaying time series forecasting model evaluation results from the fev-bench benchmark. It evaluates 30+ forecasting models using multiple metrics (SQL, MASE, WQL, WAPE) across 100 benchmark tasks.
Common Commands
# Run the Streamlit app locally
uv run streamlit run fev-leaderboard-app.py --server.port=8501 --server.address=0.0.0.0
# Generate data from the autogluon/fev repo (required once before running locally)
uv run python save_tables.py [commit] # defaults to latest main
uv run python save_tables.py --fev-repo ../fev # use an existing local checkout
# Docker build and run
docker build -t fev-leaderboard .
docker run -p 8501:8501 fev-leaderboard
Note: Use uv run prefix for all Python commands in this project.
No test or lint frameworks are configured.
Architecture
fev-leaderboard-app.py # Main entry point (Streamlit multi-page router)
save_tables.py # Generates tables/ and src/vendor/ from a single fev commit (run at Docker build)
pages/
βββ fev_bench.py # Main leaderboard: computes leaderboard/pairwise, renders them
βββ about.py # Help page with links
src/
βββ utils.py # Visualization, formatting, model metadata lookups, color palette
βββ strings.py # UI text, metric descriptions, paper citations
βββ task_groups.py # Task groupings by frequency and domain
βββ vendor/ # GENERATED, gitignored: compute logic copied from the fev repo
βββ fev_bench_compute.py
tables/ # GENERATED by save_tables.py, gitignored
βββ summaries.csv # Raw evaluation summaries (leaderboards are computed from these)
βββ models.csv # Display metadata, from the fev repo's benchmarks/fev_bench/models.yaml
βββ pivot_*.csv # Per-task error matrices (+ `_baseline_imputed` / `_leakage_imputed`)
βββ source.json # Which fev commit the tables came from
Data flow: GitHub (autogluon/fev) β save_tables.py β tables/ + vendored compute β fev_bench.py
tables/ and src/vendor/ are not committed. The Dockerfile runs save_tables.py main at build time,
so rebuilding the Space (Settings β Factory rebuild) picks up whatever is on fev main. The About
page shows which fev commit the running Space was built from.
Leaderboard and pairwise tables are not precomputed. Win rate and skill score depend on which
models are included, and the sidebar lets readers include/exclude model types, so fev_bench.py
computes them per (task group, metric, included types) and caches with st.cache_data.
Key Modules
src/utils.py: Core module containing:
get_model_metadata(): Model attributes loaded fromtables/models.csv(organization, url, model_type, zero_shot, commercial_use, display_name)ALL_METRICS: Dict with SQL, MASE, WQL, WAPE definitionsformat_leaderboard(),construct_bar_chart(),construct_pairwise_chart(),construct_pivot_table(): Styling functionsCOLORS: Custom palette (purple, gold, silver, bronze)
src/strings.py: Documentation strings for metric formulas, win rate/skill score calculations, imputation strategies
Metrics
| Metric | Type | Description |
|---|---|---|
| SQL | Probabilistic | Scaled Quantile Loss (scale-invariant) |
| MASE | Point | Mean Absolute Scaled Error (scale-invariant) |
| WQL | Probabilistic | Weighted Quantile Loss (scale-dependent) |
| WAPE | Point | Weighted Absolute Percentage Error (scale-dependent) |
Model Types
Each submission declares a model_type in benchmarks/fev_bench/models.yaml in the fev repo:
pretrained, task-specific, statistical, system, or closed-api. This drives the name color
on the leaderboard. system and closed-api models are hidden until the reader opts in from the
sidebar, because their results cannot be reproduced independently or compared to single models.
Model metadata is never edited here: it is declared by submitters in the fev repo and pulled in by
save_tables.py. Models flagged hidden: true there are dropped before any metric is computed.
Imputation Strategy
- Failed tasks: Replaced with Seasonal Naive scores
- Leaky tasks (training corpus overlap for zero-shot models): Replaced with Chronos-Bolt scores
External References
- fev-bench paper: https://arxiv.org/abs/2509.26468
- fev library docs: https://autogluon.github.io/fev/latest/
- GitHub: https://github.com/autogluon/fev