fev-bench / CLAUDE.md
shchuro's picture
Generate leaderboard data from fev main at build time
88036ff
|
Raw History Blame Contribute Delete
4.76 kB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

fev-bench Leaderboard is a Streamlit web application displaying time series forecasting model evaluation results from the fev-bench benchmark. It evaluates 30+ forecasting models using multiple metrics (SQL, MASE, WQL, WAPE) across 100 benchmark tasks.

Common Commands

# Run the Streamlit app locally
uv run streamlit run fev-leaderboard-app.py --server.port=8501 --server.address=0.0.0.0

# Generate data from the autogluon/fev repo (required once before running locally)
uv run python save_tables.py [commit]        # defaults to latest main
uv run python save_tables.py --fev-repo ../fev   # use an existing local checkout

# Docker build and run
docker build -t fev-leaderboard .
docker run -p 8501:8501 fev-leaderboard

Note: Use uv run prefix for all Python commands in this project.

No test or lint frameworks are configured.

Architecture

fev-leaderboard-app.py     # Main entry point (Streamlit multi-page router)
save_tables.py             # Generates tables/ and src/vendor/ from a single fev commit (run at Docker build)
pages/
β”œβ”€β”€ fev_bench.py           # Main leaderboard: computes leaderboard/pairwise, renders them
└── about.py               # Help page with links
src/
β”œβ”€β”€ utils.py               # Visualization, formatting, model metadata lookups, color palette
β”œβ”€β”€ strings.py             # UI text, metric descriptions, paper citations
β”œβ”€β”€ task_groups.py         # Task groupings by frequency and domain
└── vendor/                # GENERATED, gitignored: compute logic copied from the fev repo
    └── fev_bench_compute.py
tables/                    # GENERATED by save_tables.py, gitignored
β”œβ”€β”€ summaries.csv          # Raw evaluation summaries (leaderboards are computed from these)
β”œβ”€β”€ models.csv             # Display metadata, from the fev repo's benchmarks/fev_bench/models.yaml
β”œβ”€β”€ pivot_*.csv            # Per-task error matrices (+ `_baseline_imputed` / `_leakage_imputed`)
└── source.json            # Which fev commit the tables came from

Data flow: GitHub (autogluon/fev) β†’ save_tables.py β†’ tables/ + vendored compute β†’ fev_bench.py

tables/ and src/vendor/ are not committed. The Dockerfile runs save_tables.py main at build time, so rebuilding the Space (Settings β†’ Factory rebuild) picks up whatever is on fev main. The About page shows which fev commit the running Space was built from.

Leaderboard and pairwise tables are not precomputed. Win rate and skill score depend on which models are included, and the sidebar lets readers include/exclude model types, so fev_bench.py computes them per (task group, metric, included types) and caches with st.cache_data.

Key Modules

src/utils.py: Core module containing:

  • get_model_metadata(): Model attributes loaded from tables/models.csv (organization, url, model_type, zero_shot, commercial_use, display_name)
  • ALL_METRICS: Dict with SQL, MASE, WQL, WAPE definitions
  • format_leaderboard(), construct_bar_chart(), construct_pairwise_chart(), construct_pivot_table(): Styling functions
  • COLORS: Custom palette (purple, gold, silver, bronze)

src/strings.py: Documentation strings for metric formulas, win rate/skill score calculations, imputation strategies

Metrics

Metric Type Description
SQL Probabilistic Scaled Quantile Loss (scale-invariant)
MASE Point Mean Absolute Scaled Error (scale-invariant)
WQL Probabilistic Weighted Quantile Loss (scale-dependent)
WAPE Point Weighted Absolute Percentage Error (scale-dependent)

Model Types

Each submission declares a model_type in benchmarks/fev_bench/models.yaml in the fev repo: pretrained, task-specific, statistical, system, or closed-api. This drives the name color on the leaderboard. system and closed-api models are hidden until the reader opts in from the sidebar, because their results cannot be reproduced independently or compared to single models.

Model metadata is never edited here: it is declared by submitters in the fev repo and pulled in by save_tables.py. Models flagged hidden: true there are dropped before any metric is computed.

Imputation Strategy

  • Failed tasks: Replaced with Seasonal Naive scores
  • Leaky tasks (training corpus overlap for zero-shot models): Replaced with Chronos-Bolt scores

External References