Command menu

Navigate and run quick actions

Back

02Time-series machine learning

Store Demand Forecasting

An end-to-end pipeline that turns raw sales logs into reliable multi-step demand forecasts, including for products that rarely sell.

15%

RMSE reduction across model families

<50ms

to generate thousands of multi-step forecasts

30 days

strictly held-out validation window

0

data leakage, enforced by temporal masking

Overview

0%

lower forecasting error (RMSE)

The pipeline ingests transactional sales data, rebuilds a continuous daily history, engineers leak-free time-series features, and trains gradient boosting models matched to how each item actually sells. A Streamlit dashboard lets anyone upload data, pick a store and item, and see forecasts against history.

(01)Features

What it does

Robust time-series pipeline

Raw transaction logs become continuous daily records. Missing dates are reindexed and zero-filled so the statistics stay valid.

Models matched to demand

Standard LightGBM handles dense demand. LightGBM Tweedie regression and a two-stage hurdle classifier handle zero-inflated items where more than 99% of days have no sales.

Automated feature engineering

Calendar indicators, shifted lags and rolling mean, standard deviation, min and max, all built with strict temporal masking.

Interactive dashboard

Upload a dataset, choose a store and item, set a horizon, and compare predictions with historical sales in Streamlit.

(02)Architecture

How it validates

  1. 1

    Chronological split

    The final 30 days are kept completely separate from the training window.

  2. 2

    Sparsity-aware modelling

    The model choice adapts to how intermittent each item’s demand is.

  3. 3

    Fast inference

    Measured prediction time stays under 50ms for thousands of multi-step forecasts.

(03)Tech stack

Built with

Data

  • Pandas
  • NumPy

Machine learning

  • LightGBM
  • Scikit-learn

Interface

  • Streamlit

Next project

Do It For Me (DIFM)