02Time-series machine learning
Store Demand Forecasting
An end-to-end pipeline that turns raw sales logs into reliable multi-step demand forecasts, including for products that rarely sell.
15%
RMSE reduction across model families
<50ms
to generate thousands of multi-step forecasts
30 days
strictly held-out validation window
0
data leakage, enforced by temporal masking
Overview
0%
lower forecasting error (RMSE)
The pipeline ingests transactional sales data, rebuilds a continuous daily history, engineers leak-free time-series features, and trains gradient boosting models matched to how each item actually sells. A Streamlit dashboard lets anyone upload data, pick a store and item, and see forecasts against history.
(01)Features
What it does
Robust time-series pipeline
Raw transaction logs become continuous daily records. Missing dates are reindexed and zero-filled so the statistics stay valid.
Models matched to demand
Standard LightGBM handles dense demand. LightGBM Tweedie regression and a two-stage hurdle classifier handle zero-inflated items where more than 99% of days have no sales.
Automated feature engineering
Calendar indicators, shifted lags and rolling mean, standard deviation, min and max, all built with strict temporal masking.
Interactive dashboard
Upload a dataset, choose a store and item, set a horizon, and compare predictions with historical sales in Streamlit.
(02)Architecture
How it validates
- 1
Chronological split
The final 30 days are kept completely separate from the training window.
- 2
Sparsity-aware modelling
The model choice adapts to how intermittent each item’s demand is.
- 3
Fast inference
Measured prediction time stays under 50ms for thousands of multi-step forecasts.
(03)Tech stack
Built with
Data
- Pandas
- NumPy
Machine learning
- LightGBM
- Scikit-learn
Interface
- Streamlit
Next project