How this project was built
A final-year BSc Mathematics project applying Box-Jenkins time series analysis to a real operations problem: deciding when and how much to reorder.
Approach
The application is deliberately split so that no statistical work happens inside a web route. Flask handles requests and rendering only; every calculation lives in a separate package that can be imported and tested on its own.
- Reproducible by default
- The three source CSVs ship with the project, so the full pipeline runs offline without a Kaggle account.
- Parameters are tested, not assumed
- The differencing order comes from repeated ADF tests, and p and q are read from the ACF and PACF rather than hardcoded.
- Accuracy measured on unseen data
- A chronological train and test split keeps the headline error figure honest, since in-sample error always flatters a fitted model.
Analysis pipeline
Each stage below is a page in the application, so the method can be walked through in order.
-
01
Load and merge
Three CSVs joined on order and product keys, dates parsed, quantities coerced
-
02
Aggregate
Resample to month end, keep zero-demand months, trim partial first and last buckets
-
03
Describe
Mean, median, variance, skewness, kurtosis, coefficient of variation
-
04
Test stationarity
Augmented Dickey-Fuller against the unit-root null hypothesis
-
05
Difference
Apply the smallest d that rejects the null, if any is needed at all
-
06
Identify order
Read significant lags from the ACF and PACF to propose p and q
-
07
Estimate
Maximum likelihood fit, compare candidates by AIC and BIC
-
08
Forecast
Six periods ahead with a 95% prediction interval
-
09
Validate and stock
Held-out error metrics, then safety stock, reorder point and order quantity
Mathematical scope
- Time series theory. Stationarity, autocovariance, the backshift operator.
- Hypothesis testing. ADF statistic, critical values, p-value interpretation.
- Autocorrelation. ACF, PACF and Bartlett confidence bounds.
- Box-Jenkins. Identification, estimation by maximum likelihood, diagnostic checking.
- Forecast theory. h-step ahead prediction and interval construction.
- Inventory models. Wilson's EOQ formula and service-level safety stock.
- Probability. The normal distribution and z-scores behind service levels.
Tooling
| Flask 3.1 | Routing, sessions, templating |
|---|---|
| pandas 2.2 | Merging, resampling, descriptive statistics |
| statsmodels 0.14 | ADF test, ACF/PACF, ARIMA estimation |
| NumPy 2.1 | Array arithmetic and error metrics |
| SciPy 1.14 | Normal quantiles for service levels |
| Matplotlib 3.9 | All charts, rendered server side |
| MathJax 3 | Equation typesetting in the browser |