Methodology

How this project was built

A final-year BSc Mathematics project applying Box-Jenkins time series analysis to a real operations problem: deciding when and how much to reorder.

Approach

The application is deliberately split so that no statistical work happens inside a web route. Flask handles requests and rendering only; every calculation lives in a separate package that can be imported and tested on its own.

Reproducible by default
The three source CSVs ship with the project, so the full pipeline runs offline without a Kaggle account.
Parameters are tested, not assumed
The differencing order comes from repeated ADF tests, and p and q are read from the ACF and PACF rather than hardcoded.
Accuracy measured on unseen data
A chronological train and test split keeps the headline error figure honest, since in-sample error always flatters a fitted model.

Analysis pipeline

Each stage below is a page in the application, so the method can be walked through in order.

  1. 01
    Load and merge

    Three CSVs joined on order and product keys, dates parsed, quantities coerced

  2. 02
    Aggregate

    Resample to month end, keep zero-demand months, trim partial first and last buckets

  3. 03
    Describe

    Mean, median, variance, skewness, kurtosis, coefficient of variation

  4. 04
    Test stationarity

    Augmented Dickey-Fuller against the unit-root null hypothesis

  5. 05
    Difference

    Apply the smallest d that rejects the null, if any is needed at all

  6. 06
    Identify order

    Read significant lags from the ACF and PACF to propose p and q

  7. 07
    Estimate

    Maximum likelihood fit, compare candidates by AIC and BIC

  8. 08
    Forecast

    Six periods ahead with a 95% prediction interval

  9. 09
    Validate and stock

    Held-out error metrics, then safety stock, reorder point and order quantity

Mathematical scope

  • Time series theory. Stationarity, autocovariance, the backshift operator.
  • Hypothesis testing. ADF statistic, critical values, p-value interpretation.
  • Autocorrelation. ACF, PACF and Bartlett confidence bounds.
  • Box-Jenkins. Identification, estimation by maximum likelihood, diagnostic checking.
  • Forecast theory. h-step ahead prediction and interval construction.
  • Inventory models. Wilson's EOQ formula and service-level safety stock.
  • Probability. The normal distribution and z-scores behind service levels.
Full derivations

Tooling

Flask 3.1Routing, sessions, templating
pandas 2.2Merging, resampling, descriptive statistics
statsmodels 0.14ADF test, ACF/PACF, ARIMA estimation
NumPy 2.1Array arithmetic and error metrics
SciPy 1.14Normal quantiles for service levels
Matplotlib 3.9All charts, rendered server side
MathJax 3Equation typesetting in the browser