MSc Data Science research, rebuilt as a production forecasting system
Forecasting congestion in 387 cities, 24 hours ahead.
Zahma predicts the TomTom congestion index hour by hour, with ranges that hold their promised coverage, explanations for every forecast, and a live model you can push around in your browser.
The map
Pick a city to see the last week, the next 24 hours and the model's uncertainty. Colour shows the chosen measure.
How accurate is it?
Every model was trained only on data before each of five forecast origins and scored on the next 24 hours: 46,000+ real hours across 387 cities. Imputed hours are never scored.
Leaderboard
Lower MAE is better. Coverage should match the promised level.
Error by forecast hour
MAE for each hour ahead, 1 to 24
Calibrated uncertainty
Share of real values inside each model's 90% range, against the range's width
Saudi cities
Riyadh, Jeddah, Mecca, Medina, Dammam: MAE by model
What drives a forecast
SHAP values from the LightGBM model on 6,000 sampled hours: how much each input moves the prediction, on average.
Most influential inputs
Effect of local hour
Average push from the hour-of-day input, in index points
Effect of weekday
Four traffic rhythms
Cities grouped by the shape of their average week, using dynamic time warping k-means. The groups overlap (silhouette ), so treat them as tendencies, not borders.
Dissertation 2.0
My MSc dissertation (UCAM, 2025) reported R² = 0.98 for jam length. Rebuilding it showed what that number measures, and what a genuine forecast achieves.
Methods and data
The full pipeline, from raw snapshots to the model running in this page.