Zahmaزحمة

MSc Data Science research, rebuilt as a production forecasting system

Forecasting congestion in 387 cities, 24 hours ahead.

Zahma predicts the TomTom congestion index hour by hour, with ranges that hold their promised coverage, explanations for every forecast, and a live model you can push around in your browser.

The map

Pick a city to see the last week, the next 24 hours and the model's uncertainty. Colour shows the chosen measure.

How accurate is it?

Every model was trained only on data before each of five forecast origins and scored on the next 24 hours: 46,000+ real hours across 387 cities. Imputed hours are never scored.

Leaderboard

Lower MAE is better. Coverage should match the promised level.

Error by forecast hour

MAE for each hour ahead, 1 to 24

Calibrated uncertainty

Share of real values inside each model's 90% range, against the range's width

Saudi cities

Riyadh, Jeddah, Mecca, Medina, Dammam: MAE by model

What drives a forecast

SHAP values from the LightGBM model on 6,000 sampled hours: how much each input moves the prediction, on average.

Most influential inputs

Effect of local hour

Average push from the hour-of-day input, in index points

Effect of weekday

Four traffic rhythms

Cities grouped by the shape of their average week, using dynamic time warping k-means. The groups overlap (silhouette ), so treat them as tendencies, not borders.

Dissertation 2.0

My MSc dissertation (UCAM, 2025) reported R² = 0.98 for jam length. Rebuilding it showed what that number measures, and what a genuine forecast achieves.

Methods and data

The full pipeline, from raw snapshots to the model running in this page.