How a New Dataset Is Mapping Every Commute in England and Wales
1.2 billion weekly trips across England and Wales, mapped by mode, time, and purpose—finally giving planners the data they need.
1.2 billion weekly trips mapped across England and Wales with street-level precision.
In a typical week, 1.2 billion trips are made across England and Wales—each one a decision, a destination, a mode of transport. For decades, urban planners have had to guess at the details. Not anymore. A new dataset has reconstructed, with unprecedented resolution, how people move between every small neighborhood in the country, broken down by whether they’re walking to school, driving to work, or taking the train for leisure. The number? 7,264 neighborhoods, 49 million people, and 70 origin-destination matrices that capture travel by mode, time, and purpose—down to the morning rush hour. This isn’t just a technical achievement. It’s a tool that could reshape how cities plan transit, reduce congestion, and design for equity.
The Science
The study by Zhang et al. (2026) introduces a fused dataset that combines mobile network data from BT with the UK’s National Travel Survey (NTS), census demographics, and trip-rate models to generate high-resolution origin-destination (OD) matrices for England and Wales. These matrices quantify how many trips occur between each pair of Middle Layer Super Output Areas (MSOAs)—small statistical zones averaging 7,200 people—across the country. The result is a set of 70 open-access matrices, each capturing travel flows segmented by seven modes (walking, cycling, private car, motorcycle, bus, rail, subway), time periods, and, for the weekday morning peak, eight trip purposes.
The core challenge in transport planning has long been a mismatch between data sources. Surveys like the NTS offer rich detail on why people travel and how, but with limited spatial resolution due to small sample sizes—just 8,975 households in 2024, spread across thousands of zones. Meanwhile, passive data from mobile networks observe movement at scale, but lack demographic nuance and trip purpose. They also undercount children and misrepresent mode use, since a phone trace can’t distinguish a bus ride from a car trip with certainty.
The authors bridge this gap through a four-stage calibration pipeline, implemented in the open-source tool uk-travel-pipeline. First, they standardize BT’s mobile data—aggregated, anonymized records of adult trips between MSOAs—to a common geographic framework. Each trip is assigned to one of eight trip-length bands (e.g., under 1 mile, 1–2 miles) and one of nine English regions (plus Wales). Because mobile data only capture adult travelers, the team uplifts volumes at each origin using local child-to-adult population ratios from the census. The average uplift factor is 1.226, meaning children account for roughly 18.3% of trips at the origin level—though their travel patterns are assumed to mirror those of adults, a limitation acknowledged by the authors.
Next comes mode-share calibration. Mobile data infer mode from movement speed and patterns, but these inferences often misalign with real behavior. To correct this, the researchers use NTS data (Table NTS9916) to set target mode shares by region and trip length. For example, if mobile data suggest 90% of short trips in London are by car, but the NTS reports only 60%, the model adjusts downward. This calibration is applied in region–length–mode groups, ensuring that the final mode mix matches survey evidence while preserving the spatial structure of movement observed in mobile data.
The third stage splits the broad "road" category from the mobile data into four distinct modes—cycle, private car, motorcycle, and bus—using NTS-derived proportions. It also refines the time-of-day distribution of trips, aligning peak-period shares with survey data (from Transport for the North’s trip-rate tables). This ensures, for instance, that the proportion of morning peak car trips matches real-world commuting patterns, not just what mobile data happen to capture.
Finally, trip purpose is allocated using trip-production rates—how many trips people in different demographic groups are expected to make for different purposes (e.g., work, education, shopping). These rates, weighted by local population segments, create a "local prior" for purpose shares at each origin. An iterative raking process then adjusts these shares so that national totals align with NTS0502, a survey table reporting trip start times by purpose. This step ensures that, for example, 28% of weekday morning trips are for work (as observed in the survey), while still allowing variation by location—commuter towns may have higher work trip shares, university areas more education trips.
What They Found
The resulting dataset covers a 56-week window from September 2024 to September 2025, built from 229 usable weekdays and 97 weekend days of mobile observations. Two temporal groupings are released:
- Typical week: Average trips per seven-day week, aggregated across all times of day.
- Weekday AM peak: Average trips during the 07:00–09:59 morning rush.
For the typical week, seven all-purpose matrices are provided—one for each travel mode. For the AM peak, seven all-purpose matrices are joined by 56 purpose-specific ones (7 modes × 8 purposes), totaling 70 matrices. Each cell reports estimated person trips between MSOAs.
The most striking validation result is the alignment of mode shares with survey data. Before calibration, the mobile-derived base overestimated road travel at 88.8% and underestimated walking at 8.1%. After calibration, these shifted to 65.4% road and 31.0% walking—matching NTS benchmarks almost exactly. The mean absolute error in mode share across 427 region–length–mode cells dropped from 5.07 to 0.01 percentage points. For private car trips, the gap fell from 13.5 to 0.03 points; for walking, from 11.6 to less than 0.01.
Mode Share Before and After Calibration
Comparison of travel mode shares in the raw mobile data versus the calibrated dataset, showing dramatic correction of car overestimation and walking underestimation.
| Label | Value |
|---|---|
| Walking (before) | 8.1 |
| Walking (after) | 31 |
| Private Car (before) | 88.8 |
| Private Car (after) | 65.4 |
| Bus (before) | 1.2 |
| Bus (after) | 2.1 |
| Rail (before) | 1.5 |
| Rail (after) | 1.2 |
Figure 1a shows this dramatic realignment. The mobile data alone grossly overrepresent car use and undercount active travel—a well-known bias in passive mobility datasets. Calibration corrects this, anchoring the model in behavioral reality.
The purpose allocation also performs as intended. After raking to NTS0502 controls, the national share of morning peak trips by purpose matches survey evidence to within 0.01 percentage points. For instance, work trips—which dominate the AM peak—account for 58.3% of flows, education for 14.1%, and shopping/other personal business for 8.7%. These shares vary locally: in university towns like Oxford or Cambridge, education trips spike; in suburban commuter zones, work dominates.
Another key finding is the sheer density of the dataset. With 7,264 MSOAs, there are over 52 million possible origin-destination pairs. The matrices are "dense"—meaning nearly all pairs have non-zero flows—enabling fine-grained analysis. For example, planners can now estimate how many children walk to school in a specific neighborhood, or how many rail commuters enter a city from each surrounding town.
Why This Changes Things
This dataset is not just an incremental improvement. It represents a paradigm shift in transport planning—one that moves from coarse, infrequent snapshots to dynamic, granular, and open evidence.
Consider the policy implications. Local authorities in England are now required to conduct "place-based assessments" of mobility, recognizing that travel-to-work areas rarely align with administrative boundaries (Department for Transport, 2026). Yet without data at the neighborhood level, such assessments have been guesswork. Now, for the first time, a council in Milton Keynes can see exactly how many residents commute to London by rail, how many cycle to local jobs, and how many school trips could be made on foot with safer routes. This enables targeted interventions: a new bus route here, a bike lane there, a school travel plan tailored to actual demand.
The open-source nature of the pipeline is equally transformative. The uk-travel-pipeline tool allows users to re-run the model with different assumptions—say, projecting 2030 travel under new population forecasts or testing how a new rail line might shift mode shares. Because the code is transparent, assumptions are inspectable. Want to know why car trips dropped in a certain area? You can trace it back to the NTS mode share for that region and trip length.
Globally, this approach could be replicated. Many countries have mobile data and national travel surveys, but few have fused them at this scale. The UK’s MSOA framework—stable, detailed, and widely used—makes this possible. But the method could be adapted to census tracts in the US, dissemination areas in Canada, or neighborhoods in India. The key is not the geography, but the fusion logic: use passive data for spatial structure, surveys for behavioral realism.
Purpose Distribution in Weekday Morning Peak
Share of trips by purpose during the 07:00–09:59 period, showing dominance of work and education trips.
| Label | Value |
|---|---|
| Work | 58.3 |
| Education | 14.1 |
| Shopping/Personal Business | 8.7 |
| Health | 3.2 |
| Visiting Friends/Family | 4.1 |
| Social/Leisure | 5.4 |
| Accompanying/Collecting | 3.8 |
| Other | 2.4 |
The implications extend beyond transport. These matrices can feed into air quality models (by estimating vehicle kilometers), public health studies (by measuring active travel), and housing policy (by assessing accessibility). For example, a developer proposing a new housing estate can now quantify how many future residents would drive versus walk to nearby amenities—information critical for environmental impact assessments.
Critically, the dataset captures equity dimensions often missed in aggregate models. By including children and weighting by local demographics, it reveals disparities in access. A low-income neighborhood with poor bus service may show high car dependency not because residents prefer driving, but because alternatives are lacking. Planners can now see these gaps and design accordingly.
What's Next
The authors acknowledge several limitations. First, child travel is modeled as a scaled version of adult patterns, which may not reflect reality—children are more likely to be driven to school, for instance. Future versions could incorporate school travel surveys or trip-chaining data to refine this.
Second, the model assumes that mobile data accurately capture spatial flows, but coverage varies by operator and area. Rural regions with poor signal may be underrepresented. The team mitigates this by using BT’s aggregate product, which is designed to be representative, but validation remains indirect.
Third, trip purpose is allocated at the origin level, not per OD pair. This means all trips from a residential zone to any destination share the same purpose mix—reasonable for work trips, less so for leisure. A more sophisticated model might use destination land use (e.g., retail, parks) to refine purpose assignment.
Finally, the dataset is static—a snapshot of 2024–2025. Real cities evolve. The next step is to build dynamic versions that respond to events: a strike, a new bike lane, a pandemic. The pipeline is already structured for this; re-running it with updated mobile data could yield monthly or even weekly updates.
The biggest opportunity lies in integration. These matrices could be linked with real-time transit data, weather, or event calendars to create predictive models. Imagine a city dashboard that forecasts congestion every morning based on school terms, holidays, and planned events—powered by fused survey and mobile data.
For now, the release of 70 open matrices and a transparent pipeline marks a turning point. It proves that high-resolution, behaviorally grounded mobility data can be built without sacrificing privacy or accessibility. As urban populations grow and climate pressures mount, the ability to plan transport systems with precision isn’t just useful—it’s essential. This dataset doesn’t just map where people go. It helps us build places where fewer people need to drive—where walking, cycling, and transit aren’t afterthoughts, but the default.
Sign in to join the conversation.
Comments (0)
No comments yet. Be the first to share your thoughts.