Reading City Heights from Space: Why One Sensor Can't See Every Building

In a Brazilian city, an algorithm read satellite images and guessed how tall buildings are — within about five meters of the truth. That might sound unglamorous next to the usual bragging rights of machine learning, but the real achievement sits underneath: the model didn't just estimate heights, it told us why it reached its answers, neighborhood by neighborhood. And what it revealed upends a comfortable assumption many researchers have been leaning on.
For years, the received wisdom in remote sensing has been that more data — more satellite bands, more sensors, deeper neural networks — automatically means better maps. This study suggests otherwise. The same predictor that dominates in one part of a city is nearly useless in another. Footprint geometry rules for low-rise blocks. Shadow geometry takes over for tall, isolated towers. Spectral reflectance carries the day for the tallest buildings of all. No single sensor is uniformly best. The implications ripple far beyond one Brazilian metropolis: they touch how we count the world's material wealth, how we plan disaster response, and how we build the global datasets the Global South desperately needs.
The Science
The problem the researchers (Iablonovski et al., 2026) set out to solve is deceptively concrete: how tall is each building in a city? At the individual-footprint scale, this number matters enormously. It tells urban planners how much material — concrete, steel, brick — is locked up in the built environment, which is essential for "material stock accounting," the field that tracks how much of humanity's resources live in standing structures rather than in circulation. It tells disaster teams, after an earthquake or flood, which buildings are most likely to have collapsed, where victims might be trapped, and what the damage will cost.
In wealthy countries, this is often the job of airborne LiDAR — a laser scanner flown over the city that measures distance to the ground and to rooftops with centimeter precision. But LiDAR campaigns are expensive, logistically heavy, and urgent to repeat. In the Global South, where many of the fastest-growing cities on Earth are located, airborne LiDAR coverage is rare. Commercial very high-resolution satellite imagery is an option, but at prices that make city-wide, repeat coverage cost-prohibitive. So cities that arguably need building-height data most are exactly the ones that can't afford it.
The obvious workaround has been free satellite data. The European Space Agency's Sentinel-1 radar and Sentinel-2 optical satellites stream open data over the entire globe, continuously. A flurry of recent papers has used them to estimate building heights, with real success. But there's a ceiling: these products still come out too coarse-grained for material stock analysis, which needs answers at the level of individual buildings, not blocks.
This study's contribution is to push past that ceiling using two more data sources that are almost free — available to scientists under research licenses at no cost, even if not openly redistributable. TerraSAR-X StripMap is a German radar satellite offering much finer resolution than Sentinel-1. PlanetScope is a constellation of small optical satellites providing daily, meter-scale imagery of the planet. The authors folded all of it together — TerraSAR-X, PlanetScope, and Sentinel-1 — to predict building heights in a large Brazilian city, validating against a LiDAR reference dataset.
The "how" matters as much as the "what." Training any machine learning model on geographic data is complicated by spatial autocorrelation — the statistical tendency for things near each other to be similar. Buildings in the same neighborhood share soil, climate, construction era, and building codes. If a model is trained and tested on data from the same neighborhood, it can "cheat," learning to predict neighborhood identity rather than building height, and flatter its own accuracy. The researchers addressed this with a geographically weighted random forest — a model that fits locally, allowing the relationship between satellite features and building height to vary across space, and that is honest about the spatial structure of its training data.
What They Found
Against the LiDAR reference, the model achieved a root mean square error of 5.34 meters and an of 0.756. For context, measures how much of the variation in building height the model explains — 0.756 means it captures about three-quarters of the real-world differences between buildings. An RMSE of 5.34 meters means that, on average, the model's height guesses are off by roughly the height of a two-story building's worth of vertical distance. That's far from perfect, but for a free, satellite-only pipeline running across a whole city, it's a meaningful step past the coarse resolution of earlier Sentinel-only products.
The headline result, though, is not the aggregate accuracy — it's what happened when the researchers looked inside the model at its local feature importance. Machine learning models are often black boxes: they give an answer but not a reason. This work used local interpretability methods to ask, for each building, which input feature actually drove the prediction. The answer varied systematically across the urban landscape.
For low-rise buildings, the dominant predictor was footprint geometry — the shape and area of the building's outline as seen from above. This makes intuitive sense: low buildings tend to come in regular, predictable shapes, and their footprint size correlates with their height in ways the model can exploit.
For taller, more isolated structures, shadow-derived height took over. In optical satellite imagery, the length of a building's shadow tells you about its height, given knowledge of the sun's angle. Free-standing towers cast clean, readable shadows; crowded low-rise blocks don't. The model learned to lean on this signal where it was informative.
For the tallest buildings in the set, spectral reflectance — the pattern of light bouncing off roofs and facades — became the decisive predictor. Different roofing materials and heights show up in different wavelengths, and for the very tallest structures the model found that reflected light carried the most height information.
Meanwhile, the radar channels told their own story. Sentinel-1 backscatter (how strongly the ground reflects the radar pulse back to the satellite) and InSAR (interferometric synthetic aperture radar, which uses phase differences between radar images to measure surface displacement) occupied complementary spatial niches. In some urban contexts one was far more predictive; in others, the other. Crucially, and perhaps surprisingly, no single sensor was uniformly preferable across the full set of buildings.
Why This Changes Things
The most important implication is a correction to how the field has been thinking. The dominant trend in satellite-based earth observation over the last few years has been to throw everything into ever-larger neural networks and hope the model figures out the patterns on its own. These global, one-size-fits-all models are celebrated for their headline accuracy numbers. But this study shows that a global model is exactly the wrong tool for understanding why a prediction is made — and, by extension, for knowing when to trust it.
The finding that predictor dominance shifts across intra-urban contexts is a quiet warning. Imagine a disaster-response team relying on a single model to assess a city after an earthquake. If that model was trained on the global average, it may have tuned itself to the wrong signals in exactly the neighborhoods where it's needed most. A model that leans on footprint geometry might be excellent in low-rise residential areas and dangerously weak in the central business district with its isolated towers. Accuracy reported "on average" can mask catastrophic failure in the specific places that matter.
This is what the authors call "optioneering guidance" — a term borrowed from engineering, meaning the systematic evaluation of alternatives before committing to a design. For agencies building satellite-derived mapping products, the study offers a practical menu: if your target city is dominated by low-rise sprawl, invest in footprint data; if you care about tall isolated structures, prioritize shadow geometry; if you're tracking the tallest buildings, pay attention to spectral reflectance. The answer to "which sensor should I use?" is no longer a single name — it's "it depends," followed by a map of where each sensor earns its keep.
There's also a geopolitical dimension. Much of the debate about high-resolution earth observation revolves around access: commercial imagery is a tool of the wealthy, and the Global South is often locked out. This study is a small but meaningful act of leveling. Every data source it uses — TerraSAR-X StripMap and PlanetScope products under research licenses, Sentinel-1 openly — is available to academic scientists essentially anywhere. The pipeline it demonstrates isn't a proprietary secret; it's a recipe. A research group in Lagos, Jakarta, or São Paulo could replicate it with open tools and data they can actually obtain. That matters for material stock accounting — understanding the vast reserves of concrete and steel locked in the rapidly growing cities of the Global South, which is essential for any serious circular-economy or construction-demolition planning.
The disaster-assessment stakes are even more visceral. In the hours after a major earthquake, knowing which buildings are tall enough to have collapsed in particular ways helps triage search-and-rescue and estimate casualties. Free satellite data with honest, localized confidence estimates could be the difference between a city that can organize its own response and one that waits for outside help that arrives days too late.
What's Next
The study is honest about its limits. Four pages and three figures, intended for a conference (JURSE 2027), this is a proof of concept more than a finished product. The RMSE of 5.34 meters, while a real improvement, is still too coarse for fine-grained engineering decisions — you wouldn't use it to certify the structural integrity of an individual building. The geographic weighting that made the model spatially honest also makes it harder to generalize: a model fitted to one Brazilian city carries its own urban DNA, and it's not yet clear how its learned feature-importance patterns transfer to cities with different morphologies, materials, or climates.
The most exciting open questions follow directly from the paper's central finding. If predictor importance shifts so consistently and interpretably across urban contexts, can those shifts be predicted? Could a model learn to anticipate, from basic city statistics — building density, height distribution, land use — which satellite signal will be most informative before running the expensive computation? That would turn "optioneering" from a post-hoc analysis into a design tool.
There's also the question of time. Cities in the Global South are changing faster than satellite campaigns can keep up. Sentinel-1's radar is not affected by clouds or darkness, so it can produce reliable time series. The next step, logically, is longitudinal: not just "how tall is this building now," but "how fast is this city growing, and where?" A model that knows which signals to trust in which neighborhoods could track urban material accumulation year by year, feeding everything from climate-embodied-carbon accounting to infrastructure planning.
Perhaps the deepest takeaway is methodological humility. In an era of giant models and bigger claims, this paper quietly demonstrates that sometimes the most powerful thing a model can do is explain itself — to show its work, neighborhood by neighborhood, and admit that its strengths are situational. The buildings it measured in Brazil already exist, already cast their shadows, already stand as monuments to urbanization. But the way we understand them — and the billions of similar buildings across the Global South that have never been measured at all — just got a little more honest.