The Air in the Room Is Sabotaging Physics' Most Stubborn Measurement

Gravity is the one force we still can't pin down. Every other fundamental constant — the speed of light, Planck's constant, the charge of the electron — has been measured to parts per billion or better. But Newton's , the number that describes how strongly every kilogram pulls on every other kilogram, still carries a relative uncertainty of about 22 parts per million. And there's something else strange: different laboratories measuring over the past four decades have kept producing results that disagree with each other more than their own careful error bars suggest they should.
One hundred readings of should cluster around a single value, like darts thrown by a sober player. Instead they scatter wider than the individual throws can explain. Physicists have long suspected the culprit is not the instrument itself but the room around it — the air, the building, the ground water, the people walking down the corridor.
In a new preprint, Jyotirmaya Mohanta and Yutaka Shikano of the University of Tsukuba take aim at one of the most subtle of these environmental nuisances: the fact that the atmosphere is never perfectly still. Columns of air thicken and thin as pressure waves roll through the laboratory, and because every bit of air has mass, those density ripples tug slightly harder on one side of a balance than the other. The researchers show that this "atmospheric Newtonian noise" is too weak to explain the historical scatter in measurements by itself — but it is creeping right up to the edge of becoming a problem, just as physicists prepare their most ambitious next-generation experiments.
The work is ultimately a warning, delivered with the patience and rigor of a metrologist rather than a prophet. It says: you can no longer assume the air around your experiment is harmless. You need to measure it, model it, and put a number on it.
\section*{The gravitational constant: a 200-year embarrassment}
Newton's law of gravitation is the oldest and most famous equation in physics — the one about apples and moons. But its central constant, , is notoriously the most difficult one to measure. The first successful measurement came in 1798, when Henry Cavendish hung a dumbbell of lead balls from a fine quartz fiber and watched them twist toward larger masses placed nearby (he was actually measuring the density of the Earth, but the method became the template for ).
The modern descendants of Cavendish's apparatus are essentially the same idea refined to an extraordinary degree: a torsion balance with its pendulum mass, a set of source masses that move past it to create a known gravitational pull, and an exquisitely sensitive measurement of how much the fiber twists. The torque on the pendulum is proportional to , so measuring the twist means measuring gravity's strength.
The problem is that the balance twists in response to everything that is gravitational, not just the source masses. The pendulum also feels the pull of the walls, the hill outside, the water table beneath the building, and the constantly shifting blobs of air around it. Most of these environmental effects can be suppressed or accounted for, but the air is special: you cannot build a shield against gravity. A lead wall stops radiation and sound, but it does nothing to stop a changing gravitational field, because gravity passes through everything.
\section*{The Science}
Mohanta and Shikano set out to build a rigorous, metrologically complete framework for this "atmospheric Newtonian noise" — the fluctuating gravitational gradient produced by moving air masses, and how it contaminates a torsion-balance measurement of .
The heart of their approach is taking seriously the official international standard for expressing measurement uncertainty, known as the GUM (Guide to the Expression of Uncertainty in Measurement). Most discussions of environmental noise in experiments are informal — a remark in a methods section, a hand-waved estimate. The authors insist that the atmospheric contribution must be propagated through the measurement equation itself, with the same mathematical discipline applied to instrumental noise.
The measurement model is deceptively simple. A torsion balance estimates as
where is the estimated gravitational torque and is a geometrical response coefficient. Any uncertainty in the torque estimate flows directly into uncertainty in . The question is: how much torque noise does the atmosphere inject?
The torque on the pendulum comes from the gravity-gradient tensor — essentially, how quickly the gravitational acceleration changes from one spot to another. For a balance oriented along the -axis and twisting about , the dominant coupling is to the shear gradient . A fluctuating air mass creates a fluctuating , which creates a fluctuating torque:
where is a geometry factor proportional to the pendulum's quadrupole moment.
To model the atmosphere, the authors use the standard trick from gravitational-wave detector physics of replacing the three-dimensional air mass with an effective surface density on the plane below the pendulum. Each spatial Fourier mode of that density field — think of it as a ripple with a particular wavelength and direction — produces a predictable shear gradient at the height of the balanceholly. The coupling is exponentially suppressed for short wavelengths (a small ripple barely reaches up to the pendulum) and strongly angle-dependent.
They then derive, in closed form, the "transfer functions" for two benchmark geometries. The first is a simple two-mass dumbbell, which serves as a worst-case reference: if a pendulum's quadrupole moment is not minimized, its atmospheric coupling looks roughly like this. The second is a perfectly symmetric four-mass cross, which by symmetry cancels its coupling to a uniform shear gradient entirely — an idealized limit of how much rejection is possible in principle.
\section*{What They Found}
The central quantity is a "baseline factor" that quantifies how well a given geometry suppresses long-wavelength environmental ripples (here is the spatial wavenumber and the pendulum half-length). For the dumbbell, this suppression falls off as
for long wavelengths. But for the symmetric cross, the cancellation is dramatically stronger:
That extra of suppression is the mathematical expression of symmetry doing real work: a symmetric arrangement cancels not just the leading gravitational gradient but the first several orders of its spatial variation.
Baseline suppression factor vs. dipole for dumbbell and symmetric cross
Illustration of baseline suppression factors for the two benchmark geometries in arbitrary units, showing the dramatically stronger suppression of the symmetric cross at long wavelengths
| Label | Value |
|---|---|
| Dumbbell A(kl) | 700 |
| Cross A_x(kl) | 880 |
| Dumbbell long-wavelength | 620 |
| Cross long-wavelength | 8,900 |
shows how the two geometries diverge as the wavelength of the environmental ripple grows.
But there is a catch, and it is the paper's most intellectually honest point. Symmetry-based suppression is not free. If the signal — the gravitational pull of the carefully arranged source masses you want to measure — couples through the same quadrupole channel as the environmental noise, then cancelling that channel cancels both together. The perfectly symmetric cross that rejects all environmental noise would also reject the signal of interest, leaving nothing to measure.
This is the fundamental trade-off at the heart of experiments: every design move that suppresses environmental coupling must be checked against how much signal it also suppresses. The authors are explicit that their symmetric cross is not a proposal for a real experiment but an idealized benchmark that clarifies the long-wavelength limitainer — a Rosetta stone for understanding what geometry can and cannot do.
The paper's second essential finding concerns time, not space. Environmental inputs come in two flavors. The first is stationary correlated noise — fluctuations with a fixed statistical character that can be described by a power spectral density. For these, the authors show that the averaging-down of uncertainty follows a familiar-looking law: longer integration times reduce the noise, but only up to the correlation time of the input
Averaging efficiency vs. integration time for correlated noise
Normalized standard uncertainty of the sample mean for a stationary OU process with correlation time Tc, showing the plateau where additional averaging buys little once the integration time exceeds the correlation time
| Label | Value |
|---|---|
| Short integration (T << Tc) | 100 |
| Intermediate integration | 40 |
| Long integration (T >> Tc) | 12 |
. Once the measurement runs longer than how long a typical atmospheric fluctuation persists, further averaging buys nothing — a plateau that any experimenter would recognize with dread.
The second flavor is non-stationary baseline drift — the slow, wandering shifts of the balance's equilibrium angle that have no mean-reverting timescale within a single run. The authors emphasize that these two must be treated completely differently: PSD-based propagation works for the stationary part, while drift requires separate models (random-walk, , or trend models) or conservative Type-B treatment. Conflating the two is a recurring source of under-estimated uncertainty in the field.
To get actual numbers, the authors close their master equation using the global infrasound pressure-noise envelope — the "new low and high noise models" (NLNM/NHNM) that seismologists use to describe ambient pressure fluctuations around the world. Plugging realistic atmospheric pressure benchmarks into their framework, they find equivalent gradient uncertainties of roughly to for integration times of to seconds at a pendulum height of about a meter, depending on the pressure envelope and the frequency-to-wavenumber mapping assumed.
The punchline appears in
Atmospheric Newtonian noise vs. present and target uncertainty floors
Comparison of relative uncertainty levels: present CODATA reference uncertainty (~2.2e-5) versus the atmospheric Newtonian noise benchmark range (10^-6 to 10^-7 relative), showing the atmospheric background is below current floors but approaches relevance at part-per-million targets
| Label | Value |
|---|---|
| CODATA relative uncertainty (2.2e-5) | 0.000022 |
| Atmospheric NN benchmark upper | 0.000001 |
| Atmospheric NN benchmark lower | 1e-7 |
| Next-gen target (1e-6) | 0.000001 |
: at present reference uncertainty levels — the current CODATA relative standard uncertainty of about — atmospheric Newtonian noise is comfortably below the floor. It does not explain the historical scatter in published values)R, and the authors are careful to say so plainly. But as future experiments push toward relative uncertainties at or below the part-per-million level, the atmospheric background becomes a relevant — possibly dominant — Type-B contribution; it is no longer something you can wave away.
\section*{Why This Changes Things}
What makes this paper valuable is not a spectacular new discovery but a discipline it imposes on an entire measurement field. The gravitational constant has been remeasured dozens of times with increasingly ingenious apparatus, yet the results stubbornly refuse to converge. The metrology community has begun to suspect that the problem is not the instruments but the invisible, unshieldable environment they sit inside.
This work provides the missing machinery for turning that suspicion into a calculation. It gives experimenters closed-form transfer functions, benchmark geometries, and a statistically rigorous framework for propagating atmospheric noise into the uncertainty budget. It tells a future -experimenter exactly what to measure: the pressure field around their apparatus, with an array of sensors at different positions, so that the spatial correlation structure — not just the temporal spectrum — can be characterized.
The paper's honest accounting of the trade-offs is also refreshing. It does not promise that clever geometry will solve everything. Symmetric designs reject noise but also reject signal — the engineered anisotropy that real experiments need lies somewhere between the dumbbell and the idealized cross, and finding it requires knowing the detailed mass distribution of a specific apparatus. The framework's job is to make that trade-off visible and quantifiable, not to eliminate it.
There is a deeper resonance here. The story of precision measurement is the story of learning that the "noise" is never inert — it is always some other piece of physics you haven't yet thought to model. The air in the room, once an assumed blank, becomes a measured, characterized, budgeted quantity. This is how metrology advances: not by making instruments quieter in isolation, but by learning to see the environment as part of the apparatus.
\section*{What's Next}
The authors are upfront about the limits of their benchmark. Their surface-density approximation is conservative in some respects (vertical atmosphere structure reduces coupling further than the model assumes) and optimistic in others (indoor pressure fields can exceed outdoor envelopes because of building acoustics and ventilation). The numbers should be read as an order-of-magnitude envelope, not a precise site prediction — that precision requires apparatus-specific field models that this framework is designed to allow, but does not itself provide.
Nor does the paper claim to resolve the long-standing scatter in values. It explicitly concludes that atmospheric Newtonian noise alone cannot account for it. That scatter presumably has multiple causes — unmodeled hydrology, building deformation, sources of non-stationary drift yet to be characterized. What this work does is remove one convenient suspect from the list while rigorously documenting exactly how much it contributes.
The forward path is clear enough. As experiments reach for parts-per-million precision in , and as the field contemplates the possibility of a new conceptual understanding of gravity that would require a more precisely known constant, the environment around the experiment stops being scenery and becomes data. Arrays of microbarometers arranged around a torsion balance, simultaneous gravity-gradient measurements, site-specific transfer functions — these become part of the measurement apparatus itself.
There is something quietly hopeful in this. The gravitational constant is the most stubborn number in physics, scarred by decades of experiments that refuse to agree. But each refinement in understanding what contaminates it is a step toward a day when the scatter finally resolves into a single, shared value — and when that happens, gravity's constant will finally hold still under our gaze.