← News
Tech for Good Tech for Good Frontiers

A Neural Controller That Learns When to Flip the Switch — and Save 30% on Energy Costs

A Neural Controller That Learns When to Flip the Switch — and Save 30% on Energy Costs
30% Energy cost reduction
RSE SpA, Italy Location
Neural Controller Technology
District Heating Network System type

30% lower energy costs — not from new hardware, but from a neural controller that learns when to switch boilers on and off.

That’s the result reported by Nicolas Kirsch and colleagues in their 2026 paper on Differentiable Hybrid-Action Neural Feedback Control (Kirsch et al., 2026). In a high-fidelity simulation of a real district heating network in Italy, their system slashed operating expenses under dynamic electricity pricing — outperforming decades-old industrial rules by a wide margin. The secret? A single neural network that doesn’t just tune temperatures, but also decides which heaters should be running at any given moment — and does so in a way that can be trained entirely with gradients.

This isn’t incremental improvement. It’s a shift in how we think about controlling complex physical systems: not as sequences of hand-coded logic, but as learnable policies that internalize both when to act and how much to act.

The Science

District heating networks (DHNs) are vast underground circulatory systems that pump hot water to homes, hospitals, and factories across cities. They’re common in Europe and parts of Asia, where centralized heat production offers economies of scale and easier decarbonization than millions of individual gas boilers. But operating them efficiently is fiendishly hard.

At RSE SpA in northern Italy — the site of the study — engineers manage multiple heat sources: gas-fired boilers, electric heat pumps, and large insulated tanks storing thermal energy. Each has different costs, response times, and constraints. Heat pumps are efficient but slow; boilers respond quickly but emit CO₂; storage lets you arbitrage between cheap nighttime power and expensive daytime demand.

The challenge lies in the hybrid nature of control: every minute, operators must decide which units to activate (a discrete choice: on/off/mode) and at what level they should run (a continuous setting: temperature, flow rate). These decisions are coupled. Turning on a heat pump affects how much storage you need. Running boilers at partial load changes efficiency. And everything depends on forecasts of demand and electricity prices — which are uncertain.

Traditionally, such systems are managed by rule-based controllers. One common strategy: "If outdoor temperature < 10°C, turn on boiler A; if price > €0.20/kWh, avoid resistive heaters." These heuristics are simple, interpretable, and robust — but rigid. They can’t adapt to unusual weather patterns or exploit subtle price signals. Worse, they often miss arbitrage opportunities embedded in time-of-use tariffs.

Model Predictive Control (MPC) offers a smarter alternative, optimizing over a forecast horizon using mathematical models. But when discrete decisions enter the picture, the problem becomes a Mixed-Integer Nonlinear Program (MINLP) — notoriously slow to solve, especially in real time. Industrial deployments either simplify the physics or shorten the planning window, sacrificing performance.

Reinforcement learning (RL) has emerged as a promising workaround, learning policies offline so deployment is fast. But most RL methods struggle with hybrid action spaces — particularly ensuring that outputs respect physical constraints. And model-free RL requires massive amounts of trial-and-error data, risky in safety-critical infrastructure.

Enter HANC: the Hybrid-Action Neural Controller. Unlike previous approaches, HANC is designed from the ground up to handle both discrete and continuous commands within a single, differentiable framework. Its innovation isn’t just architectural — it’s philosophical. Instead of treating switching as a black box or approximating it post-hoc, HANC makes discrete decisions learnable through gradients.

The core idea builds on the Gumbel-Softmax trick — a technique from machine learning that allows sampling from categorical distributions in a nearly differentiable way. During training, instead of picking the most likely mode via argmax (which kills gradients), the network uses a soft approximation driven by injected noise. Gradients flow through this softened version. In the forward pass, however, it still picks a single, hard action — preserving physical realizability.

Crucially, HANC trains via backpropagation through time (BPTT) over full closed-loop rollouts — meaning it learns how today’s switching decision impacts system state tomorrow, next week, even next month. This long-horizon credit assignment is essential for systems with thermal inertia, where turning on a heater now might satisfy demand 12 hours later.

And unlike many learned controllers, HANC operates purely as a feedback policy: it uses only current measurements and internal memory, not future forecasts. That makes it robust to prediction errors — a major weakness of MPC.

Fig. 3: Schematic layout of the district heating network considered in this
work. Red and blue arrows indicate the nominal flow directions in the supply
and return networks, respectively. Green arrows denote the control variables,
whereas orange arrows denote the exogenous thermal-demand disturbances.
Fig. 3: Schematic layout of the district heating network considered in this work. Red and blue arrows indicate the nominal flow directions in the supply and return networks, respectively. Green arrows denote the control variables, whereas orange arrows denote the exogenous thermal-demand disturbances. Source: Nicolas Kirsch, Corrado Sgadari

Fig. 3: Schematic layout of the district heating network considered in this work. Red and blue arrows indicate the nominal flow directions in the supply and return networks, respectively. Green arrows denote the control variables, whereas orange arrows denote the exogenous thermal-demand disturbances.

What They Found

When tested in simulation against a realistic industrial baseline, HANC delivered a 30% reduction in operating cost — a staggering improvement for any engineered system, let alone one already considered mature.

This wasn’t achieved by pushing equipment harder or cutting corners. It came from smarter coordination: delaying heat pump activation until off-peak prices, pre-charging thermal storage during low-cost periods, and minimizing simultaneous use of inefficient backup boilers.

But perhaps more surprising was what happened when the researchers compared two versions of HANC: one trained deterministically (using a fixed threshold for mode selection), and another with Gumbel noise injected during training.

Both achieved similar cost savings — around 30%. But the noisy version made far fewer hard switches. Where the deterministic policy flipped units on and off dozens of times per day near decision boundaries, the stochastic variant stayed put unless the signal was strong.

According to the study, the number of daily on/off cycles dropped by an order of magnitude — from ~36 to just ~3.5 across all generation units

Switching Frequency per Day

The number of full on/off cycles per day across all generation units. Standard rule-based control switches frequently due to fixed thresholds. Deterministic HANC reduces this moderately. With Gumbel noise injection, switching drops tenfold — indicating wider decision margins and more stable operation.

Switching Frequency per Day
LabelValue
Rule-Based48
HANC Deterministic36
HANC + Gumbel Noise3.5

. This suggests that noise injection encourages wider decision margins: the policy becomes more confident in its choices, avoiding jitter near thresholds.

Think of it like driving a car with cruise control. A poorly tuned system might oscillate rapidly between accelerating and coasting as speed fluctuates around the setpoint. A well-designed one maintains steady throttle unless deviation is significant. HANC with Gumbel noise behaves like the latter — smoother, more reliable, and less stressful on machinery.

The authors attribute this to the exploration induced by noise during training. By occasionally sampling suboptimal modes, the network learns to distinguish clear regimes from ambiguous ones — effectively building hysteresis into its decision-making without explicit programming.

Even more impressive: the policy was trained on one model of the plant dynamics but evaluated on a different, higher-fidelity simulator. This simulates real-world deployment, where no model perfectly captures reality. Yet HANC generalized successfully — proof that it didn’t just memorize solutions, but learned transferable strategies.

Why This Changes Things

Energy systems are entering a new era. Decarbonization means replacing steady fossil-fuel plants with intermittent renewables. Markets are shifting from flat rates to dynamic pricing, rewarding flexibility. And digitalization is making sensors and actuators ubiquitous.

Yet our control systems lag behind. Most still rely on 20th-century logic: IF-THEN rules, fixed schedules, manual overrides. We’ve digitized the pipes, but not the intelligence.

HANC represents a step toward truly adaptive infrastructure — systems that don’t just react, but anticipate; that don’t follow scripts, but learn.

Its implications extend far beyond heating. Any system with hybrid decisions could benefit: power grids managing generator commitment, data centers switching cooling modes, manufacturing lines reconfiguring assembly paths, or autonomous vehicles choosing drive modes.

Consider electric arc furnaces in steelmaking: they consume enormous power, often under time-of-use contracts. Deciding when to melt scrap metal — and at what intensity — is exactly the kind of hybrid problem HANC excels at. A 30% reduction in electricity cost would translate to millions in savings — and lower emissions.

Or take building HVAC systems. Most commercial buildings use simple thermostats or basic automation. But with variable tariffs and rooftop solar, there’s huge value in coordinating when to cool, heat, store, or idle. HANC-style controllers could turn passive buildings into active participants in grid balancing.

What sets HANC apart from other AI controllers is its constraint-aware design. Many deep reinforcement learning systems violate physical limits unless heavily penalized. HANC avoids this by construction: its assembly layer ensures every output respects actuator bounds, logical dependencies, and coupling constraints. You can’t accidentally command a valve to open 120% — because the math won’t let you.

Moreover, it eliminates the need for online optimization solvers. Once trained, HANC runs as a simple neural net — deployable on edge devices, scalable to thousands of nodes, and immune to solver failures.

Compare that to traditional MPC: each step requires calling a solver, which may fail to converge, return suboptimal solutions, or take too long. In safety-critical settings, that unpredictability is unacceptable. HANC trades upfront training cost for rock-solid, real-time execution.

And because it learns end-to-end over long horizons, it discovers non-intuitive strategies humans might miss. For instance, the policy might learn to slightly overheat a building before a price spike, banking thermal inertia as a form of energy storage. Or it might stagger unit startups to avoid demand charges. These are subtle, emergent behaviors — invisible to modular control designs.

What's Next

No breakthrough comes without caveats. HANC still requires a reasonably accurate model for training — though not perfect. The paper shows robustness to model mismatch, but extreme discrepancies could degrade performance. There’s also the question of safety certification: regulators may hesitate to approve black-box controllers for critical infrastructure.

One path forward is interpretability. While HANC is a neural net, its structure is modular: separate branches for continuous and categorical decisions, transparent assembly logic. Future work could visualize attention weights, logit margins, or sensitivity maps to build trust.

Another direction is hierarchical integration. Rather than replace MPC entirely, HANC could serve as a fast, approximate policy within a broader supervisory framework — correcting for forecast errors or handling contingencies.

Scalability is another frontier. The current implementation focuses on a single network. But cities have multiple interconnected DHNs. Federated learning could allow decentralized training while preserving privacy and reducing communication overhead.

Perhaps most intriguing is the role of noise. The finding that Gumbel noise leads to smoother, more robust policies hints at a deeper principle: stochasticity as a regularizer for decision quality. Injecting randomness doesn’t just help exploration — it shapes the geometry of learned policies, widening gaps between choices. This insight could inform controller design across robotics, logistics, and finance.

Finally, real-world deployment awaits. Simulations are clean; field tests are messy. Sensors drift, actuators fail, weather surprises. Testing HANC on actual hardware — even in pilot form — will be the true test of its value.

But the potential is undeniable. As climate change forces us to squeeze every joule of efficiency from existing infrastructure, we can’t afford to leave 30% of savings on the table. HANC shows that sometimes, the biggest gains don’t come from building more, but from thinking smarter — letting machines learn the rhythms of energy, time, and thermodynamics in ways we never could.

The future of smart infrastructure may not be ruled by equations or experts — but by neural controllers that know when to flip the switch.