- Research Article
- 10.1016/j.enbuild.2026.117030
Adaptive Thermostat Preference Learning Using Behaviour Nudging and Multi-Armed Bandits: A Field Implementation
- Jan 01, 2026
- Energy and Buildings
- Hussein Elehwany + 6 more +6
Occupant behaviour (OB) centric controls have significant potential in advancing next-generation HVAC systems. Many OB-centric control studies solicit feedback from occupants to tackle the thermal preference learning problem. Behaviour nudging was also implemented in various systems to influence occupant behaviour to be more energy efficient. This study addresses the gap of using behaviour nudging and unsolicited occupant thermostat overrides to learn their thermal preferences. A multi-armed bandit (MAB) reinforcement learning (RL) was used to learn occupant thermal preferences from their thermostat interactions. The reward signal of the algorithm was designed to reward energy savings and penalize discomfort. The occupants were continuously nudged by slowly reducing the zone setpoint during the heating season, to encourage them to override the thermostats. The algorithm was implemented in two zones with multiple occupants in an academic facility in Ottawa, Canada, achieving energy savings of up to 12.7% compared to static setpoints.
Read more