2 papers
cs.LG2026
Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
Takumi Shioda, Kohei Terashima, Tatsuo Nagai
Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model predictive control and reinf…
cs.LG2026
Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control
Takumi Shioda, Kohei Terashima, Tatsuo Nagai
Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well requires planning hours ahead…