3 papers
cs.LG2026
AdamO: A Collapse-Suppressed Optimizer for Offline RL
Nan Qiao, Sheng Yue, Shuning Wang +1
Offline reinforcement learning (RL) can fail spectacularly when bootstrapped temporal-difference (TD) updates amplify their own errors, driving the critic toward extreme and unusab…
cs.CV2026
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
Zhiheng Wu, Tong Wang, Shuning Wang +2
Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations,…
cs.LG2026
Cloud-Edge Collaborative Large Models for Robust Photovoltaic Power Forecasting
Nan Qiao, Shuning Wang, Sijing Duan +5
Photovoltaic (PV) power forecasting in edge-enabled grids requires balancing forecasting accuracy, robustness under weather-driven distribution shifts, and strict latency constrain…