5 papers
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Zepeng Wang, Jiagao Hu, Fuhao Li +3
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been ex…
Beyond Information Redundancy: Expanding Cross-Modal Knowledge Representation for Power Load Time Series Forecasting
Yuxuan Chen, Shuo Dai, Ruoyi Xu +1
Load forecasting is pivotal for stable power systems. Conventional uni-modal methods suffer from representation drift under data scarcity. While recent multi-modal approaches attem…
Dynamic Execution Commitment of Vision-Language-Action Models
Feng Chen, Xianghui Wang, Yuxuan Chen +4
Vision-Language-Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low-level actions in a single forward pas…
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
Fuhao Li, Shaofeng You, Jiagao Hu +6
Evaluating object removal in images and videos remains challenging because the task is inherently one-to-many, yet existing metrics frequently disagree with human perception. Full-…
From Ideal to Real: Stable Video Object Removal under Imperfect Conditions
Jiagao Hu, Yuxuan Chen, Fuhao Li +4
Removing objects from videos remains difficult in the presence of real-world imperfections such as shadows, abrupt motion, and defective masks. Existing diffusion-based video inpai…