9 papers
CLIP-Guided Unsupervised Semantic-Aware Exposure Correction
Puzhen Wu, Han Weng, Quan Zheng +5
Improper exposure often leads to severe loss of details, color distortion, and reduced contrast. Exposure correction still faces two critical challenges: (1) the ignorance of objec…
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
Yunzhong Xiao, Yangmin Li, Hewei Wang +2
Agents utilizing tools powered by large language models (LLMs) or vision-language models (VLMs) have demonstrated remarkable progress in diverse tasks across text and visual modali…
Towards Understanding Camera Motions in Any Video
Zhiqiu Lin, Siyuan Cen, Daniel Jiang +12
We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, an…
Segregation and Context Aggregation Network for Real-time Cloud Segmentation
Yijie Li, Hewei Wang, Jiayi Zhang +5
Cloud segmentation from intensity images is a pivotal task in atmospheric science and computer vision, aiding weather forecasting and climate analysis. Ground-based sky/cloud segme…
RAINER: A Robust Ensemble Learning Grid Search-Tuned Framework for Rainfall Patterns Prediction
Zhenqi Li, Junhao Zhong, Hewei Wang +6
Rainfall prediction remains a persistent challenge due to the highly nonlinear and complex nature of meteorological data. Existing approaches lack systematic utilization of grid se…
CP2M: Clustered-Patch-Mixed Mosaic Augmentation for Aerial Image Segmentation
Yijie Li, Hewei Wang, Jinfeng Xu +4
Remote sensing image segmentation is pivotal for earth observation, underpinning applications such as environmental monitoring and urban planning. Due to the limited annotation dat…