6 papers
RISE: Adaptive Imagination for World Action Models
Hongbo Lu, Liang Yao, Chenghao He +5
World Action Models (WAMs) improve planning by incorporating future world evolution into action generation, yet existing methods allocate a fixed imagination budget to every scene.…
RemoteZero: Geospatial Reasoning with Zero Labels
Liang Yao, Fan Liu, Shengxiang Xu +4
Geospatial reasoning requires models to identify image regions that satisfy complex and often implicit user intents. Recent reinforcement learning approaches improve reasoning with…
RobustFlow: Towards Robust Agentic Workflow Generation
Shengxiang Xu, Jiayi Zhang, Shimin Di +5
The automated generation of agentic workflows is a promising frontier for enabling large language models (LLMs) to solve complex tasks. However, the empirical study reveals that ex…
Label-Noise Resistant Learning via Optimal Brain Damage Masking
Xinlei Zhang, Fan Liu, Chuanyi Zhang +4
Noisy labels are inevitable in real-world multimedia applications. Due to the strong memorization capacity of deep neural networks, these noisy labels cause significant performance…
UEMM-Air: Make Unmanned Aerial Vehicles Perform More Multi-modal Tasks
Liang Yao, Fan Liu, Shengxiang Xu +6
The development of multi-modal learning for Unmanned Aerial Vehicles (UAVs) typically relies on a large amount of pixel-aligned multi-modal image data. However, existing datasets f…
Boost UAV-based Ojbect Detection via Scale-Invariant Feature Disentanglement and Adversarial Learning
Fan Liu, Liang Yao, Chuanyi Zhang +4
Detecting objects from Unmanned Aerial Vehicles (UAV) is often hindered by a large number of small objects, resulting in low detection accuracy. To address this issue, mainstream a…