Publications (16)
iMaC: Translating Actions into Motion and Contact Images for Embodied World Models
Zhenyu Wu, Xiuwei Xu, Yukun Zhou +8
Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely o…
Towards Privacy-Preserving Fine-Grained Visual Classification via Hierarchical Learning from Label Proportions
Jinyi Chang, Dongliang Chang, Lei Chen +2
In recent years, Fine-Grained Visual Classification (FGVC) has achieved impressive recognition accuracy, despite minimal inter-class variations. However, existing methods heavily r…
DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection
Ruihao Xu, Yong Liu, Yansong Tang +6
With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocabulary aerial detection which i…
TCOVIS: Temporally Consistent Online Video Instance Segmentation
Junlong Li, Bingyao Yu, Yongming Rao +2
In recent years, significant progress has been made in video instance segmentation (VIS), with many offline and online methods achieving state-of-the-art performance. While offline…
ShapeGen: Robotic Data Generation for Category-Level Manipulation
Yirui Wang, Xiuwei Xu, Angyuan Ma +3
Manipulation policies deployed in uncontrolled real-world scenarios are faced with great in-category geometric diversity of everyday objects. In order to function robustly under su…
Toward Generalizable Forgery Detection and Reasoning
Yueying Gao, Dongliang Chang, Bingyao Yu +5
Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models…
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
Haotian Qin, Dongliang Chang, Yueying Gao +3
Although existing CLIP-based methods for detecting AI-generated images have achieved promising results, they are still limited by severe feature redundancy, which hinders their gen…
NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
Yifei Li, Haoyuan He, Yu Zheng +5
The accessibility surge and abuse risks of user-friendly image editing models have created an urgent need for generalizable, up-to-date methods for Image Manipulation Detection and…
F2F-AP: Flow-to-Future Asynchronous Policy for Real-time Dynamic Manipulation
Haoyu Wei, Xiuwei Xu, Ziyang Cheng +5
Asynchronous inference has emerged as a prevalent paradigm in robotic manipulation, achieving significant progress in ensuring trajectory smoothness and efficiency. However, a syst…
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
Deyi Zhu, Yuji Wang, Yong Liu +4
Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with…
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
Yanran Zhang, Wenzhao Zheng, Yifei Li +5
In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fie…
R2RDreamer: 3D-aware Data Augmentation for Spatially-generalized 2D Manipulation Policies
Xiuwei Xu, Haowen Sun, Angyuan Ma +7
Spatial generalization is critical for imitation-learned manipulation policies, but achieving it typically requires scaling demonstrations across diverse object poses, robot config…
R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
Xiuwei Xu, Angyuan Ma, Hankun Li +4
Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to work robustly under different spatial dis…
QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
Yanran Zhang, Bingyao Yu, Yu Zheng +5
The emergence of visual autoregressive (AR) models has revolutionized image generation while presenting new challenges for synthetic image detection. Unlike previous GAN or diffusi…
CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection
Ziyang Cheng, Haoyu Wei, Hang Yin +4
While decoupled control schemes for legged mobile manipulators have shown robustness, learning holistic whole-body control policies for tracking global end-effector poses remains f…
Learning Counterfactually Decoupled Attention for Open-World Model Attribution
Yu Zheng, Boyang Gong, Fanye Kong +6
In this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region part…