papers

Publications (16)

cs.RO2026

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

Zhenyu Wu, Xiuwei Xu, Yukun Zhou +8

Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely o…

cs.CV2025

Towards Privacy-Preserving Fine-Grained Visual Classification via Hierarchical Learning from Label Proportions

Jinyi Chang, Dongliang Chang, Lei Chen +2

In recent years, Fine-Grained Visual Classification (FGVC) has achieved impressive recognition accuracy, despite minimal inter-class variations. However, existing methods heavily r…

cs.CV2026

DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection

Ruihao Xu, Yong Liu, Yansong Tang +6

With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocabulary aerial detection which i…

cs.CV2023

TCOVIS: Temporally Consistent Online Video Instance Segmentation

Junlong Li, Bingyao Yu, Yongming Rao +2

In recent years, significant progress has been made in video instance segmentation (VIS), with many offline and online methods achieving state-of-the-art performance. While offline…

cs.RO2026

ShapeGen: Robotic Data Generation for Category-Level Manipulation

Yirui Wang, Xiuwei Xu, Angyuan Ma +3

Manipulation policies deployed in uncontrolled real-world scenarios are faced with great in-category geometric diversity of everyday objects. In order to function robustly under su…

cs.CV2026

Toward Generalizable Forgery Detection and Reasoning

Yueying Gao, Dongliang Chang, Bingyao Yu +5

Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models…

cs.CV2025

Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection

Haotian Qin, Dongliang Chang, Yueying Gao +3

Although existing CLIP-based methods for detecting AI-generated images have achieved promising results, they are still limited by severe feature redundancy, which hinders their gen…

cs.CV2026

NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization

Yifei Li, Haoyuan He, Yu Zheng +5

The accessibility surge and abuse risks of user-friendly image editing models have created an urgent need for generalizable, up-to-date methods for Image Manipulation Detection and…

cs.RO2026

F2F-AP: Flow-to-Future Asynchronous Policy for Real-time Dynamic Manipulation

Haoyu Wei, Xiuwei Xu, Ziyang Cheng +5

Asynchronous inference has emerged as a prevalent paradigm in robotic manipulation, achieving significant progress in ensuring trajectory smoothness and efficiency. However, a syst…

cs.CV2026

Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

Deyi Zhu, Yuji Wang, Yong Liu +4

Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with…

cs.CV2026

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

Yanran Zhang, Wenzhao Zheng, Yifei Li +5

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fie…

cs.RO2026

R2RDreamer: 3D-aware Data Augmentation for Spatially-generalized 2D Manipulation Policies

Xiuwei Xu, Haowen Sun, Angyuan Ma +7

Spatial generalization is critical for imitation-learned manipulation policies, but achieving it typically requires scaling demonstrations across diverse object poses, robot config…

cs.RO2026

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

Xiuwei Xu, Angyuan Ma, Hankun Li +4

Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to work robustly under different spatial dis…

cs.CV2025

QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection

Yanran Zhang, Bingyao Yu, Yu Zheng +5

The emergence of visual autoregressive (AR) models has revolutionized image generation while presenting new challenges for synthetic image detection. Unlike previous GAN or diffusi…

cs.RO2026

CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection

Ziyang Cheng, Haoyu Wei, Hang Yin +4

While decoupled control schemes for legged mobile manipulators have shown robustness, learning holistic whole-body control policies for tracking global end-effector poses remains f…

cs.CV2025

Learning Counterfactually Decoupled Attention for Open-World Model Attribution

Yu Zheng, Boyang Gong, Fanye Kong +6

In this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region part…