14 papers
XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition
Kun-Yu Lin, Henghui Ding, Jia-Run Du +6
Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective vi…
ProEdit: Inversion-based Editing From Prompts Done Right
Zhi Ouyang, Dian Zheng, Xiao-Ming Wu +4
Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image in…
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
Jiaming Zhou, Ke Ye, Jiayi Liu +6
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However…
CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion
Xiaotong Lin, Tianming Liang, Jian-Fang Hu +5
3D human-object interaction (HOI) anticipation aims to predict the future motion of humans and their manipulated objects, conditioned on the historical context. Generally, the arti…
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
Tianming Liang, Kun-Yu Lin, Chaolei Tan +3
Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language un…
Distilling the Unknown to Unveil Certainty
Zhilin Zhao, Longbing Cao, Yixuan Zhang +2
Out-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper pr…