6 papers
XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition
Kun-Yu Lin, Henghui Ding, Jia-Run Du +6
Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective vi…
DynProto: Dynamic Prototype Evolution for Out-of-Distribution Detection
Yanqi Wu, Xinhua Lu, Runhe Lai +4
Recent studies show that using potential out-of-distribution (OOD) labels from large corpora as auxiliary information can improve OOD detection in vision-language models (VLMs). Ho…
Distilling the Unknown to Unveil Certainty
Zhilin Zhao, Longbing Cao, Yixuan Zhang +2
Out-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper pr…
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Yi-Xing Peng, Qize Yang, Yu-Ming Tang +4
Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption d…
ParGo: Bridging Vision-Language with Partial and Global Views
An-Lan Wang, Bin Shan, Wei Shi +7
This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous work…
Task-Oriented 6-DoF Grasp Pose Detection in Clutters
An-Lan Wang, Nuo Chen, Kun-Yu Lin +2
In general, humans would grasp an object differently for different tasks, e.g., "grasping the handle of a knife to cut" vs. "grasping the blade to hand over". In the field of robot…