5 papers · 1 filter
Improving Human Image Animation via Semantic Representation Alignment
Chang Liu, Mengting Chen, Yixuan Huang +5
The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, especially when generating long…
Unsupervised Domain Adaptation via Similarity-based Prototypes for Cross-Modality Segmentation
Ziyu Ye, Chen Ju, Chaofan Ma +1
Deep learning models have achieved great success on various vision challenges, but a well-trained model would face drastic performance degradation when applied to unseen data. Sinc…
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
Haicheng Wang, Chen Ju, Weixiong Lin +4
Temporal sentence grounding aims to detect event timestamps described by the natural language query from given untrimmed videos. The existing fully-supervised setting achieves grea…
Multi-Modal Prototypes for Open-World Semantic Segmentation
Yuhuan Yang, Chaofan Ma, Chen Ju +4
In semantic segmentation, generalizing a visual system to both seen categories and novel categories at inference time has always been practically valuable yet challenging. To enabl…
DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
Haozhe Cheng, Cheng Ju, Haicheng Wang +5
As one of the fundamental video tasks in computer vision, Open-Vocabulary Action Recognition (OVAR) recently gains increasing attention, with the development of vision-language pre…