35 citations · 61 across the 15 of their papers we have counts for
10 papers · 1 filter
Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
Yicheng Xiao, Zhuoyan Luo, Yong Liu +5
Video Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as sim…
Efficient Text-Guided 3D-Aware Portrait Generation with Score Distillation Sampling on Distribution
Yiji Cheng, Fei Yin, Xiaoke Huang +5
Text-to-3D is an emerging task that allows users to create 3D content with infinite possibilities. Existing works tackle the problem by optimizing a 3D representation with guidance…
TaleCrafter: Interactive Story Visualization with Multiple Characters
Yuan Gong, Youxin Pang, Xiaodong Cun +8
Accurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable…
SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation
Zhuoyan Luo, Yicheng Xiao, Yong Liu +5
This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction pr…
Meta-Auxiliary Network for 3D GAN Inversion
Bangrui Jiang, Zhenhua Guo, Yujiu Yang
Real-world image manipulation has achieved fantastic progress in recent years. GAN inversion, which aims to map the real image to the latent code faithfully, is the first step in t…
RIFormer: Keep Your Vision Backbone Effective While Removing Token Mixer
Jiahao Wang, Songyang Zhang, Yong Liu +6
This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs),…