most citedRethinking Alignment in Video Super-Resolution Transformers

35 citations · 61 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV20231 cited

Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

Yicheng Xiao, Zhuoyan Luo, Yong Liu +5

Video Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as sim…

cs.CV2023

Efficient Text-Guided 3D-Aware Portrait Generation with Score Distillation Sampling on Distribution

Yiji Cheng, Fei Yin, Xiaoke Huang +5

Text-to-3D is an emerging task that allows users to create 3D content with infinite possibilities. Existing works tackle the problem by optimizing a 3D representation with guidance…

cs.CV20231 cited

TaleCrafter: Interactive Story Visualization with Multiple Characters

Yuan Gong, Youxin Pang, Xiaodong Cun +8

Accurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable…

cs.CV202314 cited

SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

Zhuoyan Luo, Yicheng Xiao, Yong Liu +5

This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction pr…

cs.CV2023

Meta-Auxiliary Network for 3D GAN Inversion

Bangrui Jiang, Zhenhua Guo, Yujiu Yang

Real-world image manipulation has achieved fantastic progress in recent years. GAN inversion, which aims to map the real image to the latent code faithfully, is the first step in t…

cs.CV2023

RIFormer: Keep Your Vision Backbone Effective While Removing Token Mixer

Jiahao Wang, Songyang Zhang, Yong Liu +6

This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs),…