61 citations · 111 across the 13 of their papers we have counts for
13 papers
VCoME: Verbal Video Composition with Multimodal Editing Effects
Weibo Gong, Xiaojie Jin, Xin Li +2
Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to…
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Haoji Zhang, Yiqin Wang, Yansong Tang +4
Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline…
Selective Feature Adapter for Dense Vision Transformers
Xueqing Deng, Qi Fan, Xiaojie Jin +2
Fine-tuning pre-trained transformer models, e.g., Swin Transformer, are successful in numerous downstream for dense prediction vision tasks. However, one major issue is the cost/st…
Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling
Xiaozheng Zheng, Zhuo Su, Chao Wen +2
To bridge the physical and virtual worlds for rapidly developed VR/AR applications, the ability to realistically drive 3D full-body avatars is of great significance. Although real-…
Delving Deeper into Data Scaling in Masked Image Modeling
Cheng-Ze Lu, Xiaojie Jin, Qibin Hou +3
Understanding whether self-supervised learning methods can scale with unlimited data is crucial for training large-scale models. In this work, we conduct an empirical study on the…
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Xingjian He, Sihan Chen, Fan Ma +7
Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…