activity
20162024
most citedTraining Skinny Deep Neural Networks with Iterative Hard Thresholding Methods

61 citations · 111 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CV2024

VCoME: Verbal Video Composition with Multimodal Editing Effects

Weibo Gong, Xiaojie Jin, Xin Li +2

Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to…

cs.CV2024

Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Haoji Zhang, Yiqin Wang, Yansong Tang +4

Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline…

cs.CV2023

Selective Feature Adapter for Dense Vision Transformers

Xueqing Deng, Qi Fan, Xiaojie Jin +2

Fine-tuning pre-trained transformer models, e.g., Swin Transformer, are successful in numerous downstream for dense prediction vision tasks. However, one major issue is the cost/st…

cs.CV2023

Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling

Xiaozheng Zheng, Zhuo Su, Chao Wen +2

To bridge the physical and virtual worlds for rapidly developed VR/AR applications, the ability to realistically drive 3D full-body avatars is of great significance. Although real-…

cs.CV20234 cited

Delving Deeper into Data Scaling in Masked Image Modeling

Cheng-Ze Lu, Xiaojie Jin, Qibin Hou +3

Understanding whether self-supervised learning methods can scale with unlimited data is crucial for training large-scale models. In this work, we conduct an empirical study on the…

cs.CV202311 cited

VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Xingjian He, Sihan Chen, Fan Ma +7

Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…