most citedVideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

15 citations · 47 across the 8 of their papers we have counts for

collaborators

8 papers

cs.IR20242 cited

Maximizing User Experience with LLMOps-Driven Personalized Recommendation Systems

Chenxi Shi, Penghao Liang, Yichao Wu +2

The integration of LLMOps into personalized recommendation systems marks a significant advancement in managing LLM-driven applications. This innovation presents both opportunities…

cs.CL20246 cited

Research on the Application of Deep Learning-based BERT Model in Sentiment Analysis

Yichao Wu, Zhengyu Jin, Chenxi Shi +2

This paper explores the application of deep learning techniques, particularly focusing on BERT models, in sentiment analysis. It begins by introducing the fundamental concept of se…

cs.CV2023

Speed Co-Augmentation for Unsupervised Audio-Visual Pre-training

Jiangliu Wang, Jianbo Jiao, Yibing Song +5

This work aims to improve unsupervised audio-visual pre-training. Inspired by the efficacy of data augmentation in visual contrastive learning, we propose a novel speed co-augmenta…

cs.CV20232 cited

TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale

Ziyun Zeng, Yixiao Ge, Zhan Tong +3

The ultimate goal for foundation models is realizing task-agnostic, i.e., supporting out-of-the-box usage without task-specific fine-tuning. Although breakthroughs have been made i…

cs.CV202315 cited

VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Limin Wang, Bingkun Huang, Zhiyu Zhao +5

Scale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video fo…

cs.CV20236 cited

SparseFormer: Sparse Visual Recognition via Limited Latent Tokens

Ziteng Gao, Zhan Tong, Limin Wang +1

Human visual recognition is a sparse process, where only a few salient visual cues are attended to rather than traversing every detail uniformly. However, most current vision netwo…