activity
20222024
most citedASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action Localization

5 citations · 5 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

GenRec: Unifying Video Generation and Recognition with Diffusion Models

Zejia Weng, Xitong Yang, Zhen Xing +2

Video diffusion models are able to generate high-quality videos by learning strong spatial-temporal priors on large-scale datasets. In this paper, we aim to investigate whether suc…

cs.CV2023

Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data

Zuxuan Wu, Zejia Weng, Wujian Peng +4

Despite significant results achieved by Contrastive Language-Image Pretraining (CLIP) in zero-shot image recognition, limited effort has been made exploring its potential for zero-…

cs.CV20231 cited

Towards Scalable Neural Representation for Diverse Videos

Bo He, Xitong Yang, Hanyu Wang +6

Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV, E-NeRV). W…

cs.CV20236 cited

Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight Optimization

Zejia Weng, Xitong Yang, Ang Li +2

Contrastive Language-Image Pretraining (CLIP) has demonstrated impressive zero-shot learning abilities for image understanding, yet limited effort has been made to investigate CLIP…

cs.CV2023

Vision Transformers Are Good Mask Auto-Labelers

Shiyi Lan, Xitong Yang, Zhiding Yu +3

We propose Mask Auto-Labeler (MAL), a high-quality Transformer-based mask auto-labeling framework for instance segmentation using only box annotations. MAL takes box-cropped images…

cs.CV20225 cited

ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action Localization

Bo He, Xitong Yang, Le Kang +3

Weakly-supervised temporal action localization aims to recognize and localize action segments in untrimmed videos given only video-level action labels for training. Without the bou…