activity
20212024
most citedReuse and Diffuse: Iterative Denoising for Text-to-Video Generation

9 citations · 58 across the 19 of their papers we have counts for

collaborators

19 papers

cs.CV2024

Learning to Rank Patches for Unbiased Image Redundancy Reduction

Yang Luo, Zhineng Chen, Peng Zhou +3

Images suffer from heavy spatial redundancy because pixels in neighboring regions are spatially correlated. Existing approaches strive to overcome this limitation by reducing less…

cs.CV2024

OmniVid: A Generative Framework for Universal Video Understanding

Junke Wang, Dongdong Chen, Chong Luo +4

The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution.…

cs.CV2024

MouSi: Poly-Visual-Expert Vision-Language Models

Xiaoran Fan, Tao Ji, Changhao Jiang +21

Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens. These issu…

cs.AI20248 cited

Secrets of RLHF in Large Language Models Part II: Reward Modeling

Binghai Wang, Rui Zheng, Lu Chen +24

Reinforcement Learning from Human Feedback (RLHF) has become a crucial technology for aligning language models with human values and intentions, enabling models to produce more hel…

cs.CV20234 cited

Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection

Lingchen Meng, Xiyang Dai, Jianwei Yang +7

Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to expl…

cs.CV2023

Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data

Zuxuan Wu, Zejia Weng, Wujian Peng +4

Despite significant results achieved by Contrastive Language-Image Pretraining (CLIP) in zero-shot image recognition, limited effort has been made exploring its potential for zero-…