9 citations · 58 across the 19 of their papers we have counts for
19 papers
Learning to Rank Patches for Unbiased Image Redundancy Reduction
Yang Luo, Zhineng Chen, Peng Zhou +3
Images suffer from heavy spatial redundancy because pixels in neighboring regions are spatially correlated. Existing approaches strive to overcome this limitation by reducing less…
OmniVid: A Generative Framework for Universal Video Understanding
Junke Wang, Dongdong Chen, Chong Luo +4
The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution.…
MouSi: Poly-Visual-Expert Vision-Language Models
Xiaoran Fan, Tao Ji, Changhao Jiang +21
Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens. These issu…
Secrets of RLHF in Large Language Models Part II: Reward Modeling
Binghai Wang, Rui Zheng, Lu Chen +24
Reinforcement Learning from Human Feedback (RLHF) has become a crucial technology for aligning language models with human values and intentions, enabling models to produce more hel…
Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection
Lingchen Meng, Xiyang Dai, Jianwei Yang +7
Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to expl…
Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data
Zuxuan Wu, Zejia Weng, Wujian Peng +4
Despite significant results achieved by Contrastive Language-Image Pretraining (CLIP) in zero-shot image recognition, limited effort has been made exploring its potential for zero-…