2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2025
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
Yue Zhao, Fuzhao Xue, Scott Reed +6
We introduce Quantized Language-Image Pretraining (QLIP), a visual tokenization method that combines state-of-the-art reconstruction quality with state-of-the-art zero-shot image u…
cs.CV2022★ 2 cited
Learning Video Representations from Large Language Models
Yue Zhao, Ishan Misra, Philipp Krähenbühl +1
We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual…
cs.CV2022★ 1 cited
Real-time Online Video Detection with Temporal Smoothing Transformers
Yue Zhao, Philipp Krähenbühl
Streaming video recognition reasons about objects and their actions in every frame of a video. A good streaming recognition model captures both long-term dynamics and short-term ch…