28 citations · 67 across the 8 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 14 cited
Large Language Models are Visual Reasoning Coordinators
Liangyu Chen, Bo Li, Sheng Shen +5
Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsen…
cs.CV2023★ 10 cited
LAMP: Learn A Motion Pattern for Few-Shot-Based Video Generation
Ruiqi Wu, Liangyu Chen, Tong Yang +3
With the impressive progress in diffusion-based text-to-image generation, extending such powerful generative ability to text-to-video raises enormous attention. Existing methods ei…
cs.CV2023★ 28 cited
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Bo Li, Yuanhan Zhang, Liangyu Chen +5
High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language…