2 citations · 4 across the 12 of their papers we have counts for
4 papers · 1 filter
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
Zhiyang Xu, Jiuhai Chen, Zhaojiang Lin +10
Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite thes…
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
Jingyuan Qi, Zhiyang Xu, Qifan Wang +1
We introduce Autoregressive Retrieval Augmentation (AR-RAG), a novel paradigm that enhances image generation by autoregressively incorporating knearest neighbor retrievals at the p…
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Haibo Wang, Zhiyang Xu, Yu Cheng +6
Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal groundin…
Multimodal Instruction Tuning with Conditional Mixture of LoRA
Ying Shen, Zhiyang Xu, Qifan Wang +3
Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in diverse tasks across different domains, with an increasing focus on improving their zero-shot g…