4 citations · 5 across the 2 of their papers we have counts for
3 papers
cs.CV2023
MChat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
Xiaowei Chi, Junbo Qi, Rongyu Zhang +3
While current LLM chatbots like GPT-4V bridge the gap between human instructions and visual representations to enable text-image generations, they still lack efficient alignment me…
cs.CV2023★ 1 cited
Improving Compositional Text-to-image Generation with Large Vision-Language Models
Song Wen, Guian Fang, Renrui Zhang +3
Recent advancements in text-to-image models, particularly diffusion models, have shown significant promise. However, compositional text-to-image models frequently encounter difficu…
cs.RO2023★ 4 cited
Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill
Wenzhe Cai, Siyuan Huang, Guangran Cheng +4
Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first…