1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024
E-CAR: Efficient Continuous Autoregressive Image Generation via Multistage Modeling
Zhihang Yuan, Yuzhang Shang, Hanling Zhang +7
Recent advances in autoregressive (AR) models with continuous tokens for image generation show promising results by eliminating the need for discrete tokenization. However, these m…
cs.CV2024
Personalized Multimodal Large Language Models: A Survey
Junda Wu, Hanjia Lyu, Yu Xia +24
Multimodal Large Language Models (MLLMs) have become increasingly important due to their state-of-the-art performance and ability to integrate multiple data modalities, such as tex…
cs.CV2024★ 1 cited
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
Penghui Du, Yu Wang, Yifan Sun +7
Existing methods enhance open-vocabulary object detection by leveraging the robust open-vocabulary recognition capabilities of Vision-Language Models (VLMs), such as CLIP.However,…