2 papers
cs.CV2025
EventVL: Understand Event Streams via Multimodal Large Language Model
Pengteng Li, Yunfan Lu, Pinghao Song +3
The event-based Vision-Language Model (VLM) recently has made good progress for practical vision tasks. However, most of these works just utilize CLIP for focusing on traditional p…
cs.CV2025
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Rongchang Xie, Chen Du, Ping Song +1
We introduce MUSE-VL, a Unified Vision-Language Model through Semantic discrete Encoding for multimodal understanding and generation. Recently, the research community has begun exp…