Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
Run Luo, Yunshui Li, Longze Chen +9
The development of large language models (LLMs) has significantly advanced the emergence of large multimodal models (LMMs). While LMMs have achieved tremendous success by promoting…
cs.CV2024
IP-MOT: Instance Prompt Learning for Cross-Domain Multi-Object Tracking
Run Luo, Zikai Song, Longze Chen +3
Multi-Object Tracking (MOT) aims to associate multiple objects across video frames and is a challenging vision task due to inherent complexities in the tracking environment. Most e…