2 papers
cs.CV2025
iMOVE: Instance-Motion-Aware Video Understanding
Jiaze Li, Yaya Shi, Zongyang Ma +7
Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understan…
cs.CV2025
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Jiankang Chen, Tianke Zhang, Changyi Liu +8
Multimodal visual language models are gaining prominence in open-world applications, driven by advancements in model architectures, training techniques, and high-quality data. Howe…