3 papers
cs.CV2026
ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis
Yuan Wang, Yongchao Du, Mengting Chen +4
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However,…
cs.CV2026
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
Yuan Wang, Di Huang, Yaqi Zhang +5
Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Des…
cs.CV2025
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Yujia Liang, Jile Jiao, Xuetao Feng +3
Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…