4 papers · 1 filter
GKDT: General Keypoint Detection Transformer
Changsheng Lu, Yuxin Chen, Haokun Gui +5
With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful…
Semantic Generative Tuning for Unified Multimodal Models
Songsong Yu, Yuxin Chen, Ying Shan +1
Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently…
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
Junfu Pu, Yuxin Chen, Teng Wang +1
Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detail…
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
Shengchao Zhou, Yuxin Chen, Yuying Ge +4
Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relation…