activity
20242026
collaborators

5 papers

cs.CV2026

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

Yuan Wang, Yongchao Du, Mengting Chen +4

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However,…

cs.CV2025

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

Yujia Liang, Jile Jiao, Xuetao Feng +3

Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…

cs.CV2024

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Yuan Wang, Di Huang, Yaqi Zhang +5

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Des…

cs.CV2024

Holistic-Motion2D: Scalable Whole-body Human Motion Generation in 2D Space

Yuan Wang, Zhao Wang, Junhao Gong +8

In this paper, we introduce a novel path to human motion generation by focusing on 2D space. Traditional methods have primarily generated human motions in 3D, wh…

cs.CV2024

Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning

Wenjin Hou, Shiming Chen, Shuhuang Chen +6

Generative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes, which is an effective way to advance ZSL. However, existing generative metho…