From the 1 of 8 linked papers with an AI index.
8 papers · 1 filter
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
Dayong Liu, Chao Xu, Weihong Chen +5
The paper introduces CFG-Bench, a benchmark of videos and QA pairs to evaluate how well multimodal language models can generate fine-grained action instructions and higher-order re…
Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
Yijie Qian, Juncheng Wang, Chao Xu +6
As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality…
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
Yuxiang Feng, Juncheng Wang, Chao Xu +7
Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial design, but it remains challe…
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
Yijie Qian, Juncheng Wang, Yuxiang Feng +7
Current state-of-the-art paradigms predominantly treat Text-to-Motion (T2M) generation as a direct translation problem, mapping symbolic language directly to continuous poses. Whil…
Combo: Co-speech holistic 3D human motion generation and efficient customizable adaptation in harmony
Chao Xu, Mingze Sun, Zhi-Qi Cheng +5
In this paper, we propose a novel framework, Combo, for harmonious co-speech holistic 3D human motion generation and efficient customizable adaption. In particular, we identify tha…
TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
Jiazheng Xing, Chao Xu, Yijie Qian +5
Virtual try-on focuses on adjusting the given clothes to fit a specific person seamlessly while avoiding any distortion of the patterns and textures of the garment. However, the cl…