4 papers
Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
Juncheng Wang, Zhe Hu, Chao Xu +5
Autoregressive (AR) models excel at generating temporally coherent audio by producing tokens sequentially, yet they often falter in faithfully following complex textual prompts, es…
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
Yijie Qian, Juncheng Wang, Yuxiang Feng +7
Current state-of-the-art paradigms predominantly treat Text-to-Motion (T2M) generation as a direct translation problem, mapping symbolic language directly to continuous poses. Whil…
An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
Chao Xu, Suyu Zhang, Yang Liu +11
Vision-Language-Action (VLA) models are driving a revolution in robotics, enabling machines to understand instructions and interact with the physical world. This field is exploding…
PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation
Zelin Zang, Fei Wang, Liangyu Li +4
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Recent UDA methods based on Vision Transformers (ViTs) h…