4 papers
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
Yi Yang, Zhihong Liu, Siqi Kou +9
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict te…
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
Yi Yang, Xueqi Li, Yiyang Chen +7
Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict…
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
Yushe Cao, Dianxi Shi, Xing Fu +5
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to ena…
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
Nannan Zhu, Yonghao Dong, Teng Wang +9
While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reason…