15 papers
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
Wenbo Wang, Fangyun Wei, QiXiu Li +5
Vision-language-action models have advanced robotic manipulation but remain constrained by reliance on the large, teleoperation-collected datasets dominated by the static, tabletop…
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
Sicheng Xu, Yu Deng, Shoukang Hu +5
Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a…
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
Dong Chen, Fangyun Wei, Ziyu Wan +18
We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters acro…
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
ZhiYuan Feng, Yu Deng, Ruichuan An +11
In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such re…
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
Zhiyuan Feng, Zhaolu Kang, Qijie Wang +16
Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent V…
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
Sicheng Xu, Guojun Chen, Jiaolong Yang +4
We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human…