activity
20242026
collaborators

15 papers

cs.RO2026

MobileManiBench: Simplifying Model Verification for Mobile Manipulation

Wenbo Wang, Fangyun Wei, QiXiu Li +5

Vision-language-action models have advanced robotic manipulation but remain constrained by reliance on the large, teleoperation-collected datasets dominated by the static, tabletop…

cs.CV2026

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

Sicheng Xu, Yu Deng, Shoukang Hu +5

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a…

cs.CV2026

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Dong Chen, Fangyun Wei, Ziyu Wan +18

We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters acro…

cs.AI2026

TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning

ZhiYuan Feng, Yu Deng, Ruichuan An +11

In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such re…

cs.CV2026

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

Zhiyuan Feng, Zhaolu Kang, Qijie Wang +16

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent V…

cs.CV2025

VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image

Sicheng Xu, Guojun Chen, Jiaolong Yang +4

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human…