1 citations · 1 across the 8 of their papers we have counts for
Showing 2026 · cs.CVShow all
2 papers · 2 filters
cs.CV2026
Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation
Boyuan Xiao, Bohong Chen, Yumeng Li +3
In embodied vision-language decision making tasks such as robotic manipulation and navigation, Vision-Language and Vision-Language-Action Models (VLMs & VLAs) are powerful tools wi…
cs.CV2026
EchoAvatar: Real-time Generative Avatar Animation from Audio Streams
Bohong Chen, Yumeng Li, Yinglin Xu +3
Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing ap…