4 papers
Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation
Boyuan Xiao, Bohong Chen, Yumeng Li +3
In embodied vision-language decision making tasks such as robotic manipulation and navigation, Vision-Language and Vision-Language-Action Models (VLMs & VLAs) are powerful tools wi…
DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model
Bohong Chen, Haiyang Liu
Generating realistic, dyadic talking head video requires ultra-low latency. Existing chunk-based methods require full non-causal context windows, introducing significant delays. Th…
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
Bohong Chen, Yumeng Li, Youyi Zheng +2
The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categ…
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
Bohong Chen, Yumeng Li, Yao-Xiang Ding +2
Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic fu…