3 papers
cs.CV2026
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
Yichen Peng, Jyun-Ting Song, Chen-Chieh Liao +3
Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset c…
cs.CV2026
DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
Yichen Peng, Jyun-Ting Song, Siyeol Jung +7
Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a sing…
cs.RO2026
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
Jiahui Fu, Junyu Nan, Lingfeng Sun +5
Solving long-horizon tasks requires robots to integrate high-level semantic reasoning with low-level physical interaction. While vision-language models (VLMs) and video generation…