4 papers
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
Qilang Ye, Yu Zhou, Lian He +10
Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action…
DiTraj: training-free trajectory control for video diffusion transformer
Cheng Lei, Jiayu Zhang, Yue Ma +6
Diffusion Transformers (DiT)-based video generation models with 3D full attention exhibit strong generative capabilities. Trajectory control represents a user-friendly task in the…
Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
Georgios Tziafas, Jiayun Zhang, Hamidreza Kasaei
Learning visuomotor policies from expert demonstrations is an important frontier in modern robotics research, however, most popular methods require copious efforts for collecting t…
SVC 2025: the First Multimodal Deception Detection Challenge
Xun Lin, Xiaobao Guo, Taorui Wang +5
Deception detection is a critical task in real-world applications such as security screening, fraud prevention, and credibility assessment. While deep learning methods have shown p…