4 papers · 1 filter
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi +3
Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologie…
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Noriyuki Kugo, Xiang Li, Zixin Li +9
Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. Howe…
AdaVid: Adaptive Video-Language Pretraining
Chaitanya Patel, Juan Carlos Niebles, Ehsan Adeli
Contrastive video-language pretraining has demonstrated great success in learning rich and robust video representations. However, deploying such video encoders on compute-constrain…
SocialGen: Modeling Multi-Human Social Interaction with Language Models
Heng Yu, Juze Zhang, Changan Chen +4
Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental…