4 citations · 7 across the 12 of their papers we have counts for
5 papers · 2 filters
PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning
Youngchae Chee, Hosu Lee, Sungjune Park +2
Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. Howeve…
Decoding Children's Gait Behavior
Yifan Shen, Boyi Li, Meihuan Huang +12
We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We specifically target the ambulator…
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
Junho Kim, Xu Cao, Houze Yang +6
Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with who…
Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos
Yujin Ham, Junho Kim, Vivek Boominathan +1
Egocentric "walking tour" videos provide a rich source of image data to develop rich and diverse visual models of environments around the world. However, the significant presence o…
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
Junho Kim, Hosu Lee, James M. Rehg +2
Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require…