1 citations · 1 across the 8 of their papers we have counts for
7 papers · 1 filter
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
Jiwook Han, Geo Ahn, Youngrae Kim +1
Multimodal Large Language Models (MLLMs) have shown strong performance on Video Temporal Grounding (VTG). However, their coarse recognition capabilities are insufficient for fine-g…
Do Vision Language Models Understand Human Engagement in Games?
Ziyi Wang, Qizan Guo, Rishitosh Singh +1
Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vision--language models (VLMs) can infer…
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
Ziyi Wang, Xinshun Wang, Shuang Chen +2
We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single a…
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
Ziyi Wang, Peiming Li, Xinshun Wang +3
Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skeletons. Existing methods either c…
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
Xinshun Wang, Peiming Li, Ziyi Wang +5
Human motion analysis tasks, such as temporal 3D pose estimation, motion prediction, and motion in-betweening, play an essential role in computer vision. However, current paradigms…
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
Peiming Li, Ziyi Wang, Yulin Yuan +4
Point cloud videos capture dynamic 3D motion while reducing the effects of lighting and viewpoint variations, making them highly effective for recognizing subtle and continuous hum…