10 papers
SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning
Yizhi Li, Jiawei Jiang, Guanhong Wang +2
Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-…
Hand3R: Online 4D Hand-Scene Reconstruction in the Wild
Wendi Hu, Haonan Zhou, Wenhao Hu +1
For Embodied AI, jointly reconstructing dynamic hands and the dense scene context is crucial for understanding physical interaction. However, most existing methods recover isolated…
RecurGS: Interactive Scene Modeling via Discrete-State Recurrent Gaussian Fusion
Wenhao Hu, Haonan Zhou, Zesheng Li +4
Recent advances in 3D scene representations have enabled high-fidelity novel view synthesis, yet adapting to discrete scene changes and constructing interactive 3D environments rem…
IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion
Wenhao Hu, Zesheng Li, Haonan Zhou +5
Reconstructing complete and interactive 3D scenes remains a fundamental challenge in computer vision and robotics, particularly due to persistent object occlusions and limited sens…
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai +3
The challenge of long video understanding lies in its high computational complexity and prohibitive memory cost, since the memory and computation required by transformer-based LLMs…
DSG-World: Learning a 3D Gaussian World Model from Dual State Videos
Wenhao Hu, Xuexiang Wen, Xi Li +1
Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines ar…