10 papers
Conditional Multi-Event Temporal Grounding in Long-Form Video
Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez +12
Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositi…
Unlocking Multi-Site Clinical Data: A Federated Approach to Privacy-First Child Autism Behavior Analysis
Guangyu Sun, Wenhan Wu, Zhishuai Guo +3
Automated recognition of autistic behaviors in children is essential for early intervention and objective clinical assessment. However, the development of robust models is severely…
Geo: Geometry-Guided Cross-view Geo-Localization and Image Synthesis
Yancheng Zhang, Xiaohan Zhang, Guangyu Sun +3
Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geomet…
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
Guangyu Sun, Archit Singhal, Burak Uzkent +3
Video Large Language Models (VLMs) have achieved strong performance on various vision-language tasks, yet their practical use is limited by the massive number of visual tokens prod…
EGGS: Exchangeable 2D/3D Gaussian Splatting for Geometry-Appearance Balanced Novel View Synthesis
Yancheng Zhang, Guangyu Sun, Chen Chen
Novel view synthesis (NVS) is crucial in computer vision and graphics, with wide applications in AR, VR, and autonomous driving. While 3D Gaussian Splatting (3DGS) enables real-tim…
Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)
Jing Bi, Susan Liang, Xiaofei Zhou +16
Reasoning is central to human intelligence, enabling structured problem-solving across diverse tasks. Recent advances in large language models (LLMs) have greatly enhanced their re…