7 papers · 1 filter
DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views
Fuzhen Jiang, Changyue Shi, Chuxiao Yang +3
Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as embodied AI and autonomous…
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
Yufei Yin, Jie Zheng, Qianke Meng +7
Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by the poor quality of open-vocab…
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
Yufei Yin, Yuchen Xing, Qianke Meng +3
Understanding long videos requires extracting query-relevant information from long sequences under tight compute budgets. Existing text-then-LLM pipelines lose fine-grained visual…
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
Yufei Yin, Qianke Meng, Minghao Chen +3
Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on…
REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
Changyue Shi, Minghao Chen, Yiping Mao +4
Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often s…
SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images
Xinyuan Hu, Changyue Shi, Chuxiao Yang +6
Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, exi…