11 papers
Gaga: Group Any Gaussians via 3D-aware Memory Bank
Weijie Lyu, Xueting Li, Abhijit Kundu +2
We introduce Gaga, a framework that reconstructs and segments open-world 3D scenes by leveraging inconsistent 2D masks predicted by zero-shot class-agnostic segmentation models. Co…
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5
The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies…
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei +3
Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often…
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
Deshui Miao, Xingsen Huang, Yameng Gu +3
Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models…
ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation
Hanxiao Sun, Mingxin Yang, Shuhui Yang +5
Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unseen regions. Multi-view condi…
TravExplorer: Cross-Floor Embodied Exploration via Traversability-Aware 3-D Planning
Han Zheng, Zhe Chen, Yudong Huang +4
Zero-shot Object Navigation (ZSON) has shown promise for open-vocabulary target search in unseen environments, yet most existing systems remain tied to planar representations and s…