5 papers
SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis
Yejun Zhang, Zihan Wang, Xu Ji +8
Generating photorealistic novel views from unposed images requires both 3D geometric understanding and the ability to synthesize unseen content. A natural strategy combines feed-fo…
PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos
Yihao Wang, Yang Miao, Wenshuai Zhao +8
Articulation perception aims to recover the motion and structure of articulated objects (e.g., drawers and cupboards), and is fundamental to 3D scene understanding in robotics, sim…
SceneExpander: Text-Guided 3D Scene Expansion via Free-Form View Insertion
Zijian He, Renjie Liu, Yihao Wang +5
World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workflows are inherently iterative: c…
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Hengtao Li, Pengxiang Ding, Runze Suo +8
Vision-Language-Action (VLA) models enable embodied decision-making but rely heavily on imitation learning, leading to compounding errors and poor robustness under distribution shi…
Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models
Wei Chen, Zhiyuan Li, Shuo Xin +1
This paper presents Dolphin, a novel decoder-decoder architecture for energy-efficient processing of long contexts in language models. Our approach addresses the significant energy…