9 papers
MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
Zhipeng Du, Duolikun Danier, Jan Eric Lenssen +1
In this paper, we focus on online zero-shot monocular 3D instance segmentation, a novel practical setting where existing approaches fail to perform because they rely on posed RGB-D…
Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier +4
The adoption of pre-trained visual representations (PVRs), leveraging features from large-scale vision models, has become a popular paradigm for training visuomotor policies. Howev…
SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language Segmentation
Chengxi Zeng, Yuxuan Jiang, Ge Gao +6
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ende…
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
Duolikun Danier, Ge Gao, Steven McDonagh +3
Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated vid…
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier +2
The integration of pre-trained visual representations (PVRs) has significantly advanced visuomotor policy learning. However, effectively leveraging these models remains a challenge…
GFix: Perceptually Enhanced Gaussian Splatting Video Compression
Siyue Teng, Ge Gao, Duolikun Danier +5
3D Gaussian Splatting (3DGS) enhances 3D scene reconstruction through explicit representation and fast rendering, demonstrating potential benefits for various low-level vision task…