8 papers
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs
Jiawei Li, Ziyi Liu, Weijie Shi +3
3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning…
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
Pengna Li, Kangyi Wu, Shaoqing Xu +7
Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue…
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
Fucai Ke, Zhixi Cai, Boying Li +6
Multi-view visual reasoning is essential for intelligent systems that must understand complex environments from sparse and discrete viewpoints, yet existing research has largely fo…
FlowComposer: Composable Flows for Compositional Zero-Shot Learning
Zhenqi He, Lin Li, Long Chen
Compositional zero-shot learning (CZSL) aims to recognize unseen attribute-object compositions by recombining primitives learned from seen pairs. Recent CZSL methods built on visio…
Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
Lin Li, Chuhan Zhang, Dong Zhang +3
Open-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pr…
Towards Customized Knowledge Distillation for Chip-Level Dense Image Predictions
Dong Zhang, Pingcheng Dong, Long Chen +1
It has been revealed that efficient dense image prediction (EDIP) models designed for AI chips, trained using the knowledge distillation (KD) framework, encounter two key challenge…