From the 1 of 5 linked papers with an AI index.
5 papers
Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics
Yue Chang, Rufeng Chen, Yifan Tian +7
The paper introduces JITOMA, a framework that builds 3D scene graphs on demand by using task-driven heatmaps and large language models to activate only relevant parts of the graph,…
LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation
Rufeng Chen, Yue Chang, Zili Shao +5
Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus…
PSG-Nav: Probabilistic Scene Graph Navigation via Multiverse Decision Making
Rufeng Chen, Yue Chang, Xiaqiang Tang +2
Open-vocabulary navigation requires embodied agents to manage significant perception uncertainty stemming from semantic ambiguity and model errors. However, most existing works set…
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
Yue Chang, Rufeng Chen, Zhaofan Zhang +3
Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suff…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…