1 citations · 1 across the 2 of their papers we have counts for
5 papers
Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation
Zilong Xie, Jingyu Gong, Xin Tan +2
Existing end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal L…
Learning to Rank Critical Road Segments via Heterogeneous Graphs with Origin-Destination Flow Integration
Ming Xu, Jinrong Xiang, Zilong Xie +1
Existing learning-to-rank methods for road networks often fail to incorporate origin-destination (OD) flows and route information, limiting their ability to model long-range spatia…
E2E-AFG: An End-to-End Model with Adaptive Filtering for Retrieval-Augmented Generation
Yun Jiang, Zilong Xie, Wei Zhang +2
Retrieval-augmented generation methods often neglect the quality of content retrieved from external knowledge bases, resulting in irrelevant information or potential misinformation…
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
Wanpeng Zhang, Zilong Xie, Yicheng Feng +4
Multimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We…
Hierarchical Spatial Proximity Reasoning for Vision-and-Language Navigation
Ming Xu, Zilong Xie
Most Vision-and-Language Navigation (VLN) algorithms are prone to making inaccurate decisions due to their lack of visual common sense and limited reasoning capabilities. To addres…