11 papers
MacTok: Robust Continuous Tokenization for Image Generation
Hengyu Zeng, Xin Gao, Guanghao Li +5
Continuous image tokenizers enable efficient visual generation, and those based on variational frameworks can learn smooth, structured latent representations through KL regularizat…
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
Daojie Peng, Fulong Ma, Jun Ma
Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of vis…
Annotation-Free Detection of Drivable Areas and Curbs Leveraging LiDAR Point Cloud Maps
Fulong Ma, Daojie Peng, Jun Ma
Drivable areas and curbs are critical traffic elements for autonomous driving, forming essential components of the vehicle visual perception system and ensuring driving safety. Dee…
Traffic Sign Recognition in Autonomous Driving: Dataset, Benchmark, and Field Experiment
Guoyang Zhao, Weiqing Qi, Kai Zhang +7
Traffic Sign Recognition (TSR) is a core perception capability for autonomous driving, where robustness to cross-region variation, long-tailed categories, and semantic ambiguity is…
MonoGlass3D: Monocular 3D Glass Detection with Plane Regression and Adaptive Feature Fusion
Kai Zhang, Guoyang Zhao, Jianxing Shi +3
Detecting and localizing glass in 3D environments poses significant challenges for visual perception systems, as the optical properties of glass often hinder conventional sensors f…
DuLoc: Life-Long Dual-Layer Localization in Changing and Dynamic Expansive Scenarios
Haoxuan Jiang, Peicong Qian, Yusen Xie +3
LiDAR-based localization serves as a critical component in autonomous systems, yet existing approaches face persistent challenges in balancing repeatability, accuracy, and environm…