activity
20242026
collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

MacTok: Robust Continuous Tokenization for Image Generation

Hengyu Zeng, Xin Gao, Guanghao Li +5

Continuous image tokenizers enable efficient visual generation, and those based on variational frameworks can learn smooth, structured latent representations through KL regularizat…

cs.CV2026

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

Daojie Peng, Fulong Ma, Jun Ma

Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of vis…

cs.CV2026

Annotation-Free Detection of Drivable Areas and Curbs Leveraging LiDAR Point Cloud Maps

Fulong Ma, Daojie Peng, Jun Ma

Drivable areas and curbs are critical traffic elements for autonomous driving, forming essential components of the vehicle visual perception system and ensuring driving safety. Dee…

cs.CV2026

Traffic Sign Recognition in Autonomous Driving: Dataset, Benchmark, and Field Experiment

Guoyang Zhao, Weiqing Qi, Kai Zhang +7

Traffic Sign Recognition (TSR) is a core perception capability for autonomous driving, where robustness to cross-region variation, long-tailed categories, and semantic ambiguity is…

cs.CV2025

TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition

Guoyang Zhao, Fulong Ma, Weiqing Qi +4

Traffic sign is a critical map feature for navigation and traffic control. Nevertheless, current methods for traffic sign recognition rely on traditional deep learning models, whic…

cs.CV2025

FisheyeDepth: A Real Scale Self-Supervised Depth Estimation Model for Fisheye Camera

Guoyang Zhao, Yuxuan Liu, Weiqing Qi +3

Accurate depth estimation is crucial for 3D scene comprehension in robotics and autonomous vehicles. Fisheye cameras, known for their wide field of view, have inherent geometric be…