9 papers
DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection
Shiyi Mu, Zichong Gu, Zhiqi Ai +2
Stereo-based 3D obstacle perception for autonomous driving is currently constrained by an imbalanced triplet: deployment cost, detection accuracy, and open-set adaptability. While…
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
Yiling Gao, Hongchen Wei, Zhenzhong Chen
In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To a…
StereoDETR: Stereo-based Transformer for 3D Object Detection
Shiyi Mu, Zichong Gu, Zhiqi Ai +3
Compared to monocular 3D object detection, stereo-based 3D methods offer significantly higher accuracy but still suffer from high computational overhead and latency. The state-of-t…
Visual Bridge: Universal Visual Perception Representations Generating
Yilin Gao, Shuguang Dou, Junzhou Li +4
Recent advances in diffusion models have achieved remarkable success in isolated computer vision tasks such as text-to-image generation, depth estimation, and optical flow. However…
Knowledge Transfer from Interaction Learning
Yilin Gao, Kangyi Chen, Zhongxing Peng +2
Current visual foundation models (VFMs) face a fundamental limitation in transferring knowledge from vision language models (VLMs), while VLMs excel at modeling cross-modal interac…
Enhanced Fingerprint-based Positioning With Practical Imperfections: Deep learning-based approaches
Shugong Xu, Jun Jiang, Wenjun Yu +7
High-precision positioning is vital for cellular networks to support innovative applications such as extended reality, unmanned aerial vehicles (UAVs), and industrial Internet of T…