7 papers
FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving
Guolei Huang, Tengfei She, Yuxuan Lu +3
Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate…
VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction
Junhong Lin, Xianda Guo, Kangli Wang +4
Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides observations with ambiguous depth…
: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-Tuning
Yuqi Ye, Zijian Zhang, Junhong Lin +3
Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However…
AnyPcc: Compressing Any Point Cloud with a Single Universal Model
Kangli Wang, Qianxi Yi, Yuqi Ye +2
Generalization remains a critical challenge in deep learning-based point cloud geometry compression. While existing methods perform well on standard benchmarks, their performance c…
Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood Intervals
Changhao Peng, Yuqi Ye, Wei Gao
Gaussian and Laplacian entropy models are proved effective in learned point cloud attribute compression, as they assist in arithmetic coding of latents. However, we demonstrate thr…
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
Tao Feng, Zhiyuan Zhao, Yifan Xie +4
We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that r…