5 papers · 1 filter
Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models
Ruchen Liu, Yi Yang, Yiming Xu +3
LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show tha…
HiMAP: History-aware Map-occupancy Prediction with Fallback
Yiming Xu, Yi Yang, Hao Cheng +1
Accurate motion forecasting is critical for autonomous driving, yet most predictors rely on multi-object tracking (MOT) with identity association, assuming that objects are correct…
Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
Yi Yang, Yiming Xu, Timo Kaiser +3
In this report, we present our solution to the MOT25-Spatiotemporal Action Grounding (MOT25-StAG) Challenge. The aim of this challenge is to accurately localize and track multiple…
Gap Completion in Point Cloud Scene occluded by Vehicles using SGC-Net
Yu Feng, Yiming Xu, Yan Xia +2
Recent advances in mobile mapping systems have greatly enhanced the efficiency and convenience of acquiring urban 3D data. These systems utilize LiDAR sensors mounted on vehicles t…
Controllable Diverse Sampling for Diffusion Based Motion Behavior Forecasting
Yiming Xu, Hao Cheng, Monika Sester
In autonomous driving tasks, trajectory prediction in complex traffic environments requires adherence to real-world context conditions and behavior multimodalities. Existing method…