11 papers
Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation
Haozhe Wang, Jintao Cheng, Weibin Li +1
Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP's spati…
SIPTraj: Map-Free End-to-End Trajectory Prediction via Physics-Guided Scene Interaction
Feifei Liu, Zejun Wei, Haozhe Wang +6
Trajectory prediction of surrounding agents is a prerequisite for safe planning and decision making in autonomous driving. Without high-definition (HD) maps, sensor-derived bird's-…
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
Yu Han, Zhiwei Huang, Yanting Zhang +5
Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in t…
OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding
Xiaoyu Tang, Jun Dong, Jintao Cheng +1
Remote sensing visual grounding (RSVG) aims to localize specific targets in remote sensing images using natural language expressions. However, existing methods are restricted to si…
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
Jintao Cheng, Haozhe Wang, Weibin Li +7
Vision-Language-Action (VLA) models have rapidly advanced embodied intelligence, enabling robots to execute complex, instruction-driven tasks. However, as model capacity and visual…
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
Hong Jia, Weibin Li, Jingyao Wu +6
Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classif…