collaborators

13 papers

cs.CV2026

Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation

Haozhe Wang, Jintao Cheng, Weibin Li +1

Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP's spati…

cs.RO2026

SIPTraj: Map-Free End-to-End Trajectory Prediction via Physics-Guided Scene Interaction

Feifei Liu, Zejun Wei, Haozhe Wang +6

Trajectory prediction of surrounding agents is a prerequisite for safe planning and decision making in autonomous driving. Without high-definition (HD) maps, sensor-derived bird's-…

cs.CV2026

OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding

Xiaoyu Tang, Jun Dong, Jintao Cheng +1

Remote sensing visual grounding (RSVG) aims to localize specific targets in remote sensing images using natural language expressions. However, existing methods are restricted to si…

cs.CV2026

VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models

Jintao Cheng, Haozhe Wang, Weibin Li +7

Vision-Language-Action (VLA) models have rapidly advanced embodied intelligence, enabling robots to execute complex, instruction-driven tasks. However, as model capacity and visual…

cs.SD2026

Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models

Hong Jia, Weibin Li, Jingyao Wu +6

Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classif…

cs.CV2025

Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather

Zhijian He, Feifei Liu, Yuwei Li +4

Multi-modal 3D object detection is important for reliable perception in robotics and autonomous driving. However, its effectiveness remains limited under adverse weather conditions…