12 papers
Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation
Haozhe Wang, Jintao Cheng, Weibin Li +1
Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP's spati…
Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition
Jintao Cheng, Weibin Li, Zhijian He +3
Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either depend on extensive supervised tr…
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
Jintao Cheng, Haozhe Wang, Weibin Li +7
Vision-Language-Action (VLA) models have rapidly advanced embodied intelligence, enabling robots to execute complex, instruction-driven tasks. However, as model capacity and visual…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
Hong Jia, Weibin Li, Jingyao Wu +6
Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classif…
AutoHealth: An Uncertainty-Aware Multi-Agent System for Autonomous Health Data Modeling
Tong Xia, Weibin Li, Gang Liu +1
LLM-based agents have demonstrated strong potential for autonomous machine learning, yet their applicability to health data remains limited. Existing systems often struggle to gene…