collaborators

12 papers

cs.CV2026

Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation

Haozhe Wang, Jintao Cheng, Weibin Li +1

Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP's spati…

cs.CV2026

Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition

Jintao Cheng, Weibin Li, Zhijian He +3

Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either depend on extensive supervised tr…

cs.CV2026

VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models

Jintao Cheng, Haozhe Wang, Weibin Li +7

Vision-Language-Action (VLA) models have rapidly advanced embodied intelligence, enabling robots to execute complex, instruction-driven tasks. However, as model capacity and visual…

cs.CL2026

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.SD2026

Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models

Hong Jia, Weibin Li, Jingyao Wu +6

Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classif…

cs.AI2026

AutoHealth: An Uncertainty-Aware Multi-Agent System for Autonomous Health Data Modeling

Tong Xia, Weibin Li, Gang Liu +1

LLM-based agents have demonstrated strong potential for autonomous machine learning, yet their applicability to health data remains limited. Existing systems often struggle to gene…