2 citations · 2 across the 8 of their papers we have counts for
8 papers
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
Hong Jia, Weibin Li, Jingyao Wu +6
Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classif…
AutoHealth: An Uncertainty-Aware Multi-Agent System for Autonomous Health Data Modeling
Tong Xia, Weibin Li, Gang Liu +1
LLM-based agents have demonstrated strong potential for autonomous machine learning, yet their applicability to health data remains limited. Existing systems often struggle to gene…
One Step Is Enough: Dispersive MeanFlow Policy Optimization
Guowei Zou, Haitao Wang, Hejun Wu +3
Real-time robotic control demands fast action generation. However, existing generative policies based on diffusion and flow matching require multi-step sampling, fundamentally limi…
DM1: MeanFlow with Dispersive Regularization for 1-Step Robotic Manipulation
Guowei Zou, Haitao Wang, Hejun Wu +3
The ability to learn multi-modal action distributions is indispensable for robotic manipulation policies to perform precise and robust control. Flow-based generative models have re…
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
Jintao Cheng, Weibin Li, Jiehao Luo +5
Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundat…