2 citations · 3 across the 4 of their papers we have counts for
5 papers
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
Rujiao Long, Yang Li, Xingyao Zhang +7
Exploration capacity shapes both inference-time performance and reinforcement learning (RL) training for large (vision-) language models, as stochastic sampling often yields redund…
Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
Zibin Liu, Cheng Zhang, Xi Zhao +6
Large Language Model (LLM) agents are increasingly deployed to automate complex workflows in mobile and desktop environments. However, current model-centric agent architectures str…
RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models
Tianqianjin Lin, Xi Zhao, Xingyao Zhang +5
Reinforcement learning (RL) can refine the reasoning abilities of large language models (LLMs), but critically depends on a key prerequisite: the LLM can already generate high-util…
MobiAgent: A Systematic Framework for Customizable Mobile Agents
Cheng Zhang, Erhu Feng, Xi Zhao +7
With the rapid advancement of Vision-Language Models (VLMs), GUI-based mobile agents have emerged as a key development direction for intelligent mobile systems. However, existing a…