12 papers
Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration
Jingbo Wang, Sendong Zhao, Haochun Wang +2
The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling diverse agents to integrate unique experti…
Continual Test-Time Adaptation for Object Detection with Adaptive Monitoring and Randomized Restoration
Shilei Cao, Juepeng Zheng, Yan Liu +5
Real-world application models are commonly deployed in dynamic environments, where the target domain distribution undergoes temporal changes. Continual Test-Time Adaptation (CTTA)…
Revisiting the Reliability of Language Models in Instruction-Following
Jianshuo Dong, Yutong Zhang, Yan Liu +4
Advanced LLMs have achieved near-ceiling instruction-following accuracy on benchmarks such as IFEval. However, these impressive scores do not necessarily translate to reliable serv…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Achieving Time Series Reasoning Requires Rethinking Model Design, Tasks Formulation, and Evaluation
Yaxuan Kong, Yiyuan Yang, Shiyu Wang +7
Understanding time series data is fundamental to many real-world applications. Recent work explores multimodal large language models (MLLMs) to enhance time series understanding wi…
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
Zehua Zhao, Zhixian Huang, Junren Li +28
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and mis…