4 papers · 1 filter
ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis
Hao Wang, Jindong Han, Wei Fan +1
Climate research is pivotal for mitigating global environmental crises, yet the accelerating volume of multi-scale datasets and the complexity of analytical tools have created sign…
ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices
Dezhi Kong, Zhengzhao Feng, Qiliang Liang +12
Multimodal large language models (MLLMs) have made significant progress in mobile agent development, yet their capabilities are predominantly confined to a reactive paradigm, where…
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
Shu Zhou, Rui Ling, Junan Chen +3
Scaling test-time compute through extended chains of thought has become a dominant paradigm for improving large language model reasoning. However, existing research implicitly assu…
Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration
Hang Lv, Hongchao Gu, Ruiqing Yang +5
Generative listwise reranking leverages global context for superior retrieval but is plagued by intrinsic position bias, where models exhibit structural sensitivity to input order…