activity
20182026
most citedDon't Make Your LLM an Evaluation Benchmark Cheater

18 citations · 98 across the 75 of their papers we have counts for

collaborators
Showing 2025Show all

25 papers · 1 filter

cs.CV2025

PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents

Zikang Liu, Junyi Li, Wayne Xin Zhao +3

Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) promise human-like interaction with software applications, yet long-horizon tasks remain c…

cs.AI2025

Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design

Yihong Liu, Junyi Li, Wayne Xin Zhao +2

Combinatorial optimization problems are traditionally tackled with handcrafted heuristic algorithms, which demand extensive domain expertise and significant implementation effort.…

cs.AI2025

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework

Jie Chen, Jinhao Jiang, Yingqian Min +4

Large reasoning models (LRMs) have exhibited strong performance on complex reasoning tasks, with further gains achievable through increased computational budgets at inference. Howe…

cs.CL2025

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Jia Deng, Jie Chen, Zhipeng Chen +7

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs). Unlike traditiona…

cs.CV2025

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

Yifan Li, Kun Zhou, Wayne Xin Zhao +2

As scaling up training data has significantly improved the general multimodal capabilities of Large Vision-Language Models (LVLMs), they still suffer from the hallucination issue,…

cs.CL2025

Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Jia Deng, Jie Chen, Zhipeng Chen +2

Recently, reinforcement learning with verifiable rewards (RLVR) has been widely used for enhancing the reasoning abilities of large language models (LLMs). A core challenge in RLVR…