activity
20232025
most citedMCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

3 citations · 11 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models

Zhiyuan Hu, Yibo Wang, Hanze Dong +5

Large reasoning models (LRMs) already possess a latent capacity for long chain-of-thought reasoning. Prior work has shown that outcome-based reinforcement learning (RL) can inciden…

cs.CL2025

BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation

Bo Pang, Hanze Dong, Jiacheng Xu +3

Large language models (LLMs), such as o1 from OpenAI, have demonstrated remarkable reasoning capabilities. o1 generates a long chain-of-thought (LongCoT) before answering a questio…

cs.CL2025

Reward-Guided Speculative Decoding for Efficient LLM Reasoning

Baohao Liao, Yuhui Xu, Hanze Dong +5

We introduce Reward-Guided Speculative Decoding (RSD), a novel framework aimed at improving the efficiency of inference in large language models (LLMs). RSD synergistically combine…

cs.CL2024

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Yiheng Xu, Zekun Wang, Junli Wang +6

Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a uni…

cs.CL2024

XForecast: Evaluating Natural Language Explanations for Time Series Forecasting

Taha Aksu, Chenghao Liu, Amrita Saha +3

Time series forecasting aids decision-making, especially for stakeholders who rely on accurate predictions, making it very important to understand and explain these models to ensur…

cs.CL2024

MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs

Lei Wang, Shan Dong, Yuhui Xu +6

Recent large language models (LLMs) have demonstrated versatile capabilities in long-context scenarios. Although some recent benchmarks have been developed to evaluate the long-con…