most citedMCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

2 citations · 2 across the 8 of their papers we have counts for

collaborators

8 papers

cs.SE2025

LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering

Jielin Qiu, Zuxin Liu, Zhiwei Liu +18

As large language models (LLMs) evolve into sophisticated autonomous agents capable of complex software development tasks, evaluating their real-world capabilities becomes critical…

cs.LG2025

GeoGNN: Quantifying and Mitigating Semantic Drift in Text-Attributed Graphs

Liangwei Yang, Jing Ma, Jianguo Zhang +11

Graph neural networks (GNNs) on text--attributed graphs (TAGs) typically encode node texts using pretrained language models (PLMs) and propagate these embeddings through linear nei…

cs.LG2025

xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning

Cheng Qian, Zuxin Liu, Shirley Kokane +10

Modern LLM deployments confront a widening cost-performance spectrum: premium models deliver strong reasoning but are expensive, while lightweight models are economical yet brittle…

cs.LG2025

CoDA: Coding LM via Diffusion Adaptation

Haolin Chen, Shiyu Wang, Can Qin +12

Diffusion language models promise bidirectional context and infilling capabilities that autoregressive coders lack, yet practical systems remain heavyweight. We introduce CoDA, a 1…

cs.AI2025

UserRL: Training Interactive User-Centric Agent via Reinforcement Learning

Cheng Qian, Zuxin Liu, Akshara Prabhakar +10

Reinforcement learning (RL) has shown promise in training agentic models that move beyond static benchmarks to engage in dynamic, multi-turn interactions. Yet, the ultimate value o…

cs.SE2025

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Jielin Qiu, Zuxin Liu, Zhiwei Liu +14

The emergence of long-context language models with context windows extending to millions of tokens has created new opportunities for sophisticated code understanding and software d…