collaborators

7 papers

cs.CL2026

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model lea…

cs.LG2025

Xmodel-2.5: 1.3B Data-Efficient Reasoning SLM

Yang Liu, Xiaolong Zhong, Ling Jiang

Large language models deliver strong reasoning and tool-use skills, yet their computational demands make them impractical for edge or cost-sensitive deployments. We present \textbf…

cs.CL2025

ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?

Haoxin Wang, Xianhan Peng, Xucheng Huang +5

In this paper, we introduce ECom-Bench, the first benchmark framework for evaluating LLM agent with multimodal capabilities in the e-commerce customer support domain. ECom-Bench fe…

cs.CL2025

MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service

Yizhe Huang, Yang Liu, Ruiyu Zhao +3

Large Language Model-based agents(LLM-based agents) are increasingly deployed in customer service, yet they often forget across sessions, repeat errors, and lack mechanisms for con…

cs.CL2025

Survey of Specialized Large Language Model

Chenghan Yang, Ruiyu Zhao, Yang Liu +1

The rapid evolution of specialized large language models (LLMs) has transitioned from simple domain adaptation to sophisticated native architectures, marking a paradigm shift in AI…

cs.CL2025

MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents

Ming Gong, Xucheng Huang, Chenghan Yang +4

Recent advances in large language models (LLMs) have enabled new applications in e-commerce customer service. However, their capabilities remain constrained in complex, multimodal…