1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.AI2026
Beyond Quantity: Trajectory Diversity Scaling for Code Agents
Guhong Chen, Chenghao Sun, Cheng Fu +16
As code large language models (LLMs) evolve into tool-interactive agents via the Model Context Protocol (MCP), their generalization is increasingly limited by low-quality synthetic…
cs.CL2024
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
Xinghua Zhang, Haiyang Yu, Cheng Fu +2
In the realm of large language models (LLMs), the ability of models to accurately follow instructions is paramount as more agents and applications leverage LLMs for construction, w…
cs.HC2024★ 1 cited
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
Chongyan Sun, Ken Lin, Shiwei Wang +3
This paper introduces LalaEval, a holistic framework designed for the human evaluation of domain-specific large language models (LLMs). LalaEval proposes a comprehensive suite of e…