Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Beyond Quantity: Trajectory Diversity Scaling for Code Agents
Guhong Chen, Chenghao Sun, Cheng Fu +16
As code large language models (LLMs) evolve into tool-interactive agents via the Model Context Protocol (MCP), their generalization is increasingly limited by low-quality synthetic…
cs.AI2026
Structuring Reasoning for Complex Rules Beyond Flat Representations
Zhihao Yang, Ancheng Xu, Jingpeng Li +11
Large language models (LLMs) face significant challenges when processing complex rule systems, as they typically treat interdependent rules as unstructured textual data rather than…
cs.AI2025
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
Jiahao Zhao, Luxin Xu, Minghuan Tan +4
Numerous medical systems powered by Large Language Models (LLMs) have achieved remarkable progress in diverse healthcare tasks. However, research on their medication safety remains…