From the 1 of 11 linked papers with an AI index.
8 papers · 1 filter
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
Yanyu Chen, Yue Li, Yongyi Cui +2
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards…
GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection
Zheng Wu, Chengcheng Han, Zhengxi Lu +5
Despite the rapid progress of multimodal large language models in building Graphical User Interface (GUI) agents, their real-world task completion is fundamentally bottlenecked by…
A Principle-Driven Adaptive Policy for Group Cognitive Stimulation Dialogue for Elderly with Cognitive Impairment
Jiyue Jiang, Yanyu Chen, Pengan Chen +7
Cognitive impairment is becoming a major public health challenge. Cognitive Stimulation Therapy (CST) is an effective intervention for cognitive impairment, but traditional methods…
TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents
Yanyu Chen, Jiyue Jiang, Jiahong Liu +3
The evaluation of Deep Research Agents is a critical challenge, as conventional outcome-based metrics fail to capture the nuances of their complex reasoning. Current evaluation fac…
DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein Design
Yanting Li, Jiyue Jiang, Zikang Wang +11
Inverse Protein Folding (IPF) is a critical subtask in the field of protein design, aiming to engineer amino acid sequences capable of folding correctly into a specified three-dime…