From the 1 of 6 linked papers with an AI index.
6 papers
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
Huanxi Liu, Kun Hu, Jiaqi Liao +6
The paper introduces MCPEvol-Bench, a benchmark that tests how well large language model agents adapt to changing tool interfaces and functionalities in Model Context Protocol (MCP…
ParetoPilot: Zero-Surrogate Offline Multi-Objective Optimization via Infer-Perturb-Guide Diffusion
Ruiqing Sun, Sen Yang, Dawei Feng +3
Offline multi-objective optimization (Offline MOO) seeks Pareto-optimal designs from static datasets without additional environment interactions. Existing generative methods typica…
MAny: Merge Anything for Multimodal Continual Instruction Tuning
Zijian Gao, Wangwang Jia, Xingxing Zhang +6
Multimodal Continual Instruction Tuning (MCIT) is essential for sequential task adaptation of Multimodal Large Language Models (MLLMs) but is severely restricted by catastrophic fo…
Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities
Xu Zhang, Xudong Gong, Jiacheng Qin +5
Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation, limiting targeted model im…
Diffusion-based Evolutionary Optimization for 3D Multi-Objective Molecular Generation
Ruiqing Sun, Dawei Feng, Sen Yang +5
Optimizing conflicting molecular properties while strictly adhering to complex 3D structural constraints constitutes a challenging Constrained Multi-Objective Optimization Problem…
Pay More Attention to the Robustness of Prompt for Instruction Data Mining
Qiang Wang, Dawei Feng, Xu Zhang +4
Instruction tuning has emerged as a paramount method for tailoring the behaviors of LLMs. Recent work has unveiled the potential for LLMs to achieve high performance through fine-t…