agent adaptability 1dynamic tool evolution 1LLM agents 1model context protocol 1tool-use benchmarking 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
Huanxi Liu, Kun Hu, Jiaqi Liao +6
The paper introduces MCPEvol-Bench, a benchmark that tests how well large language model agents adapt to changing tool interfaces and functionalities in Model Context Protocol (MCP…
cs.CV2026
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
Cailin Zhuang, Ailin Huang, Yaoqi Hu +12
Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, exi…
cs.CV2025
Step1X-Edit: A Practical Framework for General Image Editing
Shiyu Liu, Yucheng Han, Peng Xing +21
In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has i…