works on

From the 1 of 10 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models

Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai +2

While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods…

cs.AI2026

MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

Huanxi Liu, Kun Hu, Jiaqi Liao +6

The paper introduces MCPEvol-Bench, a benchmark that tests how well large language model agents adapt to changing tool interfaces and functionalities in Model Context Protocol (MCP…

cs.AI2026

Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities

Xu Zhang, Xudong Gong, Jiacheng Qin +5

Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation, limiting targeted model im…

cs.AI2025

Pay More Attention to the Robustness of Prompt for Instruction Data Mining

Qiang Wang, Dawei Feng, Xu Zhang +4

Instruction tuning has emerged as a paramount method for tailoring the behaviors of LLMs. Recent work has unveiled the potential for LLMs to achieve high performance through fine-t…

cs.AI2024

Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Yuanzhao Zhai, Tingkai Yang, Kele Xu +4

Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents st…