works on

From the 1 of 28 linked papers with an AI index.

collaborators

28 papers

cs.LG2026

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Bingzhen Liu, Xiaomeng Fan, Yuwei Wu +4

Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle…

cs.AI2026

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Chenrui Shi, Yuwei Wu, Yang Liu +5

The paper introduces an Interactive Reward Agent that evaluates GUI task completion by proposing conditions and verifying them using system, application, and GUI tools, and demonst…

cs.RO2026

A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models

Haoyu Zhang, Yuwei Wu, Jin Chen +6

Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. Ho…

cs.CV2026

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

Xiaomeng Fan, Wei Wu, Yuwei Wu +9

Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generati…

cs.CV2026

Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification

Haoyu Zhang, Yuwei Wu, Pengxiang Li +6

In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs,…

cs.CV2026

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning

Pengxiang Li, Zhi Gao, Bofei Zhang +8

Multimodal agents, which integrate a controller e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks.…