19 papers
Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents
Bingzhen Liu, Xiaomeng Fan, Yuwei Wu +4
Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle…
Reliability-Prioritized Fine-Grained Generation in Multimodal Large
Xiaomeng Fan, Wei Wu, Yuwei Wu +9
Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generati…
Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification
Haoyu Zhang, Yuwei Wu, Pengxiang Li +6
In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs,…
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
Wei Wu, Xiaomeng Fan, Yuwei Wu +4
Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features fro…
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
Xintong Zhang, Xiaowen Zhang, Jingrong Wu +8
Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text…
Facial Expression Generation Aligned with Human Preference for Natural Dyadic Interaction
Xu Chen, Rui Gao, Xinjie Zhang +5
Achieving natural dyadic interaction requires generating facial expressions that are emotionally appropriate and socially aligned with human preference. Human feedback offers a com…