11 papers
Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents
Bingzhen Liu, Xiaomeng Fan, Yuwei Wu +4
Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle…
Reliability-Prioritized Fine-Grained Generation in Multimodal Large
Xiaomeng Fan, Wei Wu, Yuwei Wu +9
Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generati…
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
Wei Wu, Xiaomeng Fan, Yuwei Wu +4
Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features fro…
Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
Mingliang Zhai, Hansheng Liang, Xiaomeng Fan +6
Embodied Question Answering (EQA) requires agents to explore 3D environments to obtain observations and answer questions related to the scene. Existing methods leverage VLMs to dir…
Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
Xiaomeng Fan, Yuchuan Mao, Zhi Gao +3
Open-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data. Existing methods estimate the distri…
Adaptive Model Ensemble for Continual Learning
Yuchuan Mao, Zhi Gao, Xiaomeng Fan +3
Model ensemble is an effective strategy in continual learning, which alleviates catastrophic forgetting by interpolating model parameters, achieving knowledge fusion learned from d…