5 papers · 1 filter
SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments
Yundaichuan Zhan, Minghe Gao, Zhongqi Yue +7
Recent works have explored integrating Vision-Language Models (VLMs) with classical planners that rely on symbolic representations of planning problems to generate long-horizon pla…
Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
Weile Chen, Bingchen Miao, Qifan Yu +6
Recent advances in Multimodal Large Language Models (MLLMs) have led to promising progress in web agents. However, existing web agents often rely on handcrafted execution pipelines…
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
Zhaoyu Fan, Kaihang Pan, Mingze Zhou +7
Knowledge editing enables multimodal large language models (MLLMs) to efficiently update outdated or incorrect information. However, existing benchmarks primarily emphasize cogniti…
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
Jiang Liu, Bolin Li, Haoyuan Li +13
Efficient multimodal large language models (EMLLMs), in contrast to multimodal large language models (MLLMs), reduce model size and computational costs and are often deployed on re…
AlignLLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
Hongzhe Huang, Jiang Liu, Zhewen Yu +8
Recent advances in Multi-modal Large Language Models (MLLMs), such as LLaVA-series models, are driven by massive machine-generated instruction-following data tuning. Such automatic…