6 papers
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation
Zihan Su, Hongyang Wei, Kangrui Cen +4
Unified Multimodal Models (UMMs) integrate both visual understanding and generation within a single framework. Their ultimate aspiration is to create a cycle where understanding an…
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
Guanhua Chen, Chuyue Huang, Yutong Yao +4
Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fine-grained user queries and…
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
Guanhua Chen, Yutong Yao, Shenghe Sun +5
Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA)…
Learning Agentic Policy from Action Guidance
Yuxiang Ji, Zengbin Wang, Yong Wang +6
Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training signals emerge only within its…
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
Zeguan Xiao, Xuanzhe Xu, Yun Chen +4
Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns.…
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
Yuxiang Ji, Yong Wang, Ziyu Ma +6
The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches le…