11 papers
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy
Yicheng Xu, Jiangning Zhang, Zhucun Xue +5
In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to example selection and formatting. In u…
ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image Models
Dong Han, Yong Li
With the advance of generative AI, the text-to-image (T2I) model has the ability to generate various contents. However, T2I models still can generate unsafe contents. To alleviate…
HDINO: A Concise and Efficient Open-Vocabulary Detector
Hao Zhang, Yiqun Wang, Qinran Lin +2
Despite the growing interest in open-vocabulary object detection in recent years, most existing methods rely heavily on manually curated fine-grained training datasets as well as r…
Securing the Floor and Raising the Ceiling: A Merging-based Paradigm for Multi-modal Search Agents
Zhixiang Wang, Jingxuan Xu, Dajun Chen +3
Recent advances in Vision-Language Models (VLMs) have motivated the development of multi-modal search agents that can actively invoke external search tools and integrate retrieved…
Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
Kaiyuan Li, Xiaoyue Chen, Chen Gao +2
Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image toke…
AirScape: An Aerial Generative World Model with Motion Controllability
Baining Zhao, Rongze Tang, Mingyuan Jia +9
How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general s…