6 papers
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
Wenfang Sun, Hao Chen, Yingjun Du +2
Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to…
GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
Jie Ou, Shuaihong Jiang, Yingjun Du +1
Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, DoRA, and HiRA, enable lightweight adaptation of large pre-trained models via low-rank updates. However, existing PEFT…
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
Wenfang Sun, Yingjun Du, Gaowen Liu +2
We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which lea…
CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation
Jie Liu, Pan Zhou, Yingjun Du +4
In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods ofte…
Prompt Diffusion Robustifies Any-Modality Prompt Learning
Yingjun Du, Gaowen Liu, Yuzhang Shang +3
Foundation models enable prompt-based classifiers for zero-shot and few-shot learning. Nonetheless, the conventional method of employing fixed prompts suffers from distributional s…
IPO: Interpretable Prompt Optimization for Vision-Language Models
Yingjun Du, Wenfang Sun, Cees G. M. Snoek
Pre-trained vision-language models like CLIP have remarkably adapted to various downstream tasks. Nonetheless, their performance heavily depends on the specificity of the input tex…