activity
20242026
collaborators

6 papers

cs.CV2026

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

Wenfang Sun, Hao Chen, Yingjun Du +2

Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to…

cs.LG2025

GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning

Jie Ou, Shuaihong Jiang, Yingjun Du +1

Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, DoRA, and HiRA, enable lightweight adaptation of large pre-trained models via low-rank updates. However, existing PEFT…

cs.CV2025

QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain

Wenfang Sun, Yingjun Du, Gaowen Liu +2

We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which lea…

cs.AI2025

CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation

Jie Liu, Pan Zhou, Yingjun Du +4

In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods ofte…

cs.LG2024

Prompt Diffusion Robustifies Any-Modality Prompt Learning

Yingjun Du, Gaowen Liu, Yuzhang Shang +3

Foundation models enable prompt-based classifiers for zero-shot and few-shot learning. Nonetheless, the conventional method of employing fixed prompts suffers from distributional s…

cs.LG2024

IPO: Interpretable Prompt Optimization for Vision-Language Models

Yingjun Du, Wenfang Sun, Cees G. M. Snoek

Pre-trained vision-language models like CLIP have remarkably adapted to various downstream tasks. Nonetheless, their performance heavily depends on the specificity of the input tex…