collaborators

5 papers

cs.CL2025

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning

Yeyuan Wang, Dehong Gao, Rujiao Long +4

Direct Preference Optimization (DPO) has gained significant attention for its simplicity and computational efficiency in aligning large language models (LLMs). Recent advancements…

cs.CV2025

Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models

Bin Li, Dehong Gao, Yeyuan Wang +4

Despite the significant success of Large Vision-Language models(LVLMs), these models still suffer hallucinations when describing images, generating answers that include non-existen…

cs.CV2024

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models

Yeyuan Wang, Dehong Gao, Bin Li +7

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks.…

cs.CV2024

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

Yeyuan Wang, Dehong Gao, Lei Yi +4

Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coar…

cs.LG2024

MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning

Yufei Ma, Zihan Liang, Huangyu Dai +9

The growing demand for larger-scale models in the development of \textbf{L}arge \textbf{L}anguage \textbf{M}odels (LLMs) poses challenges for efficient training within limited comp…