collaborators

6 papers

cs.CV2025

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

Haoyu Zhang, Qiaohui Chu, Meng Liu +3

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Mode…

cs.LG2025

Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models

Wei Chen, Xin Yan, Bin Wen +4

Although multimodal large language models (MLLMs) exhibit remarkable reasoning capabilities on complex multimodal understanding tasks, they still suffer from the notorious hallucin…

cs.CV2025

TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs

Yunxiao Wang, Meng Liu, Wenqi Liu +7

Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address thi…

cs.CV2025

InstructEngine: Instruction-driven Text-to-Image Alignment

Xingyu Lu, Yuhang Hu, YiFan Zhang +9

Reinforcement Learning from Human/AI Feedback (RLHF/RLAIF) has been extensively utilized for preference alignment of text-to-image models. Existing methods face certain limitations…

cs.CV2025

iMOVE: Instance-Motion-Aware Video Understanding

Jiaze Li, Yaya Shi, Zongyang Ma +7

Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understan…

cs.CV2025

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Jiankang Chen, Tianke Zhang, Changyi Liu +8

Multimodal visual language models are gaining prominence in open-world applications, driven by advancements in model architectures, training techniques, and high-quality data. Howe…