8 papers
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
Zhenpeng Huang, Jiaqi Li, Zihan Jia +6
We present LongVPO, a novel two-stage Direct Preference Optimization framework that enables short-context vision-language models to robustly understand ultra-long videos without an…
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
Yicheng Xiao, Lin Song, Rui Yang +4
Recent advances have highlighted the benefits of scaling language models to enhance performance across a wide range of NLP tasks. However, these approaches still face limitations i…
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Yicheng Xiao, Lin Song, Yukang Chen +7
Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks. We introduce MindOmni, a unified multimodal large language model that addre…
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Yicheng Xiao, Lin Song, Rui Yang +6
With the advancement of language models, unified multimodal understanding and generation have made significant strides, with model architectures evolving from separated components…
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
Yuxin Yang, Yinan Zhou, Yuxin Chen +8
Composed Image Retrieval (CIR) aims to retrieve target images from a gallery based on a reference image and modification text as a combined query. Recent approaches focus on balanc…
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
Cheng Cheng, Lin Song, Di An +4
Autoregressive (AR) image generators offer a language-model-friendly approach to image generation by predicting discrete image tokens in a causal sequence. However, unlike diffusio…