activity
20242026
collaborators

8 papers

cs.CV2026

LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization

Zhenpeng Huang, Jiaqi Li, Zihan Jia +6

We present LongVPO, a novel two-stage Direct Preference Optimization framework that enables short-context vision-language models to robustly understand ultra-long videos without an…

cs.CL2025

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation

Yicheng Xiao, Lin Song, Rui Yang +4

Recent advances have highlighted the benefits of scaling language models to enhance performance across a wide range of NLP tasks. However, these approaches still face limitations i…

cs.AI2025

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Yicheng Xiao, Lin Song, Yukang Chen +7

Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks. We introduce MindOmni, a unified multimodal large language model that addre…

cs.CV2025

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

Yicheng Xiao, Lin Song, Rui Yang +6

With the advancement of language models, unified multimodal understanding and generation have made significant strides, with model architectures evolving from separated components…

cs.CV2025

DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval

Yuxin Yang, Yinan Zhou, Yuxin Chen +8

Composed Image Retrieval (CIR) aims to retrieve target images from a gallery based on a reference image and modification text as a combined query. Recent approaches focus on balanc…

cs.CV2025

From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation

Cheng Cheng, Lin Song, Di An +4

Autoregressive (AR) image generators offer a language-model-friendly approach to image generation by predicting discrete image tokens in a causal sequence. However, unlike diffusio…