collaborators

11 papers

cs.CV2026

A Survey: Spatiotemporal Consistency in Video Generation

Zhiyu Yin, Kehai Chen, Xuefeng Bai +7

Video generation aims to produce temporally coherent sequences of visual frames, representing a pivotal advancement in Artificial Intelligence Generated Content (AIGC). Compared to…

cs.CL2026

Evaluating and Steering Modality Preferences in Multimodal Large Language Model

Yu Zhang, Jinlong Ma, Yongshuai Hou +5

Multi-modal large language models (MLLMs) have achieved remarkable success on complex multi-modal tasks. However, it remains insufficiently explored whether they exhibit $\textbf{m…

cs.CV2026

Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models

Yingjie Zhu, Xuefeng Bai, Kehai Chen +6

Large Vision-Language Models (LVLMs) have achieved remarkable success across a wide range of multimodal tasks, yet their robustness to spatial variations remains insufficiently und…

cs.LG2026

Improving Value-based Process Verifier via Structural Prior Injection

Zetian Sun, Dongfang Li, Baotian Hu +2

In the Large Language Model(LLM) reasoning scenario, people often estimate state value via Monte Carlo sampling. Though Monte Carlo estimation is an elegant method with less induct…

cs.CL2025

KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model

Xinping Zhao, Xinshuo Hu, Zifei Shan +14

Recent advancements in Large Language Models (LLMs)-based text embedding models primarily focus on data scaling or synthesis, yet limited exploration of training techniques and dat…

cs.CL2025

SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models

Jun Rao, Yunjie Liao, Xuebo Liu +6

Existing alignment methods for preference optimization of large language models (LLMs) aim to enhance model performance by utilizing pairs of positive and negative samples. However…