activity
20242026
most citedEvaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis

2 citations · 3 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

Jiuzhou Lin, Junlong Wu, Fei Zuo +11

Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typical…

cs.CV2026

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

Bingzheng Qu, Kehai Chen, Xuefeng Bai +1

Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-sp…

cs.CV2026

Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent

Yihong Tang, Kehai Chen, Xuefeng Bai +1

The advancement of Multimodal Large Language Models (MLLMs) has expanded Role-Playing Agents (RPAs) into visually grounded environments. However, human vision is inherently subject…

cs.CV2026

Mitigating Multimodal Hallucination via Phase-wise Self-reward

Yu Zhang, Chuyang Sun, Kehai Chen +3

Large Vision-Language Models (LVLMs) still struggle with vision hallucination, where generated responses are inconsistent with the visual input. Existing methods either rely on lar…

cs.CV2026

Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance

Yingjie Zhu, Xuefeng Bai, Kehai Chen +4

Reasoning over table images remains challenging for Large Vision-Language Models (LVLMs) due to complex layouts and tightly coupled structure-content information. Existing solution…

cs.CV2026

VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics

Zhiyu Yin, Zhipeng Liu, Kehai Chen +5

While current video generation focuses on text or image conditions, practical applications like video editing and vlogging often need to seamlessly connect separate clips. In our w…