11 citations · 21 across the 24 of their papers we have counts for
10 papers · 1 filter
Online Self-Calibration Against Hallucination in Vision-Language Models
Minghui Chen, Chenxu Yang, Hengjie Zhu +3
Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment…
EasyVideoR1: Easier RL for Video Understanding
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve i…
Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
Jinhu Fu, Yihang Lou, Qingyi Si +3
Large Vision-Language Models (LVLMs) have achieved impressive performance across multimodal understanding and reasoning tasks, yet their internal safety mechanisms remain opaque an…
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
Xiao Wang, Qingyi Si, Jianlong Wu +3
Multimodal Large Language Models (MLLMs) have revolutionized video understanding, yet are still limited by context length when processing long videos. Recent methods compress video…
Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
Peize Li, Qingyi Si, Peng Fu +2
Retrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "ret…
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
Xiao Wang, Qingyi Si, Jianlong Wu +3
Video Large Language Models (VideoLLMs) have made significant strides in video understanding but struggle with long videos due to the limitations of their backbone LLMs. Existing s…