collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

Bowen Liu, Shuning Wang, Xinpeng Ding +3

Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual…

cs.CV2026

DenseMLLM: Standard Multimodal LLMs for Dense Prediction

Yi Li, Hongze Shen, Lexiang Tang +6

Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in high-level visual understanding. However, extending these models to fine-grained dense predic…

cs.CV2026

MedHorizon: Towards Long-context Medical Video Understanding in the Wild

Bodong Du, Bowen Liu, Yang Yu +8

Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understand…

cs.CV2026

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

Runhui Huang, Xinpeng Ding, Chunwei Wang +7

High-resolution inputs enable Large Vision-Language Models (LVLMs) to discern finer visual details, enhancing their comprehension capabilities. To reduce the training and computati…

cs.CV2025

Token Activation Map to Visually Explain Multimodal LLMs

Yi Li, Hualiang Wang, Xinpeng Ding +2

Multimodal large language models (MLLMs) are broadly empowering various fields. Despite their advancements, the explainability of MLLMs remains less explored, hindering deeper unde…

cs.CV2025

PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning

Xinpeng Ding, Kui Zhang, Jianhua Han +3

Direct Preference Optimization (DPO) helps reduce hallucinations in Video Multimodal Large Language Models (VLLMs), but its reliance on offline preference data limits adaptability…