activity
20202026
most citedCobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

6 citations · 20 across the 26 of their papers we have counts for

collaborators
Showing cs.CVShow all

18 papers · 1 filter

cs.CV2025

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

Fengyuan Dai, Zifeng Zhuang, Yufei Huang +4

Diffusion models have emerged as powerful generative tools across various domains, yet tailoring pre-trained models to exhibit specific desirable properties remains challenging. Wh…

cs.CV2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

Yang Liu, Ming Ma, Xiaomin Yu +5

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…

cs.CV2025

Exploring the Evolution of Physics Cognition in Video Generation: A Survey

Minghui Lin, Xiang Wang, Yishan Wang +8

Recent advancements in video generation have witnessed significant progress, especially with the rapid advancement of diffusion models. Despite this, their deficiencies in physical…

cs.CV2024

Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration

Yuhang Han, Xuyang Liu, Zihan Zhang +6

The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-worl…

cs.CV2024

ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification

Can Cui, Siteng Huang, Wenxuan Song +3

To address the occlusion issues in person Re-Identification (ReID) tasks, many methods have been proposed to extract part features by introducing external spatial information. Howe…

cs.CV2024★ 1 cited

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

Yang Liu, Pengxiang Ding, Siteng Huang +3

Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video…