activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV202625 cited

Efficient Multimodal Large Language Models: A Survey

Yizhang Jin, Jian Li, Yexin Liu +10

In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning.…

cs.CV2025

Can video generation replace cinematographers? Research on the cinematic language of generated video

Xiaozhe Li, Kai WU, Siyi Yang +12

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing…

cs.CV2024

Exploring Real&Synthetic Dataset and Linear Attention in Image Restoration

Yuzhen Du, Teng Hu, Jiangning Zhang +6

Image restoration (IR) aims to recover high-quality images from degraded inputs, with recent deep learning advancements significantly enhancing performance. However, existing metho…

cs.CV2024

CustAny: Customizing Anything from A Single Example

Lingjie Kong, Kai Wu, Xiaobin Hu +8

Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, i…

cs.CV2024

NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models

Kai Wu, Boyuan Jiang, Zhengkai Jiang +5

Multimodal large language models (MLLMs) contribute a powerful mechanism to understanding visual information building on large language models. However, MLLMs are notorious for suf…