5 papers · 1 filter
Efficient Multimodal Large Language Models: A Survey
Yizhang Jin, Jian Li, Yexin Liu +10
In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning.…
Can video generation replace cinematographers? Research on the cinematic language of generated video
Xiaozhe Li, Kai WU, Siyi Yang +12
Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing…
Exploring Real&Synthetic Dataset and Linear Attention in Image Restoration
Yuzhen Du, Teng Hu, Jiangning Zhang +6
Image restoration (IR) aims to recover high-quality images from degraded inputs, with recent deep learning advancements significantly enhancing performance. However, existing metho…
CustAny: Customizing Anything from A Single Example
Lingjie Kong, Kai Wu, Xiaobin Hu +8
Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, i…
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
Kai Wu, Boyuan Jiang, Zhengkai Jiang +5
Multimodal large language models (MLLMs) contribute a powerful mechanism to understanding visual information building on large language models. However, MLLMs are notorious for suf…