5 papers
Efficient Multimodal Large Language Models: A Survey
Yizhang Jin, Jian Li, Yexin Liu +10
In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning.…
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
Kaifeng Wu, Junyan Wu, Qiang Liu +2
Long-document topic segmentation plays an important role in information retrieval and document understanding, yet existing methods still show clear shortcomings in ultra-long text…
Can video generation replace cinematographers? Research on the cinematic language of generated video
Xiaozhe Li, Kai WU, Siyi Yang +12
Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing…
Exploring Real&Synthetic Dataset and Linear Attention in Image Restoration
Yuzhen Du, Teng Hu, Jiangning Zhang +6
Image restoration (IR) aims to recover high-quality images from degraded inputs, with recent deep learning advancements significantly enhancing performance. However, existing metho…
CustAny: Customizing Anything from A Single Example
Lingjie Kong, Kai Wu, Xiaobin Hu +8
Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, i…