From the 1 of 12 linked papers with an AI index.
6 papers · 1 filter
Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +3
The paper introduces GMM-EVA, a training-free framework that uses Gaussian Mixture Models to detect event-level structures in long videos and allocate visual tokens by selecting on…
Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation
Siyi Chen, Shaowei Liu, Yixuan Jia +4
Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its suc…
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
Chaojie Yang, Tian Li, Yue Zhang +1
Diffusion Transformer (DiT) architectures have significantly advanced Text-to-Image (T2I) generation but suffer from prohibitive computational costs and deployment barriers. To add…
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +5
Large Vision-Language Models (LVLMs) suffer from serious hallucination problems, where the model-generated responses are inconsistent with the visual inputs. Existing hallucination…
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
Yuxin Yang, Yinan Zhou, Yuxin Chen +8
Composed Image Retrieval (CIR) aims to retrieve target images from a gallery based on a reference image and modification text as a combined query. Recent approaches focus on balanc…
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
Shengkai Zhang, Nianhong Jiao, Tian Li +4
We propose an effective method for inserting adapters into text-to-image foundation models, which enables the execution of complex downstream tasks while preserving the generalizat…