12 papers
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Xinkui Zhao, Enbo Chen, Yifan Zhang +4
Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However,…
Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value
Yixian Xu, Shengjie Luo, Liwei Wang +2
Diffusion models have achieved remarkable success in generative modeling. Despite more stable training, the loss of diffusion models is not indicative of absolute data-fitting qual…
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
Zhaoyang Li, Dongjun Qian, Kai Su +6
Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, exis…
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
Xinkui Zhao, Zuxin Wang, Yifan Zhang +6
The rapid development of multimodal large-language models (MLLMs) has significantly expanded the scope of visual language reasoning, enabling unified systems to interpret and descr…
S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation
Jiechao Gao, Chang Liu, Yuangang Li
Radiology Report Generation (RRG) aims to automatically generate diagnostic reports from radiology images. To achieve this, existing methods have leveraged the powerful cross-modal…
UnZipLoRA: Separating Content and Style from a Single Image
Chang Liu, Viraj Shah, Aiyu Cui +1
This paper introduces UnZipLoRA, a method for decomposing an image into its constituent subject and style, represented as two distinct LoRAs (Low-Rank Adaptations). Unlike existing…