5 papers
The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
Yiheng Tao, Kaiwen Cheng, Yao Lu +2
Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context performance. We identify that existing approache…
CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models
Ji Guo, Xiaolong Qin, Cencen Liu +3
Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their applications become increasi…
Video Generation with Predictive Latents
Yian Zhao, Feng Wang, Qiushan Guo +4
Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency an…
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Zijun Liao, Yian Zhao, Xin Shan +5
Real-time object detection has achieved substantial progress through meticulously designed architectures and optimization strategies. However, the pursuit of high-speed inference v…
UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
Ailing Zhang, Lina Lei, Dehong Kong +7
Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major f…