9 papers
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
Xichen Ding, Jianzhe Gao, Cong Pan +2
Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. W…
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
Jie Qin, Jiancheng Huang, Limeng Qiao +1
Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understand…
LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models
Yiming Hao, Mutian Xu, Chongjie Ye +4
Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to th…
RunawayEvil: Jailbreaking the Image-to-Video Generative Models
Songping Wang, Rufan Qian, Yueming Lyu +5
Image-to-Video (I2V) generation synthesizes dynamic visual content from image and text inputs, providing significant creative control. However, the security of such multimodal syst…
MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection
Shengtian Yang, Yue Feng, Yingshi Liu +2
Video Anomaly Detection (VAD) aims to locate unusual activities or behaviors within videos. Recently, offline VAD has garnered substantial research attention, which has been invigo…
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
Yifan Chang, Jie Qin, Limeng Qiao +4
Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind…