1 citations · 1 across the 5 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety
Shikai Qiu, Xiaowen Xu, Benlei Cui +55
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI s…
cs.CV2024★ 1 cited
Text-Video Multi-Grained Integration for Video Moment Montage
Zhihui Yin, Ye Ma, Xipeng Cao +3
The proliferation of online short video platforms has driven a surge in user demand for short video editing. However, manually selecting, cropping, and assembling raw footage into…
cs.CV2024
Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
Shengyuan Liu, Bo Wang, Ye Ma +6
Existing subject-driven text-to-image generation models suffer from tedious fine-tuning steps and struggle to maintain both text-image alignment and subject fidelity. For generatin…