Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
Spiros Baxevanakis, Platon Karageorgis, Ioannis Dravilas +1
Training Vision Transformers (ViTs) presents significant challenges, one of which is the emergence of artifacts in attention maps, hindering their interpretability. Darcet et al. (…
cs.CV2026
BLM-Guard: Explainable Multimodal Ad Moderation with Chain-of-Thought and Policy-Aligned Rewards
Yiran Yang, Zhaowei Liu, Yuan Yuan +10
Short-video platforms now host vast multimodal ads whose deceptive visuals, speech and subtitles demand finer-grained, policy-driven moderation than community safety filters. We pr…
cs.CV2024
Mobius: A High Efficient Spatial-Temporal Parallel Training Paradigm for Text-to-Video Generation Task
Yiran Yang, Jinchao Zhang, Ying Deng +1
Inspired by the success of the text-to-image (T2I) generation task, many researchers are devoting themselves to the text-to-video (T2V) generation task. Most of the T2V frameworks…