11 papers · 1 filter
ShutterMuse: Capture-Time Photography Guidance with MLLMs
Jiayu Li, Yixiao Fang, Tianyu Hu +5
Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post-hoc crop prediction…
ImageAttributionBench: How Far Are We from Generalizable Attribution?
Tingshu Mou, Zhipeng Wei, Chao Gong +2
The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image provenance and misinformation…
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
Tingshu Mou, Jiabo He, Renying Wang +5
Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leavin…
OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens
Yiying Yang, Wei Cheng, Sijin Chen +5
OmniLottie is a versatile framework that generates high quality vector animations from multi-modal instructions. For flexible motion and visual content control, we focus on Lottie,…
RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
Xinchang Wang, Yunhao Chen, Yuechen Zhang +4
Recent image generators produce photo-realistic content that undermines the reliability of downstream recognition systems. As visual appearance cues become less pronounced, appeara…
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
Jiaming Zhang, Xin Wang, Xingjun Ma +3
Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data through joint embedding spaces.…