9 papers · 1 filter
ProCrop: Learning Aesthetic Image Cropping from Professional Compositions
Ke Zhang, Tianyu Ding, Jiachen Jiang +4
Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require an…
OFER: Occluded Face Expression Reconstruction
Pratheba Selvaraju, Victoria Fernandez Abrevaya, Timo Bolkart +4
Reconstructing 3D face models from a single image is an inherently ill-posed problem, which becomes even more challenging in the presence of occlusions. In addition to fewer availa…
Analyzing and Mitigating Model Collapse in Rectified Flow Models
Huminhao Zhu, Fangyikang Wang, Tianyu Ding +2
Training with synthetic data is becoming increasingly inevitable as synthetic content proliferates across the web, driven by the remarkable performance of recent deep generative mo…
Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach
Jiwei Guan, Tianyu Ding, Longbing Cao +3
Vision-language pretraining (VLP) with transformers has demonstrated exceptional performance across numerous multimodal tasks. However, the adversarial robustness of these models h…
CaesarNeRF: Calibrated Semantic Representation for Few-shot Generalizable Neural Rendering
Haidong Zhu, Tianyu Ding, Tianyi Chen +3
Generalizability and few-shot learning are key challenges in Neural Radiance Fields (NeRF), often due to the lack of a holistic understanding in pixel-level rendering. We introduce…
FORA: Fast-Forward Caching in Diffusion Transformer Acceleration
Pratheba Selvaraju, Tianyu Ding, Tianyi Chen +2
Diffusion transformers (DiT) have become the de facto choice for generating high-quality images and videos, largely due to their scalability, which enables the construction of larg…