8 papers · 1 filter
Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models
Yifeng Wu, Huimin Huang, Ruiluo Wu +5
Multimodal Large Language Models (MLLMs) often lose track of the right image regions during fine-grained spatial reasoning, because a textual query rarely carries any explicit geom…
Learning Latent Proxies for Controllable Single-Image Relighting
Haoze Zheng, Zihao Wang, Xianfeng Wu +5
Single-image relighting is highly under-constrained: small illumination changes can produce large, nonlinear variations in shading, shadows, and specularities, while geometry and m…
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
Xianfeng Wu, Yajing Bai, Minghan Li +5
Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic envi…
ScalingAR: Scaling Confidence for Autoregressive Image Generation
Harold Haodong Chen, Xianfeng Wu, Wen-Jie Shu +4
Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregressive (AR) image generation rem…
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models
Xuran Ma, Yexin Liu, Yaofu Liu +5
Recent advances in diffusion models have demonstrated remarkable capabilities in video generation. However, the computational intensity remains a significant challenge for practica…
Temporal Regularization Makes Your Video Generator Stronger
Harold Haodong Chen, Haojian Huang, Xianfeng Wu +5
Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and dive…