9 papers
Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation
Minyan Luo, Yuxin Zhang, Yifei Li +5
Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engag…
Out of the box age estimation through facial imagery: A Comprehensive Benchmark of Vision-Language Models vs. out-of-the-box Traditional Architectures
Simiao Ren, Xingyu Shen, Ankit Raj +8
Facial age estimation plays a critical role in content moderation, age verification, and deepfake detection. However, no prior benchmark has systematically compared modern vision-l…
GEN3D: Generating Domain-Free 3D Scenes from a Single Image
Yuxin Zhang, Ziyu Lu, Hongbo Duan +5
Despite recent advancements in neural 3D reconstruction, the dependence on dense multi-view captures restricts their broader applicability. Additionally, 3D scene generation is vit…
LumiSculpt: Enabling Consistent Portrait Lighting in Video Generation
Yuxin Zhang, Dandan Zheng, Biao Gong +5
Lighting plays a pivotal role in ensuring the naturalness and aesthetic quality of video generation. However, the impact of lighting is deeply coupled with other factors of videos,…
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
Yuxi Zhang, Yueting Li, Xinyu Du +1
Generating images from rhetorical languages remains a critical challenge for text-to-image models. Even state-of-the-art (SOTA) multimodal large language models (MLLM) fail to gene…
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Ziming Cheng, Binrui Xu, Lisheng Gong +14
With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously…