4 papers
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
Shuo Lu, Jianjie Cheng, Yinuo Xu +16
Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the…
StyMam: A Mamba-Based Generator for Artistic Style Transfer
Zhou Hong, Ning Dong, Yicheng Di +8
Image style transfer aims to integrate the visual patterns of a specific artistic style into a content image while preserving its content structure. Existing methods mainly rely on…
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
Run Ling, Ke Cao, Jian Lu +15
Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity.…
RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation
Run Ling, Wenji Wang, Yuting Liu +12
Personalized image generation is crucial for improving the user experience, as it renders reference images into preferred ones according to user visual preferences. Although effect…