collaborators

5 papers

cs.CV2025

SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency

Quanjian Song, Donghao Zhou, Jingyu Lin +5

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus o…

cs.CV2025

Medical Large Vision Language Models with Multi-Image Visual Ability

Xikai Yang, Juzheng Miao, Yuchen Yuan +4

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in process…

cs.CV2025

MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

Donghao Zhou, Jiancheng Huang, Jinbin Bai +5

Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllab…

cs.CV2025

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

Ziyu Guo, Ray Zhang, Hao Chen +4

The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In…

cs.CV2024

Point Cloud Understanding via Attention-Driven Contrastive Learning

Yi Wang, Jiaze Wang, Ziyu Guo +5

Recently Transformer-based models have advanced point cloud understanding by leveraging self-attention mechanisms, however, these methods often overlook latent information in less…