activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation

Minheng Ni, Zhengyuan Yang, Yaowen Zhang +9

We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce vis…

cs.CV2025

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

Bowen Dong, Minheng Ni, Zitong Huang +3

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…

cs.CV2025

Personalized Image Generation with Deep Generative Models: A Decade Survey

Yuxiang Wei, Yiheng Zheng, Yabo Zhang +4

Recent advancements in generative models have significantly facilitated the development of personalized content creation. Given a small set of images with user-specific concept, pe…

cs.CV2024

MR-GDINO: Efficient Open-World Continual Object Detection

Bowen Dong, Zitong Huang, Guanglei Yang +2

Open-world (OW) recognition and detection models show strong zero- and few-shot adaptation abilities, inspiring their use as initializations in continual learning methods to improv…

cs.CV2024

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Minheng Ni, Yutao Fan, Lei Zhang +1

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in re…

cs.CV2024

MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation

Yuxiang Wei, Zhilong Ji, Jinfeng Bai +3

Text-to-image (T2I) diffusion models have shown significant success in personalized text-to-image generation, which aims to generate novel images with human identities indicated by…