8 papers
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
Minheng Ni, Zhengyuan Yang, Yaowen Zhang +9
We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce vis…
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
Minheng Ni, Yutao Fan, Zhengyuan Yang +6
Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However,…
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
Zixian Guo, Ming Liu, Qilong Wang +4
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end tra…
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Bowen Dong, Minheng Ni, Zitong Huang +3
Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
Minheng Ni, Lei Zhang, Zihan Chen +4
Unthinking execution of human instructions in robotic manipulation can lead to severe safety risks, such as poisonings, fires, and even explosions. In this paper, we present respon…
Personalized Image Generation with Deep Generative Models: A Decade Survey
Yuxiang Wei, Yiheng Zheng, Yabo Zhang +4
Recent advancements in generative models have significantly facilitated the development of personalized content creation. Given a small set of images with user-specific concept, pe…