collaborators

6 papers

cs.CV2025

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

Bowen Dong, Minheng Ni, Zitong Huang +3

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…

cs.AI2025

Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving

Zixian Guo, Ming Liu, Qilong Wang +4

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end tra…

cs.CV2025

Personalized Image Generation with Deep Generative Models: A Decade Survey

Yuxiang Wei, Yiheng Zheng, Yabo Zhang +4

Recent advancements in generative models have significantly facilitated the development of personalized content creation. Given a small set of images with user-specific concept, pe…

cs.CV2024

MR-GDINO: Efficient Open-World Continual Object Detection

Bowen Dong, Zitong Huang, Guanglei Yang +2

Open-world (OW) recognition and detection models show strong zero- and few-shot adaptation abilities, inspiring their use as initializations in continual learning methods to improv…

cs.RO2024

Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy

Minheng Ni, Lei Zhang, Zihan Chen +4

Unthinking execution of human instructions in robotic manipulation can lead to severe safety risks, such as poisonings, fires, and even explosions. In this paper, we present respon…

cs.CV2024

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Minheng Ni, Yutao Fan, Lei Zhang +1

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in re…