6 citations · 27 across the 33 of their papers we have counts for
7 papers · 1 filter
CRANE: Knowledge Editing for Reasoning MLLMs
Han Huang, Hao Wang, Mengqi Zhang +3
The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought (CoT) reasoning before producing answers, has introduced a new challen…
MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation
Wenqing Tian, Hanyi Mao, Zhaocheng Liu +4
Subject-driven image generation is increasingly expected to support fine-grained control over multiple entities within a single image. In multi-reference workflows, users may provi…
Chatting with Images for Introspective Visual Thinking
Junfei Wu, Jian Guan, Qiang Liu +4
Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss of fine-grained visual informat…
AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
Mingqing Zhang, Zhuoning Xu, Peijie Wang +6
Accurate crop disease diagnosis is essential for sustainable agriculture and global food security. Existing methods, which primarily rely on unimodal models such as image-based cla…
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
Han Huang, Yuqi Huo, Zijia Zhao +6
Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-tex…
Interpretable Multimodal Out-of-context Detection with Soft Logic Regularization
Huanhuan Ma, Jinghao Zhang, Qiang Liu +2
The rapid spread of information through mobile devices and media has led to the widespread of false or deceptive news, causing significant concerns in society. Among different type…