5 papers
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang +18
Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, cu…
Neural-Driven Image Editing
Pengfei Zhou, Jie Xia, Xiaopeng Peng +15
Traditional image editing typically relies on manual prompting, making it labor-intensive and inaccessible to individuals with limited motor control or language abilities. Leveragi…
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
Pengfei Zhou, Fanrui Zhang, Xiaopeng Peng +17
Multimodal reasoning, which integrates language and visual cues into problem solving and decision making, is a fundamental aspect of human intelligence and a crucial step toward ar…
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
Zhaopan Xu, Pengfei Zhou, Jiaxin Ai +6
Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, where the identification of process errors is vital for improving this ability. Recent…
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
Zhaopan Xu, Pengfei Zhou, Weidong Tang +7
Multimodal large language models (MLLMs) have achieved remarkable success in vision-language tasks, but their reliance on vast, internet-sourced data raises significant privacy and…