7 papers
HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
Hezhen Hu, Wangbo Zhao, Lanqing Guo +6
In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving both photorealism and generaliz…
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
Ruina Hu, Chen Wang, Lai Wei +5
Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only re…
Neural-Driven Image Editing
Pengfei Zhou, Jie Xia, Xiaopeng Peng +15
Traditional image editing typically relies on manual prompting, making it labor-intensive and inaccessible to individuals with limited motor control or language abilities. Leveragi…
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang +18
Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, cu…
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
Zhaopan Xu, Pengfei Zhou, Weidong Tang +7
Multimodal large language models (MLLMs) have achieved remarkable success in vision-language tasks, but their reliance on vast, internet-sourced data raises significant privacy and…
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
Pengfei Zhou, Fanrui Zhang, Xiaopeng Peng +17
Multimodal reasoning, which integrates language and visual cues into problem solving and decision making, is a fundamental aspect of human intelligence and a crucial step toward ar…