From the 2 of 8 linked papers with an AI index.
4 papers · 1 filter
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning
Cheng Tang, Junzhi Ning, Min Cen +9
The paper presents SIVA-RL, a framework that uses sample-wise visual interventions to align sensitivity and invariance in multimodal reinforcement learning models, leading to bette…
EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
Jiashi Lin, Changhong Jiang, Xiangru Lin +12
The paper proposes EvoGraph-R1, a framework that lets a retrieval agent dynamically evolve multimodal knowledge hypergraphs through actions like retrieval, web search, and graph ed…
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
Yi Xin, Siqi Luo, Tianxiang Xu +13
Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and ef…
MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations
Jiyao Liu, Junzhi Ning, Chenglong Ma +14
Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images in…