1 paper
Jindi Guo, Chaozheng Huang, Xi Fang
We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text directly from visual context.…