1 paper
Parth Thakkar, Ankush Agarwal, Prasad Kasu +2
While Multi-modal Large Language Models (MLLMs) have shown impressive capabilities in document understanding tasks, their ability to locate and reason about fine-grained details wi…