Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
Prasham Titiya, Jainil Trivedi, Chitta Baral +1
Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remain difficult for Multimodal Large Langu…
cs.CV2025
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
Shikhhar Siingh, Abhinav Rawat, Chitta Baral +1
Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevan…