2 citations · 8 across the 27 of their papers we have counts for
4 papers · 1 filter
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Bhargava Kumar +4
Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, i…
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
Hitesh Laxmichand Patel, Amit Agarwal, Bhargava Kumar +2
Accurate barcode detection and decoding in Identity documents is crucial for applications like security, healthcare, and education, where reliable data extraction and verification…
MVTamperBench: Evaluating Robustness of Vision-Language Models
Amit Agarwal, Srikant Panda, Angeline Charles +8
Multimodal Large Language Models (MLLMs), are recent advancement of Vision-Language Models (VLMs) that have driven major advances in video understanding. However, their vulnerabili…
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
Amit Agarwal, Hitesh Patel, Priyaranjan Pattnayak +3
The development of robust Document AI models has been constrained by limited access to high-quality, labeled datasets, primarily due to data privacy concerns, scarcity, and the hig…