most citedSurvey of Large Multimodal Model Datasets, Application Categories and Taxonomy

2 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.AI2025

FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal +7

Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such…

cs.AI2025

Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation

Priyaranjan Pattnayak, Amit Agarwal, Hansa Meghwani +2

Retrieval-Augmented Generation (RAG) systems and large language model (LLM)-powered chatbots have significantly advanced conversational AI by combining generative capabilities with…

cs.LG2025

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

Eunsu Kim, Haneul Yoo, Guijin Son +3

As large language models (LLMs) continue to advance, the need for up-to-date and well-organized benchmarks becomes increasingly critical. However, many existing datasets are scatte…

cs.CL2025

Tokenization Matters: Improving Zero-Shot NER for Indic Languages

Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Amit Agarwal

Tokenization is a critical component of Natural Language Processing (NLP), especially for low resource languages, where subword segmentation influences vocabulary structure and dow…

cs.AI20242 cited

Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy

Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Bhargava Kumar +4

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, i…

cs.CL20242 cited

Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts

Amit Agarwal, Hitesh Patel, Priyaranjan Pattnayak +3

The development of robust Document AI models has been constrained by limited access to high-quality, labeled datasets, primarily due to data privacy concerns, scarcity, and the hig…