2 citations · 4 across the 6 of their papers we have counts for
6 papers
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal +7
Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such…
Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation
Priyaranjan Pattnayak, Amit Agarwal, Hansa Meghwani +2
Retrieval-Augmented Generation (RAG) systems and large language model (LLM)-powered chatbots have significantly advanced conversational AI by combining generative capabilities with…
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
Eunsu Kim, Haneul Yoo, Guijin Son +3
As large language models (LLMs) continue to advance, the need for up-to-date and well-organized benchmarks becomes increasingly critical. However, many existing datasets are scatte…
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Amit Agarwal
Tokenization is a critical component of Natural Language Processing (NLP), especially for low resource languages, where subword segmentation influences vocabulary structure and dow…
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Bhargava Kumar +4
Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, i…
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
Amit Agarwal, Hitesh Patel, Priyaranjan Pattnayak +3
The development of robust Document AI models has been constrained by limited access to high-quality, labeled datasets, primarily due to data privacy concerns, scarcity, and the hig…