collaborators

7 papers

cs.AI2026

Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph

Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik

Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors…

cs.AI2026

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik

Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfa…

cs.CL2025

AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda

Mohd Nauman, Sravan Gvm, Vijay Devane +7

Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic…

cs.CL2025

BhashaBench V1: A Comprehensive Benchmark for the Quadrant of Indic Domains

Vijay Devane, Mohd Nauman, Bhargav Patel +14

The rapid advancement of large language models(LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-ag…

cs.CL2025

The Art of Breaking Words: Rethinking Multilingual Tokenizer Design

Aamod Thakur, Ajay Nagpal, Atharva Savarkar +7

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Mode…

cs.CL2025

Intent Aware Context Retrieval for Multi-Turn Agricultural Question Answering

Abhay Vijayvargia, Ajay Nagpal, Kundeshwar Pundalik +5

Indian farmers often lack timely, accessible, and language-friendly agricultural advice, especially in rural areas with low literacy. To address this gap in accessibility, this pap…