4 papers
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
Evangelia Christakopoulou, Vivekkumar Patel, Hemanth Velaga +3
Large-scale commercial search systems optimize for relevance to drive successful sessions that help users find what they are looking for. To maximize relevance, we leverage two com…
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
Aditya Maheshwari, Amit Gajkeshwar, Kaushal Sharma +1
As Large Language Models (LLMs) becomes a popular source for religious knowledge, it is important to know if it treats different groups fairly. This study is the first to measure h…
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
Ayush Maheshwari, Kaushal Sharma, Vivek Patel +1
While large language models excel on high-resource multilingual tasks, low- and extremely low-resource Indic languages remain severely under-evaluated. We present IndicParam, a hum…
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
Ayush Maheshwari, Kaushal Sharma, Vivek Patel +1
Large language models have been widely evaluated on tasks such as comprehension, summarization, code generation, etc. However, their performance on graduate-level, culturally groun…