6 papers
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Benchmarking the Generality of Vision-Language-Action Models
Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka +6
Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices r…
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
Jinu Nyachhyon, Mridul Sharma, Prajwal Thapa +1
The Nepali language has distinct linguistic features, especially its complex script (Devanagari script), morphology, and various dialects,which pose a unique challenge for Natural…
Development of Pre-Trained Transformer-based Models for the Nepali Language
Prajwal Thapa, Jinu Nyachhyon, Mridul Sharma +1
Transformer-based pre-trained language models have dominated the field of Natural Language Processing (NLP) for quite some time now. However, the Nepali language, spoken by approxi…
Confirmation bias: A challenge for scalable oversight
Gabriel Recchia, Chatrik Singh Mangat, Jinu Nyachhyon +4
Scalable oversight protocols aim to empower evaluators to accurately verify AI models more capable than themselves. However, human evaluators are subject to biases that can lead to…
Local Herb Identification Using Transfer Learning: A CNN-Powered Mobile Application for Nepalese Flora
Prajwal Thapa, Mridul Sharma, Jinu Nyachhyon +1
Herb classification presents a critical challenge in botanical research, particularly in regions with rich biodiversity such as Nepal. This study introduces a novel deep learning a…