8 papers
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
Parth Agarwal, Navya Kommuri, Trizal Garg +7
Cricket is the second most popular sport worldwide, with billions of fans seeking advanced statistical insights unavailable through standard web searches. Although LLMs have advanc…
Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization
Gaurish Trivedi, Alakh Sharma, Kartikey Singh Bhandari +4
Standard trust-region methods constrain policy updates via Kullback-Leibler (KL) divergence. However, KL controls only an average divergence and does not directly prevent rare, lar…
SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control
Adithya Chittem, Aishna Shrivastava, Sai Tarun Pendela +2
Large language models (LLMs) have gained significant traction across a wide range of fields in recent years. There is also a growing expectation for them to display human-like pers…
Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
Chinmay Pushkar, Sanchit Kabra, Dhruv Kumar +1
Large Language Models (LLMs) have demonstrated significant potential in automated software security, particularly in vulnerability detection. However, existing benchmarks primarily…
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
Sankalp Tattwadarshi Swain, Anshika Krishnatray, Dhruv Kumar +1
Existing evaluation studies on linguistic competence of large language models (LLM agents) have focused primarily on vocabulary learning, morphological rule induction, syntactic ge…
HAEPO: History-Aggregated Exploratory Policy Optimization
Gaurish Trivedi, Alakh Sharma, Kartikey Singh Bhandari +3
Exploration is essential in modern learning, from reinforcement learning environments with small neural policies to large language models (LLMs). Existing work, such as DPO, levera…