4 papers
Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment
Tanav Singh Bajaj, Nikhil Singh, Karan Anand +1
As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety properties of individual models will…
REAMS: Reasoning Enhanced Algorithm for Maths Solving
Eishkaran Singh, Tanav Singh Bajaj, Siddharth Nayak
The challenges of solving complex university-level mathematics problems, particularly those from MIT, and Columbia University courses, and selected tasks from the MATH dataset, rem…
First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI
Sourav Banerjee, Anush Mahajan, Ayushi Agarwal +1
Natural Language Inference (NLI) tasks require identifying the relationship between sentence pairs, typically classified as entailment, contradiction, or neutrality. While the curr…
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Sourav Banerjee, Ayushi Agarwal, Eishkaran Singh
The pursuit of leaderboard rankings in Large Language Models (LLMs) has created a fundamental paradox: models excel at standardized tests while failing to demonstrate genuine langu…