17 papers · 1 filter
RECOM: A Validity Discrimination Tradeoff in Automatic Metrics for Open Ended Reddit Question Answering
Pushwitha Krishnappa, Amit Das, Vinija Jain +2
Automatic metrics are the default for evaluating LLM-generated text, yet a metric is quietly asked to do two jobs: tell genuine content alignment from surface coincidence (validity…
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
Anusa Saha, Tanmay Joshi, Vinija Jain +2
LLMs are multilingual by training, yet their lingua franca is often English, reflecting English language dominance in pretraining. Other languages remain in parametric memory but a…
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
Partha Pratim Saha, Samarth Raina, Mayur Parvatikar +4
Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally. Aligned systems still fa…
Findings of the Counter Turing Test: AI-Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +16
The growing capability of large language models to produce fluent, contextually coherent text has created mounting pressure on the systems and institutions responsible for ensuring…
A Comprehensive Dataset for Human vs. AI Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +17
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustwo…
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
Pushwitha Krishnappa, Amit Das, Vinija Jain +2
Large Language Models (LLMs) are increasingly deployed for open-domain question answering, yet their alignment with human perspectives on temporally recent information remains unde…