4 papers · 1 filter
Findings of the Counter Turing Test: AI-Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +16
The growing capability of large language models to produce fluent, contextually coherent text has created mounting pressure on the systems and institutions responsible for ensuring…
A Comprehensive Dataset for Human vs. AI Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +17
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustwo…
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
Danush Khanna, Gurucharan Marthi Krishna Kumar, Basab Ghosh +5
Adversarial threats against LLMs are escalating faster than current defenses can adapt. We expose a critical geometric blind spot in alignment: adversarial prompts exploit latent c…
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
Abhilekh Borah, Chhavi Sharma, Danush Khanna +12
Alignment is no longer a luxury, it is a necessity. As large language models (LLMs) enter high-stakes domains like education, healthcare, governance, and law, their behavior must r…