10 papers
Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
Riju Marwah, Ritvik Garimella, Vishal Pallagani +3
Autoregressive language models frequently degrade during long-horizon generation, producing repetitive text, losing instruction adherence, and exhibiting unstable entropy. Despite…
Findings of the Counter Turing Test: AI-Generated Image Detection
Rajarshi Roy, Nasrin Imanpour, Ashhar Aziz +16
The rapid advancements in generative AI technologies, such as Stable Diffusion, DALL-E, and Midjourney, have significantly transformed the creation of synthetic visual content. Whi…
Findings of the Counter Turing Test: AI-Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +16
The growing capability of large language models to produce fluent, contextually coherent text has created mounting pressure on the systems and institutions responsible for ensuring…
A Comprehensive Dataset for Human vs. AI Generated Image Detection
Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai +17
Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also en…
A Comprehensive Dataset for Human vs. AI Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +17
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustwo…
SAAG: Structured Agent Assessment and Grounding
Ritvik Garimella, Vedant Khandelwal, Anvi Kohli +1
Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema w…