activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models

Nilesh Nayan, Aishwarya Sampath Kumar, Rishiraj Girmal +5

Safety benchmarks assume that test-condition behavior predicts deployment behavior, an assumption that fails if models detect evaluation cues and adapt. This opens a gap between be…

cs.CL2025

Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis

Anushka Yadav, Isha Nalawade, Srujana Pillarichety +7

The emergence of reasoning models and their integration into practical AI chat bots has led to breakthroughs in solving advanced math, deep search, and extractive question answerin…

cs.CL2024

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text

Reshmi Ghosh, Tianyi Yao, Lizzy Chen +5

Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to con…

cs.CL2024

Leveraging Language Models to Detect Greenwashing

Avalon Vinella, Margaret Capetz, Rebecca Pattichis +3

In recent years, climate change repercussions have increasingly captured public interest. Consequently, corporations are emphasizing their environmental efforts in sustainability r…

cs.CL2024

Quantifying reliance on external information over parametric knowledge during Retrieval Augmented Generation (RAG) using mechanistic analysis

Reshmi Ghosh, Rahul Seetharaman, Hitesh Wadhwa +6

Retrieval Augmented Generation (RAG) is a widely used approach for leveraging external context in several natural language applications such as question answering and information r…

cs.CL2024

From RAGs to rich parameters: Probing how language models utilize external knowledge over parametric information for factual queries

Hitesh Wadhwa, Rahul Seetharaman, Somyaa Aggarwal +6

Retrieval Augmented Generation (RAG) enriches the ability of language models to reason using external context to augment responses for a given user prompt. This approach has risen…