1 citations · 1 across the 7 of their papers we have counts for
9 papers · 1 filter
From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks
Nilanjana Das, Mathew Dawit, Aman Chadha +1
Jailbreak attacks expose a persistent failure mode in safety-aligned LLMs: models can be pushed into harmful behavior, but the internal representations enabling this shift remain p…
Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
Yash Saxena, Ankur Padia, Mandar S Chaudhary +3
Retrieval-Augmented Generation (RAG) systems deployed in sensitive domains must provide interpretable evidence selection and robust safeguards against data poisoning, yet current a…
Structure-Aware Chunking for Tabular Data in Retrieval-Augmented Generation
Pooja Guttal, Varun Magotra, Vasudeva Mahavishnu +3
Tabular documents such as CSV and Excel files are widely used in enterprise data pipelines, yet existing chunking strategies for retrieval-augmented generation (RAG) are primarily…
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba +4
Do LLMs genuinely incorporate external definitions, or do they primarily rely on their parametric knowledge? To address these questions, we conduct controlled experiments across mu…
Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution
Yash Saxena, Raviteja Bommireddy, Ankur Padia +1
Trustworthy Large Language Models (LLMs) must cite human-verifiable sources in high-stakes domains such as healthcare, law, academia, and finance, where even small errors can have…
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
Nilanjana Das, Edward Raff, Aman Chadha +1
As the AI systems become deeply embedded in social media platforms, we've uncovered a concerning security vulnerability that goes beyond traditional adversarial attacks. It becomes…