9 papers
Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment
Haokai Zhao, Yunze Xiao, Weihao Xuan +3
Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-pro…
CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning
Rahul Markasserithodi, Aditya Joshi, Yuekang Li +3
Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can bypass safety filters even on f…
A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents
Bao Gia Doan, Aditya Joshi, Pantelis Elinas +4
RAG-based question-answering (QA) in specialist domains faces a cold-start problem: lack of evaluative benchmarks and absence of labeled data for post-training. We present DoRA (Do…
TeamUp: Semantic Project Matching and Team Formation for Learning at Scale
Dhruv Gulwani, Basem Suleiman, Aditya Joshi +1
Project-based learning improves student engagement and learning outcomes, yet allocating students to appropriately challenging projects while forming cognitively diverse teams rema…
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
Fang Wu, Haokai Zhao, Da Xing +17
Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In th…
Harnessing Test-time Adaptation for NLU tasks Involving Dialects of English
Duke Nguyen, Aditya Joshi, Flora Salim
Test-time domain adaptation (TTDA) is an excellent method which helps generalize models across domains, tasks, and distributions without the use of labeled datasets. Thus, TTDA is…