activity
20242026
collaborators

6 papers

cs.LG2026

Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework

David Huang, Jaewon Chang, Avidan Shah +2

The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection classifiers. When new jailbreak…

cs.GT2025

Accelerated Preference Elicitation with LLM-Based Proxies

David Huang, Francisco Marmolejo-Cossío, Edwin Lock +1

Bidders in combinatorial auctions face significant challenges when describing their preferences to an auctioneer. Classical work on preference elicitation focuses on query-based te…

cs.CL2025

Improving LLM Safety Alignment with Dual-Objective Optimization

Xuandong Zhao, Will Cai, Tianneng Shi +4

Existing training-time safety alignment techniques for large language models (LLMs) remain vulnerable to jailbreak attacks. Direct preference optimization (DPO), a widely deployed…

cs.AI2025

Measuring General Intelligence with Generated Games

Vivek Verma, David Huang, William Chen +2

We present gg-bench, a collection of game environments designed to evaluate general reasoning capabilities in language models. Unlike most static benchmarks, gg-bench is a data gen…

cs.CL2025

SpeechVerse: A Large-scale Generalizable Audio Language Model

Nilaksh Das, Saket Dingliwal, Srikanth Ronanki +14

Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have f…

q-fin.CP2024

Enhanced Momentum with Momentum Transformers

Max Mason, Waasi A Jagirdar, David Huang +1

The primary objective of this research is to build a Momentum Transformer that is expected to outperform benchmark time-series momentum and mean-reversion trading strategies. We ex…