4 papers
Measuring General Intelligence with Generated Games
Vivek Verma, David Huang, William Chen +2
We present gg-bench, a collection of game environments designed to evaluate general reasoning capabilities in language models. Unlike most static benchmarks, gg-bench is a data gen…
Improving LLM Safety Alignment with Dual-Objective Optimization
Xuandong Zhao, Will Cai, Tianneng Shi +4
Existing training-time safety alignment techniques for large language models (LLMs) remain vulnerable to jailbreak attacks. Direct preference optimization (DPO), a widely deployed…
Accelerated Preference Elicitation with LLM-Based Proxies
David Huang, Francisco Marmolejo-Cossío, Edwin Lock +1
Bidders in combinatorial auctions face significant challenges when describing their preferences to an auctioneer. Classical work on preference elicitation focuses on query-based te…
Enhanced Momentum with Momentum Transformers
Max Mason, Waasi A Jagirdar, David Huang +1
The primary objective of this research is to build a Momentum Transformer that is expected to outperform benchmark time-series momentum and mean-reversion trading strategies. We ex…