collaborators

12 papers

cs.AI2026

Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models

John Scoville, Shengzhuang Chen, Yejin Bang +2

Recent meta-reasoning frameworks improve LLM reasoning by wrapping chain-of-thought generation in an iterative control loop, allowing more effective backtracking, termination of re…

cs.AI2026

Aligning Language Model Benchmarks with Pairwise Preferences

Marco Gutierrez, Xinyi Leng, Hannah Cyberey +3

Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real…

cs.LG2026

PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

Hainiu Xu, Italo Luis da Silva, Jiangnan Ye +7

Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective. While existing safety rese…

cs.LG2026

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training

Lukas Thede, Stefan Winzeck, Zeynep Akata +1

Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfortunately, post-training is known…

cs.AI2026

Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings

Andrew M. Bean, Nabeel Seedat, Shengzhuang Chen +1

The prohibitive cost of evaluating large language models (LLMs) on comprehensive benchmarks necessitates the creation of small yet representative data subsets (i.e., tiny benchmark…

cs.AI2026

To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands

Fangyi Yu, Nabeel Seedat, Jonathan Richard Schwarz +1

Language models deployed in high-stakes professional settings face conflicting demands from users, institutional authorities, and professional norms. How models act when these dema…