10 papers
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
Erica Zhang, Fangzhao Zhang, Aneesh Pappu +5
Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonical testbed for agentic languag…
PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting
Hao Wu, Fan Xu, Yuxu Lu +9
Coupled spatiotemporal forecasting is important for predicting the future evolution of multiple interacting dynamical systems, such as in climate models. However, existing methods…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
Optimizer-Induced Mode Connectivity: From AdamW to Muon
Fangzhao Zhang, Sungyoon Kim, Erica Zhang +2
Mode connectivity has been widely studied, yet the role of the optimizer remains underexplored. We revisit it through optimizer-induced implicit regularization, asking how connecti…
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
Rania Elbadry, Ahmed Heakl, Fan Zhang +4
Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a structural one: temporal drift,…
Learning When to Trust LLM Priors: A Validated Framework for Semantic Prior Integration
Erica Zhang, Naomi Sagan, Danny Tse +3
Large language models (LLMs) encode rich semantic knowledge that can be useful for supervised learning, but their outputs are unreliable as statistical priors: they may be noisy, m…