4 papers
Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis
Jiayu Fu, Mourad Heddaya, Chenhao Tan
Numerous math benchmarks exist to evaluate LLMs' mathematical capabilities. However, most involve extensive manual effort and are difficult to scale. Consequently, they cannot keep…
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
Cheril Shah, Akshit Agarwal, Kanak Garg +1
Bilateral negotiation is a complex, context-sensitive task in which human negotiators dynamically adjust anchors, pacing, and flexibility to exploit power asymmetries and informal…
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
Mourad Heddaya, Kyle MacMillan, Anup Malani +2
This paper introduces CaseSumm, a novel dataset for long-context summarization in the legal domain that addresses the need for longer and more complex datasets for summarization ev…
Causal Micro-Narratives
Mourad Heddaya, Qingcheng Zeng, Chenhao Tan +2
We present a novel approach to classify causal micro-narratives from text. These narratives are sentence-level explanations of the cause(s) and/or effect(s) of a target subject. Th…