6 papers
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Haichuan Wang, Tao Lin, Lingkai Kong +3
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…
Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation
Dingyi Yang, Junqi Zhao, Xue Li +2
Cognitive anthropology suggests that the distinction of human intelligence lies in the ability to infer other individuals' knowledge states and understand their intentions. In comp…
JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
Ce Chi, Xing Wang, Zhendong Wang +9
In this work, we present JT-DA-8B (JiuTian Data Analyst 8B), a specialized large language model designed for complex table reasoning tasks across diverse real-world scenarios. To a…
From Best Responses to Learning: Investment Efficiency in Dynamic Environment
Ce Li, Qianfan Zhang, Weiqiang Zheng
We study the welfare of a mechanism in a dynamic environment where a learning investor can make a costly investment to change her value. In many real-world problems, the common ass…
Information Design with Unknown Prior
Ce Li, Tao Lin
Information designers, such as online platforms, often do not know the beliefs of their receivers. We design learning algorithms so that the information designer can learn the rece…
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Ce Li, Xiaofan Liu, Zhiyan Song +10
The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large l…