7 papers
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Congchao Wang, Diwakar Singh, Qiaozi Gao +3
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement lear…
Advancing Creative Physical Intelligence in Large Multimodal Models
Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10
Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded…
AcquisitionSynthesis: Targeted Data Generation using Acquisition Functions
Ishika Agarwal, Sofia Stoica, Emre Can Acikgoz +4
Data quality remains a critical bottleneck in developing capable, competitive models. Researchers have explored many ways to generate top quality samples. Some works rely on reject…
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10
Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains under…
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
Luca Viano, Ruida Zhou, Yifan Sun +4
The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
Xingang Guo, Yaxin Li, Xiangyi Kong +62
Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. Ho…