7 papers
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research
Haokun Liu, Filbert Aurelian Tjiaranata, Chenhao Tan
AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published research has become both harder and…
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
Cristina Garbacea, Heran Wang, Chenhao Tan
With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important chal…
Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis
Jiayu Fu, Mourad Heddaya, Chenhao Tan
Numerous math benchmarks exist to evaluate LLMs' mathematical capabilities. However, most involve extensive manual effort and are difficult to scale. Consequently, they cannot keep…
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
Haokun Liu, Sicong Huang, Jingyu Hu +2
There is growing interest in hypothesis generation with large language models (LLMs). However, fundamental questions remain: what makes a good hypothesis, and how can we systematic…
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
Mingxuan Li, Hanchen Li, Chenhao Tan
Large language models (LLMs) have demonstrated great potential for automating the evaluation of natural language generation. Previous frameworks of LLM-as-a-judge fall short in two…
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
Haokun Liu, Yangqiaoyu Zhou, Mingxuan Li +2
AI holds promise for transforming scientific processes, including hypothesis generation. Prior work on hypothesis generation can be broadly categorized into theory-driven and data-…