327 citations · 1.7k across the 333 of their papers we have counts for
36 papers · 1 filter
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
Kevin Qinghong Lin, Siyuan Hu, Pan Lu +14
Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through care…
Building Legal Reward Models for Grounding and Abstention
Rilton Franzone, Valentin Noël, Puyu Wang +2
Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is in…
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect
Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) oft…
CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Hanwen Xing, Pengyun Wang, BingXu Meng +8
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Chen Tang, Yizhou Wang, Jianyu Wu +26
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and pe…
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
Zhenhao Chen, Yongqiang Chen, Chenxi Liu +7
Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally relies on uncovering causal relati…