8 papers
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
Haoyue Yang, Zhangxiao Shen, Fan Ding +8
Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human…
FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization
Haijian Lu, Wei Wang, Jing Liu
Autoformalization aims to produce formal statements that compile and faithfully preserve the intended meaning of informal mathematics. Yet standard single-output evaluation protoco…
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
Xu Ouyang, Deyi Liu, Yuhang Cai +5
Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and q…
Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented Generation
Yuhao Wang, Ruiyang Ren, Yucheng Wang +4
Long-form question answering (LFQA) requires open-ended long-form responses that synthesize coherent, factually grounded content from multi-source evidence. This makes reinforcemen…
Emulating Clinician Cognition via Self-Evolving Deep Clinical Research
Ruiyang Ren, Yuhao Wang, Yunsen Liang +8
Clinical diagnosis is a complex cognitive process, grounded in dynamic cue acquisition and continuous expertise accumulation. Yet most current artificial intelligence (AI) systems…
Semantic Energy: Detecting LLM Hallucination Beyond Entropy
Huan Ma, Jiadong Pan, Jing Liu +7
Large Language Models (LLMs) are being increasingly deployed in real-world applications, but they remain susceptible to hallucinations, which produce fluent yet incorrect responses…