13 papers
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
Fengxian Ji, Yuke Li, Jingpu Yang +8
However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this que…
ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
Jie Gong, Maowei Jiang, Zhiwei Liu +14
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily asses…
LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents
Jingpu Yang, Fengxian Ji, Zhengzhao Lai +8
Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challeng…
Overview of the ClinicalSkillQA 2026 Shared Task on Continuous Perception and Procedural Reasoning in Clinical Skill Assessment
Xiyang Huang, Renxiong Wei, Yihuai Xu +10
This paper presents an overview of the ClinicalSkillQA 2026 shared task, which was organized with the BioNLP Workshop at ACL 2026. The goal of this shared task is to evaluate conti…
Credibility Governance: A Social Mechanism for Collective Self-Correction under Weak Truth Signals
Wanying He, Yanxi Lin, Ziheng Zhou +5
Online platforms increasingly rely on opinion aggregation to allocate real-world attention and resources, yet common signals such as engagement votes or capital-weighted commitment…
MisSpans: Fine-Grained False Span Identification in Cross-Domain Fake News
Zhiwei Liu, Paul Thompson, Jiaqi Rong +5
Online misinformation is increasingly pervasive, yet most existing benchmarks and methods evaluate veracity at the level of whole claims or paragraphs using coarse binary labels, o…