4 papers
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
Tianjun Pan, Xuan Lin, Wenyan Yang +7
Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these…
Softmax Linear Attention: Reclaiming Global Competition
Mingwei Xu, Xuan Lin, Xinnan Guo +2
While linear attention reduces the quadratic complexity of standard Transformers to linear time, it often lags behind in expressivity due to the removal of softmax normalization. T…
AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
Xuan Lin, Long Chen, Yile Wang
Large Language Models (LLMs) have shown promise in assisting molecular property prediction tasks but often rely on human-crafted prompts and chain-of-thought templates. While recen…
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
Shisong Chen, Qian Zhu, Wenyan Yang +15
Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capa…