activity
20242026
collaborators

15 papers

cs.CL2026

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Ming Li, Chenguang Wang, Xirui Li +5

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when…

cs.CL2026

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

Chenguang Wang, Ming Li, Xinyue Zeng +4

Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on c…

cs.AI2026

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

Sajib Acharjee Dip, Dawei Zhou, Liqing Zhang

Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-corr…

cs.CL2026

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

Han Chen, Ming Li, Chenguang Wang +4

Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item meaningfully distinguishes higher-…

cs.AI2026

MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science

Junkai Zhang, Jingru Gan, Xiaoxuan Wang +8

Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To fill this gap, we introduce MatSc…

cs.AI2026

METASYMBO: Multi-Agent Language-Guided Metamaterial Discovery via Symbolic Latent Evolution

Jianpeng Chen, Wangzhi Zhan, Dongqi Fu +5

Metamaterial discovery seeks microstructured materials whose geometry induces targeted mechanical behavior. Existing inverse-design methods can efficiently generate candidates, but…