collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Ming Li, Chenguang Wang, Xirui Li +5

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when…

cs.CL2026

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions

Chenrui Fan, Yize Cheng, Ming Li +3

Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple…

cs.CL2026

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

Han Chen, Ming Li, Hong Jiao +1

Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typical…

cs.CL2026

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

Chenguang Wang, Ming Li, Xinyue Zeng +4

Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on c…

cs.CL2026

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

Han Chen, Ming Li, Chenguang Wang +4

Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item meaningfully distinguishes higher-…

cs.CL2026

When is Your LLM Steerable?

Chenrui Fan, Yize Cheng, Ming Li +2

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, m…