13 papers
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Ming Li, Chenguang Wang, Xirui Li +5
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when…
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
Chenrui Fan, Yize Cheng, Ming Li +3
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple…
Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling
Han Chen, Ming Li, Hong Jiao +1
Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typical…
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Chenguang Wang, Ming Li, Xinyue Zeng +4
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on c…
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
Qitong Wang, Yijun Liang, Ming Li +2
Vision-Language Navigation (VLN) enables robots to follow natural-language instructions in visually grounded environments, serving as a key capability for embodied robotic systems.…
Sharp Pre-Schwarzian Norm Bounds for Ma-Minda Starlike Classes
Ming Li, Mei Luo
In this paper, we develop a unified framework to evaluate the pre-Schwarzian norm for the Ma-Minda starlike class. We present a direct, general computational approach. As an applic…