collaborators

6 papers

cs.CY2026

Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests

Jin-Guo Liu, Shang-Qi Lu, Xin-Ran Shi +2

This paper is an experience report on a 13-week Test-Driven, AI-Assisted (TDAA) redesign of DSAA 3071, Theory of Computation, an upper-level course at the Hong Kong University of S…

cs.AI2026

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning

Liuji Chen, Dianxing Tang, Xing Shi +4

Agentic reinforcement learning can induce tool abuse, where models overuse external tools even for queries solvable by internal reasoning. Existing approaches mitigate this issue w…

cs.CL2026

Large Language Models Could Be Rote Learners

Yuyang Xu, Renjun Hu, Haochao Ying +3

Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…

cs.LG2025

SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression

Yuyang Xu, Yi Cheng, Haochao Ying +5

Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…

cs.IR2025

Behavior Modeling Space Reconstruction for E-Commerce Search

Yejing Wang, Chi Zhang, Xiangyu Zhao +8

Delivering superior search services is crucial for enhancing customer experience and driving revenue growth. Conventionally, search systems model user behaviors by combining user p…

cs.CL2025

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

Renjun Hu, Yi Cheng, Libin Meng +4

The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…