activity
20212026
most citedTextBox: A Unified, Modularized, and Extensible Framework for Text Generation

4 citations · 9 across the 12 of their papers we have counts for

collaborators

13 papers

cs.CL2026

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics

Mengyuan Sun, Yu Li, Zhuohao Yu +2

Rubric-based evaluation is a promising paradigm for judging large language model (LLM) outputs, yet self-generated rubrics lag human-annotated criteria on hard instances. We argue…

cs.LG2026

From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism

Zhuohao Yu, Zhiwei Steven Wu, Adam Block

Inference-time compute scaling has emerged as a powerful paradigm for improving language model performance on a wide range of tasks, but the question of how best to use the additio…

cs.CL2026

SteerRM: Debiasing Reward Models via Sparse Autoencoders

Mengyuan Sun, Zhuohao Yu, Weizheng Gu +2

Reward models (RMs) are critical components of alignment pipelines, yet they exhibit biases toward superficial stylistic cues, preferring better-presented responses over semantical…

cs.LG2026

What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?

Weizheng Gu, Chengze Li, Zhuohao Yu +6

Large language models are increasingly evaluated as interactive agents, yet standard agent benchmarks conflate two qualitatively distinct sources of success: semantic tool-use and…

cs.LG2026

Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning

Jiahao Zhang, Lujing Zhang, Keltin Grimes +3

A recurring challenge in preference fine-tuning (PFT) is handling (i.e., cyclic) preferences. Intransitive preferences often stem from either

cs.AI20251 cited

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

Yidong Wang, Yunze Song, Tingyuan Zhu +11

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundam…