collaborators

9 papers

cs.CL2026

BiCAA: Bidirectional Credit Assignment for Search-Augmented Agent

Yibin Huang, Bin Xu, Hailong Cao +1

Multi-step search is a fundamental capability for search agents, enabling them to iteratively acquire, refine, and integrate external evidence for complex reasoning QA. However, va…

cs.CL2026

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

Hui Huang, Yancheng He, Wei Liu +7

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models i…

cs.CL2026

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

Hongli Zhou, Hui Huang, Wei Liu +8

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-quality human preference annotation…

cs.CL2026

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

Hongli Zhou, Hui Huang, Ziqing Zhao +10

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concern…

cs.LG2025

ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning

Zihao Feng, Xiaoxue Wang, Bowen Wu +4

While reinforcement learning (RL) is increasingly used for LLM-based tool learning, its efficiency is often hampered by an overabundance of simple samples that provide diminishing…

cs.CL2025

Cross-Domain Bilingual Lexicon Induction via Pretrained Language Models

Qiuyu Ding, Zhiqiang Cao, Hailong Cao +1

Bilingual Lexicon Induction (BLI) is generally based on common domain data to obtain monolingual word embedding, and by aligning the monolingual word embeddings to obtain the cross…