collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

BiCAA: Bidirectional Credit Assignment for Search-Augmented Agent

Yibin Huang, Bin Xu, Hailong Cao +1

Multi-step search is a fundamental capability for search agents, enabling them to iteratively acquire, refine, and integrate external evidence for complex reasoning QA. However, va…

cs.CL2026

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Yangda Peng, Yunjia Qi, Hao Peng +11

Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliability of LaaJ for rubric sco…

cs.CL2026

StoryAlign: Evaluating and Training Reward Models for Story Generation

Haotian Xia, Hao Peng, Yunjia Qi +4

Story generation aims to automatically produce coherent, structured, and engaging narratives. Although large language models (LLMs) have significantly advanced text generation, sto…

cs.CL2026

On the Paradoxical Interference between Instruction-Following and Task Solving

Yunjia Qi, Hao Peng, Xintong Shi +5

Instruction following aims to align Large Language Models (LLMs) with human intent by specifying explicit constraints on how tasks should be performed. However, we reveal a counter…

cs.CL2025

Evaluating Hydro-Science and Engineering Knowledge of Large Language Models

Shiruo Hu, Wenbo Shan, Yingjia Li +16

Hydro-Science and Engineering (Hydro-SE) is a critical and irreplaceable domain that secures human water supply, generates clean hydropower energy, and mitigates flood and drought…

cs.CL2025

WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection

Guanzhong He, Zhen Yang, Jinxin Liu +3

Search agents have achieved significant advancements in enabling intelligent information retrieval and decision-making within interactive environments. Although reinforcement learn…