activity
20242026
collaborators

5 papers

cs.CL2026

Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents

Jiayi Wu, Ruobing Xie, Zeqian Huang +6

Search agents achieve strong question-answering performance through multi-turn interactions with search engines, with Group Relative Policy Optimization (GRPO) being a widely used…

cs.LG2026

UCS: Estimating Unseen Coverage for Improved In-Context Learning

Jiayi Xin, Xiang Li, Evan Qiang +4

In-context learning (ICL) performance depends critically on which demonstrations are placed in the prompt, yet most existing selectors prioritize heuristic notions of relevance or…

cs.CL2025

Let's Be Self-generated via Step by Step: A Curriculum Learning Approach to Automated Reasoning with Large Language Models

Kangyang Luo, Zichen Ding, Zhenmin Weng +5

While Chain of Thought (CoT) prompting approaches have significantly consolidated the reasoning capabilities of large language models (LLMs), they still face limitations that requi…

cs.CL2024

PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization

Jiayi Wu, Hengyi Cai, Lingyong Yan +5

The emergence of Retrieval-augmented generation (RAG) has alleviated the issues of outdated and hallucinatory content in the generation of large language models (LLMs), yet it stil…

cs.CL2024

Cross-model Control: Improving Multiple Large Language Models in One-time Training

Jiayi Wu, Hao Sun, Hengyi Cai +5

The number of large language models (LLMs) with varying parameter scales and vocabularies is increasing. While they deliver powerful performance, they also face a set of common opt…