works on

From the 1 of 63 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CLShow all

43 papers · 1 filter

cs.CL2026

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

Kunbin Xu, Xingzuo Li, Xuefeng Bai +1

Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed thro…

cs.CL2026

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

Hongbin Zhang, Junhao Liu, Xuefeng Bai +3

The paper introduces DualAnchor, a training framework for gloss-free sign language translation that preserves large language model priors with token-level prior anchoring and enhan…

cs.CL2026

Agentic Tool Use in Large Language Models

Jinchao Hu, Meizhi Zhong, Kehai Chen +2

Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for information retrieval, computation and e…

cs.CL2026

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

Mufan Xu, Kehai Chen, Jiahao Hu +4

Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-turn interactions while leve…

cs.CL2026

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

Yihong Tang, Kehai Chen, Liang Yue +2

Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning capabilities of Large Language…

cs.CL2026

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

Yu Zhang, Mufan Xu, Xuefeng Bai +4

Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large langua…