From the 1 of 63 linked papers with an AI index.
43 papers · 1 filter
Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning
Kunbin Xu, Xingzuo Li, Xuefeng Bai +1
Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed thro…
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
Hongbin Zhang, Junhao Liu, Xuefeng Bai +3
The paper introduces DualAnchor, a training framework for gloss-free sign language translation that preserves large language model priors with token-level prior anchoring and enhan…
Agentic Tool Use in Large Language Models
Jinchao Hu, Meizhi Zhong, Kehai Chen +2
Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for information retrieval, computation and e…
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
Mufan Xu, Kehai Chen, Jiahao Hu +4
Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-turn interactions while leve…
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents
Yihong Tang, Kehai Chen, Liang Yue +2
Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning capabilities of Large Language…
Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
Yu Zhang, Mufan Xu, Xuefeng Bai +4
Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large langua…