most citedEfficient Reinforcement Finetuning via Adaptive Curriculum Learning

1 citations · 1 across the 1 of their papers we have counts for

collaborators

9 papers

cs.LG20261 cited

Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Taiwei Shi, Yiyang Wu, Linxin Song +2

Reinforcement finetuning (RFT) has shown great potential for enhancing the mathematical reasoning capabilities of large language models (LLMs), but it is often sample- and compute-…

cs.CL2026

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

Meihua Dang, Linxin Song, Honghua Zhang +3

Generations from large language models often fail to conform to desired constraints such as JSON schema. Existing locally constrained decoding (LCD) approaches enforce constraints…

cs.LG2026

Skill Reuse as Compression in Agentic RL

Zhikun Xu, Yu Feng, Jacob Dineen +3

Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successfu…

cs.CY2026

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

Saleh Afroogh, Syed Ishtiaque Ahmed, Petra Ahrweiler +46

This study provides a cross-disciplinary examination of Explainable Artificial Intelligence (XAI) approaches-focusing on deep neural networks (DNNs) and large language models (LLMs…

cs.CV2026

Video-Based Reward Modeling for Computer-Use Agents

Linxin Song, Jieyu Zhang, Huanxin Sheng +6

Computer-using agents (CUAs) are becoming increasingly capable; however, it remains difficult to scale evaluation of whether a trajectory truly fulfills a user instruction. In this…

cs.AI2026

MED-COPILOT: A Medical Assistant Powered by GraphRAG and Similar Patient Case Retrieval

Shuheng Chen, Namratha Patil, Haonan Pan +4

Clinical decision-making requires synthesizing heterogeneous evidence, including patient histories, clinical guidelines, and trajectories of comparable cases. While large language…