Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning
Chang Tian, Matthew B. Blaschko, Mingzhe Xing +3
Reinforcement learning (RL) has become a key technique for enhancing the reasoning abilities of large language models (LLMs), with policy-gradient algorithms dominating the post-tr…
cs.AI2025
A Generic Method for Fine-grained Category Discovery in Natural Language Texts
Chang Tian, Matthew B. Blaschko, Wenpeng Yin +3
Fine-grained category discovery using only coarse-grained supervision is a cost-effective yet challenging task. Previous training methods focus on aligning query samples with posit…