4 papers
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
Renren Jin, Pengzhi Gao, Yuqi Ren +6
Reinforcement learning with verifiable rewards (RLVR) has emerged as a prominent paradigm for enhancing the reasoning capabilities of large language models (LLMs). However, the ent…
TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation
Renren Jin, Tianhao Shen, Xinwei Wu +9
Conducting supervised and preference fine-tuning of large language models (LLMs) requires high-quality datasets to improve their ability to follow instructions and align with human…
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
Wuwei Huang, Dexin Wang, Deyi Xiong
In end-to-end speech translation, acoustic representations learned by the encoder are usually fixed and static, from the perspective of the decoder, which is not desirable for deal…
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
Wuwei Huang, Renren Jin, Wen Zhang +3
Recent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end…