From the 1 of 4 linked papers with an AI index.
4 papers
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim +8
The paper proposes using audio-aware large language models to give fine‑grained feedback on text‑to‑audio generation, improving how well the generated audio follows multi‑event and…
Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
Jiajun Fan, Roger Ren, Jingyuan Li +6
The role of reasoning in Audio Large Language Models remains widely underexplored, as introducing a reasoning process often degrades rather than improves performance during inferen…
Group Relative Policy Optimization for Speech Recognition
Prashanth Gurunath Shivakumar, Yile Gu, Ankur Gandhe +1
Speech Recognition has seen a dramatic shift towards adopting Large Language Models (LLMs). This shift is partly driven by good scalability properties demonstrated by LLMs, ability…
Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
Guan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav +4
While textless Spoken Language Models (SLMs) have shown potential in end-to-end speech-to-speech modeling, they still lag behind text-based Large Language Models (LLMs) in terms of…