3 papers
cs.SD2025
HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
Rongkun Xue, Yazhe Niu, Shuai Hu +3
Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantiz…
cs.CV2025
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Yufei Zhan, Ziheng Wu, Yousong Zhu +10
Despite notable advancements in multimodal reasoning, leading Multimodal Large Language Models (MLLMs) still underperform on vision-centric multimodal reasoning tasks in general sc…
cs.LG2024
Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
Jinouwen Zhang, Rongkun Xue, Yazhe Niu +4
Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement le…