5 papers
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
Yuan Gong, Xionghui Wang, Jie Wu +3
In this paper, we introduce OneReward, a unified reinforcement learning framework that enhances the model's generative capabilities across multiple tasks under different evaluation…
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
Edson Araujo, Andrew Rouditchenko, Yuan Gong +7
Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations t…
State-Space Large Audio Language Models
Saurabhchand Bhati, Yuan Gong, Leonid Karlinsky +3
Large Audio Language Models (LALM) combine the audio perception models and the Large Language Models (LLM) and show a remarkable ability to reason about the input audio, infer the…
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
Alexander H. Liu, Qirui Wang, Yuan Gong +1
Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence…
PLMM: Personal Large Language Models on Mobile Devices
Yuanhao Gong
Inspired by Federated Learning, in this paper, we propose personal large models that are distilled from traditional large language models but more adaptive to local users' personal…