collaborators

5 papers

cs.CV2025

OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning

Yuan Gong, Xionghui Wang, Jie Wu +3

In this paper, we introduce OneReward, a unified reinforcement learning framework that enhances the model's generative capabilities across multiple tasks under different evaluation…

cs.MM2025

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

Edson Araujo, Andrew Rouditchenko, Yuan Gong +7

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations t…

eess.AS2024

State-Space Large Audio Language Models

Saurabhchand Bhati, Yuan Gong, Leonid Karlinsky +3

Large Audio Language Models (LALM) combine the audio perception models and the Large Language Models (LLM) and show a remarkable ability to reason about the input audio, infer the…

eess.AS2024

A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation

Alexander H. Liu, Qirui Wang, Yuan Gong +1

Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence…

cs.CV2024

PLMM: Personal Large Language Models on Mobile Devices

Yuanhao Gong

Inspired by Federated Learning, in this paper, we propose personal large models that are distilled from traditional large language models but more adaptive to local users' personal…