papers

Publications (33)

cs.CV2025

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

Yuanrui Wang, Cong Han, Yafei Li +8

Text-to-image generation has greatly advanced content creation, yet accurately rendering visual text remains a key challenge due to blurred glyphs, semantic drift, and limited styl…

eess.AS2023

HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

Yinghao Aaron Li, Cong Han, Xilin Jiang +1

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these network…

cs.CL2021

Improving Conversational Recommendation Systems' Quality with Context-Aware Item Meta Information

Bowen Yang, Cong Han, Yu Li +2

Conversational recommendation systems (CRS) engage with users by inferring user preferences from dialog history, providing accurate recommendations, and generating appropriate resp…

eess.AS2020

Real-time binaural speech separation with preserved spatial cues

Cong Han, Yi Luo, Nima Mesgarani

Deep learning speech separation algorithms have achieved great success in improving the quality and intelligibility of separated speech from mixed audio. Most previous methods focu…

cs.CL2023

Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions

Yinghao Aaron Li, Cong Han, Xilin Jiang +1

Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic pros…

cs.MM2025

LongCat-Flash-Omni Technical Report

Meituan LongCat Team, Bairui Wang, Bayan +129

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curricu…

eess.AS2022

StyleTTS-VC: One-Shot Voice Conversion by Knowledge Transfer from Style-Based TTS Models

Yinghao Aaron Li, Cong Han, Nima Mesgarani

One-shot voice conversion (VC) aims to convert speech from any source speaker to an arbitrary target speaker with only a few seconds of reference speech from the target speaker. Th…

cs.LG2022

Extensible Proxy for Efficient NAS

Yuhong Li, Jiajie Li, Cong Han +3

Neural Architecture Search (NAS) has become a de facto approach in the recent trend of AutoML to design deep neural networks (DNNs). Efficient or near-zero-cost NAS proxies are fur…

eess.AS2021

Group Communication with Context Codec for Lightweight Source Separation

Yi Luo, Cong Han, Nima Mesgarani

Despite the recent progress on neural network architectures for speech separation, the balance between the model size, model complexity and model performance is still an important…

cs.GT2020

Incentive Mechanism Design for ROI-constrained Auto-bidding

Bin Li, Xiao Yang, Daren Sun +4

Auto-bidding plays an important role in online advertising and has become a crucial tool for advertisers and advertising platforms to meet their performance objectives and optimize…

cs.CV2026

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Vorch Team, Xiaoyu Chen, Yang Ding +25

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…

cs.CV2026

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

Cong Han, Xiaohan Lan, Haibo Qiu +1

Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pivotal research frontier. The…

eess.AS2022

Multi-Channel Speech Denoising for Machine Ears

Cong Han, E. Merve Kaya, Kyle Hoefer +2

This work describes a speech denoising system for machine ears that aims to improve speech intelligibility and the overall listening experience in noisy environments. We recorded a…

q-bio.NC2026

Power-Law Scaling in the Classification Performance of Small-Scale Spiking Neural Networks

Zhengdi Zhang, Cong Han, Wenjun Xia

This paper investigates the classification capability of small-scale spiking neural networks based on the Leaky Integrate-and-Fire (LIF) neuron model. We analyze the relationship b…

eess.AS2023

StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Yinghao Aaron Li, Cong Han, Nima Mesgarani

Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturali…

eess.AS2023

Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network

Cong Han, Nima Mesgarani

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT)…

eess.AS2020

Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording

Cong Han, Yi Luo, Chenda Li +8

Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by us…

eess.AS2021

Dual-Path Modeling for Long Recording Speech Separation in Meetings

Chenda Li, Zhuo Chen, Yi Luo +6

The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straight…

eess.AS2024

StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Yinghao Aaron Li, Xilin Jiang, Cong Han +1

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often fac…

eess.AS2023

Improved Decoding of Attentional Selection in Multi-Talker Environments with Self-Supervised Learned Speech Representation

Cong Han, Vishal Choudhari, Yinghao Aaron Li +1

Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listene…

eess.AS2025

Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience

Xilin Jiang, Cong Han, Yinghao Aaron Li +1

In daily life, we encounter a variety of sounds, both desirable and undesirable, with limited control over their presence and volume. Our work introduces "Listen, Chat, and Remix"…

eess.AS2023

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Yinghao Aaron Li, Cong Han, Vinay S. Raghavan +2

In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human…

cs.SD2024

Unsupervised Multi-channel Separation and Adaptation

Cong Han, Kevin Wilson, Scott Wisdom +1

A key challenge in machine learning is to generalize from training data to an application domain of interest. This work generalizes the recently-proposed mixture invariant training…

eess.AS2023

SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs

Yinghao Aaron Li, Cong Han, Nima Mesgarani

In recent years, large-scale pre-trained speech language models (SLMs) have demonstrated remarkable advancements in various generative speech modeling applications, such as text-to…

eess.AS2020

Ultra-Lightweight Speech Separation via Group Communication

Yi Luo, Cong Han, Nima Mesgarani

Model size and complexity remain the biggest challenges in the deployment of speech enhancement and separation systems on low-resource devices such as earphones and hearing aids. A…

cs.SD2026

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Xilin Jiang, Qiaolin Wang, Junkai Wu +30

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…

cs.CV2023

Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network

Cong Han, Yujie Zhong, Dengjie Li +2

Recently, the open-vocabulary semantic segmentation problem has attracted increasing attention and the best performing methods are based on two-stream networks: one stream for prop…

eess.AS2020

Rethinking the Separation Layers in Speech Separation Networks

Yi Luo, Zhuo Chen, Cong Han +3

Modules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules gene…

eess.AS2020

Distortion-controlled Training for End-to-end Reverberant Speech Separation with Auxiliary Autoencoding Loss

Yi Luo, Cong Han, Nima Mesgarani

The performance of speech enhancement and separation systems in anechoic environments has been significantly advanced with the recent progress in end-to-end neural network architec…

eess.AS2019

FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing

Yi Luo, Enea Ceolini, Cong Han +2

Beamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called \textit{neural beamformers},…

eess.AS2024

Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation

Xilin Jiang, Cong Han, Nima Mesgarani

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with q…

eess.AS2024

Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis

Xilin Jiang, Yinghao Aaron Li, Adrian Nicolas Florea +2

It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in mu…

eess.AS2023

Exploring Self-Supervised Contrastive Learning of Spatial Sound Event Representation

Xilin Jiang, Cong Han, Yinghao Aaron Li +1

In this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and…