papers

Publications (86)

eess.AS2021

OSSEM: one-shot speaker adaptive speech enhancement using meta learning

Cheng Yu, Szu-Wei Fu, Tsun-An Hsieh +2

Although deep learning (DL) has achieved notable progress in speech enhancement (SE), further research is still required for a DL-based SE system to adapt effectively and efficient…

math.AP2017

Global weak solution to the viscous two-fluid model with finite energy

Alexis Vasseur, Huanyao Wen, Cheng Yu

In this paper, we prove the existence of global weak solutions to the compressible two-fluid Navier-Stokes equations in three dimensional space. The pressure depends on two differe…

math.AP2013

Almost sure existence of Navier-Stokes Equations with randomized data in the whole space

Robin Ming Chen, Dehua Wang, Song Yao +1

This paper considers the supercritical Navier-Stokes equations posed in the whole space , with suitably randomized initial data, in the weak solution setting. The global weak…

cs.CV2026

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Tianhong Zhou, Mingyang Han, Boyu Li +8

Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…

math.AP2017

Energy conservation for the weak solutions of the compressible Navier-Stokes equations

Cheng Yu

In this paper, we prove the energy conservation for the weak solutions of the compressible Navier-Stokes equations for any time , under certain conditions. The results hold fo…

cs.CL2026

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

Zanlin Ni, Shenzhi Wang, Yang Yue +8

Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility i…

physics.acc-ph2026

Demonstration of High-Gain Harmonic Lasing in a Terahertz Free-Electron Laser

Yin Kang, Cheng Yu, Yue Wang +26

Compact Free-Electron Lasers (FELs) offering broad, continuous spectral tunability are traditionally constrained by fixed-parameter magnetic structures and the necessity for high-e…

cs.AI2026

Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem

Weixun Wang, XiaoXiao Xu, Wanhe An +86

Agentic crafting requires LLMs to operate in real-world environments over multiple turns by taking actions, observing outcomes, and iteratively refining artifacts. Despite its impo…

cs.RO2026

JODA: Composable Joint Dynamics for Articulated Objects

Tianhong Gao, Cheng Yu, Yinghao Xu +1

Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamical effects that govern realisti…

eess.AS2024

AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling

Vahid Ahmadi Kalkhorani, Cheng Yu, Anurag Kumar +3

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an audiovisual (AV) system for speech enhancement, target…

physics.acc-ph2025

First Lasing and Stable Operation of a Direct-Amplification Enabled Harmonic Generation Free-Electron laser

Zheng Qi, Junhao Liu, Lanpeng Ni +27

Seeded free-electron lasers (FELs) capable of operating at repetition rates up to the MHz level are in high demand for advanced time-resolved spectroscopies, which require both ful…

math.AP2016

A new proof to the energy conservation for the Navier-Stokes equations

Cheng Yu

In this paper we give a new proof to the energy conservation for the weak solutions of the incompressible Navier-Stokes equations. This result was first proved by Shinbrot. The new…

cs.SD2025

Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition

Juncheng Wang, Chao Xu, Cheng Yu +4

Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos. Following the perspective of extracting essential si…

cs.CV2025

Understanding Diffusion Models via Code Execution

Cheng Yu

Diffusion models have achieved remarkable performance in generative modeling, yet their theoretical foundations are often intricate, and the gap between mathematical formulations i…

math.AP2018

Onsager's energy conservation for inhomogeneous Euler equations

Robin Ming Chen, Cheng Yu

This paper addresses the problem of energy conservation for the two- and three-dimensional density-dependent Euler equations. Two types of sufficient conditions on the regularity o…

cs.SD2021

Improving Perceptual Quality by Phone-Fortified Perceptual Loss using Wasserstein Distance for Speech Enhancement

Tsun-An Hsieh, Cheng Yu, Szu-Wei Fu +2

Speech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information,…

math.AP2026

Finite-Energy Weak Solutions to the Quantum Isothermal Euler System via a Logarithmic Schrödinger Approximation

Cheng Yu

This paper investigates the collisionless quantum hydrodynamic, or quantum Euler, system in \(\mathbb{T}^3\) with the linear pressure law \(P(ρ)=ρ\). Since this pressure is assoc…

eess.AS2020

Time-Domain Multi-modal Bone/air Conducted Speech Enhancement

Cheng Yu, Kuo-Hsuan Hung, Syu-Siang Wang +3

Previous studies have proven that integrating video signals, as a complementary modality, can facilitate improved performance for speech enhancement (SE). However, video clips usua…

cs.GR2023

Simulating Parametric Thin Shells by Bicubic Hermite Elements

Xingyu Ni, Xuwen Chen, Cheng Yu +2

In this study, we present the bicubic Hermite element method (BHEM), a new computational framework devised for the elastodynamic simulation of parametric thin-shell structures. The…

eess.AS2023

Using fine-tuning and min lookahead beam search to improve Whisper

Andrea Do, Oscar Brown, Zhengjie Wang +4

The performance of Whisper in low-resource languages is still far from perfect. In addition to a lack of training data on low-resource languages, we identify some limitations in th…

math.AP2019

Global Existence of Entropy-Weak Solutions to the Compressible Navier-Stokes Equations with Non-Linear Density Dependent Viscosities

Didier Bresch, Alexis Vasseur, Cheng Yu

In this paper, we extend considerably the global existence results of entropy-weak solutions related to compressible Navier-Stokes system with density dependent viscosities obtaine…

eess.AS2019

Increasing Compactness Of Deep Learning Based Speech Enhancement Models With Parameter Pruning And Quantization Techniques

Jyun-Yi Wu, Cheng Yu, Szu-Wei Fu +3

Most recent studies on deep learning based speech enhancement (SE) focused on improving denoising performance. However, successful SE applications require striking a desirable bala…

eess.AS2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe +3

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models…

math.AP2011

Global weak solution and large-time behavior for the compressible flow of liquid crystals

Dehua Wang, Cheng Yu

The three-dimensional equations for the compressible flow of liquid crystals are considered. An initial-boundary value problem is studied in a bounded domain with large data. The e…

math.AP2018

Global weak solutions to compressible Navier-Stokes-Vlasov-Boltzmann systems for spray dynamics

Irene M. Gamba, Cheng Yu

This work concerns the global existence of the weak solutions to a system of partial differential equations modeling the evolution of particles in the fluid. That system is given b…

eess.AS2021

HASA-net: A non-intrusive hearing-aid speech assessment network

Hsin-Tien Chiang, Yi-Chiao Wu, Cheng Yu +4

Without the need of a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. Recently, deep neural network (DNN) models hav…

cs.CV2026

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Lianghua Huang, Zhi-Fan Wu, Wei Wang +22

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. W…

cs.SD2022

Speech Recovery for Real-World Self-powered Intermittent Devices

Yu-Chen Lin, Tsun-An Hsieh, Kuo-Hsuan Hung +4

The incompleteness of speech inputs severely degrades the performance of all the related speech signal processing applications. Although many researches have been proposed to addre…

cs.CV2026

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

Hanqing Yang, Qiang Zhou, Yongchao Du +6

Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planni…

eess.AS2020

Waveform-based Voice Activity Detection Exploiting Fully Convolutional networks with Multi-Branched Encoders

Cheng Yu, Kuo-Hsuan Hung, I-Fan Lin +3

In this study, we propose an encoder-decoder structured system with fully convolutional networks to implement voice activity detection (VAD) directly on the time-domain waveform. T…

cs.SD2024

Cross-Utterance Conditioned VAE for Speech Generation

Yang Li, Cheng Yu, Guangzhi Sun +8

Speech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In res…

cs.AI2026

AutoPKG: An Automated Framework for Dynamic E-commerce Product-Attribute Knowledge Graph Construction

Pollawat Hongwimol, Haoning Shang, Chutong Wang +6

Product attribute extraction in e-commerce is bottlenecked by ontologies that are inconsistent, incomplete, and costly to maintain. We present AutoPKG, a multi-agent Large Language…

math.AP2021

Dissipative solutions to the compressible isentropic Navier-Stokes equations

Liang Guo, Fucai Li, Cheng Yu

The existence of dissipative solutions to the compressible isentropic Navier-Stokes equations was established in this paper. This notion was inspired by the concept of dissipative…

cs.CL2024

Dynamic Depth Decoding: Faster Speculative Decoding for LLMs

Oscar Brown, Zhengjie Wang, Andrea Do +2

The acceleration of Large Language Models (LLMs) with speculative decoding provides a significant runtime improvement without any loss of accuracy. Currently, EAGLE-2 is the state-…

math.AP2021

Inviscid limit of the inhomogeneous incompressible Navier-Stokes equations under the weak Kolmogorov hypothesis in

Dixi Wang, Cheng Yu, Xinhua Zhao

In this paper, we consider the inviscid limit of inhomogeneous incompressible Navier-Stokes equations under the weak Kolmogorov hypothesis in . In particular, we firs…

cs.CV2020

Shaping Deep Feature Space towards Gaussian Mixture for Visual Classification

Weitao Wan, Jiansheng Chen, Cheng Yu +3

The softmax cross-entropy loss function has been widely used to train deep models for various tasks. In this work, we propose a Gaussian mixture (GM) loss function for deep neural…

cs.CV2024

FaceChain-FACT: Face Adapter with Decoupled Training for Identity-preserved Personalization

Cheng Yu, Haoyu Xie, Lei Shang +4

In the field of human-centric personalized image generation, the adapter-based method obtains the ability to customize and generate portraits by text-to-image training on facial da…

math.AP2026

Universality in the Low Mach number limit via a convex integration framework

Robin Ming Chen, Alexis Vasseur, Dehua Wang +1

We study the low Mach number limit of the compressible Euler equations through the lens of convex integration. For any prescribed weak solution of the incompressible Euler eq…

stat.ME2023

Matrix GARCH Model: Inference and Application

Cheng Yu, Dong Li, Feiyu Jiang +1

Matrix-variate time series data are largely available in applications. However, no attempt has been made to study their conditional heteroskedasticity that is often observed in eco…

math.AP2021

Global ill-posedness for a dense set of initial data to the Isentropic system of gas dynamics

Robin Ming Chen, Alexis F. Vasseur, Cheng Yu

In dimension and , we show that for any initial datum belonging to a dense subset of the energy space, there exist infinitely many global-in-time admissible weak solutions…

math.AP2015

Existence of Global Weak Solutions for 3D Degenerate Compressible Navier-Stokes Equations

Alexis F. Vasseur, Cheng Yu

In this paper, we prove the existence of global weak solutions for 3D compressible Navier-Stokes equations with degenerate viscosity. The method is based on the Bresch and Desjardi…

math.AP2015

Global weak solutions to compressible quantum Navier-Stokes equations with damping

Alexis F. Vasseur, Cheng Yu

The global-in-time existence of weak solutions to the barotropic compressible quantum Navier-Stokes equations with damping is proved for large data in three dimensional space. The…

cs.SD2021

SEOFP-NET: Compression and Acceleration of Deep Neural Networks for Speech Enhancement Using Sign-Exponent-Only Floating-Points

Yu-Chen Lin, Cheng Yu, Yi-Te Hsu +3

Numerous compression and acceleration strategies have achieved outstanding results on classification tasks in various fields, such as computer vision and speech signal processing.…

cs.CV2020

Multi-Scale Networks for 3D Human Pose Estimation with Inference Stage Optimization

Cheng Yu, Bo Wang, Bo Yang +1

Estimating 3D human poses from a monocular video is still a challenging task. Many existing methods' performance drops when the target person is occluded by other objects, or the m…

math.AP2013

Global weak solutions to the inhomogeneous Navier-Stokes-Vlasov equations

Dehua Wang, Cheng Yu

A fluid-particle system of the inhomogeneous Navier-Stokes equations and Vlasov equation in the three dimensional space is considered in this paper. The coupling arises from the dr…

cs.CV2026

Wan-Streamer v0.2: Higher Resolution, Same Latency

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23

We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…

cs.CY2026

Safety Degradation in AI Agents

Cheng Yu, Benedikt Stroebl, Diyi Yang +1

Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the integration of LLM…

eess.AS2020

Speech Enhancement based on Denoising Autoencoder with Multi-branched Encoders

Cheng Yu, Ryandhimas E. Zezario, Syu-Siang Wang +5

Deep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model gener…

math.AP2013

Global weak solution for a coupled compressible Navier-Stokes and Q-tensor system

Dehua Wang, Xiang Xu, Cheng Yu

In this paper, we study a coupled compressible Navier-Stokes/Q-tensor system modeling the nematic liquid crystal flow in a three-dimensional bounded spatial domain. The existence a…

physics.geo-ph2019

Modeling of fracture geometry alteration and fracture flow evolution under geostress and water-rock interaction

Cheng Yu

A coupled mech-hydro-chemical model for rock geometry alteration of fractures under water-rock interaction (WRI) and geostress is developed. Processes including WRI, asperity defor…

q-fin.PM2025

Tensor dynamic conditional correlation model: A new way to pursuit "Holy Grail of investing"

Cheng Yu, Zhoufan Zhu, Ke Zhu

Style investing creates asset classes (or the so-called "styles") with low correlations, aligning well with the principle of "Holy Grail of investing" in terms of portfolio selecti…

cs.CV2026

Artificial Intelligence for Detecting Fetal Orofacial Clefts and Advancing Medical Education

Yuanji Zhang, Yuhao Huang, Haoran Dou +28

Orofacial clefts are among the most common congenital craniofacial abnormalities, yet accurate prenatal detection remains challenging due to the scarcity of experienced specialists…

cs.SI2025

Evaluating and Improving Large Language Models for Competitive Program Generation

Minnan Wei, Ziming Li, Xiang Chen +5

Context: Due to the demand for strong algorithmic reasoning, complex logic implementation, and strict adherence to input/output formats and resource constraints, competitive progra…

math.AP2011

Incompressible limit for the compressible flow of liquid crystals

Dehua Wang, Cheng Yu

The connection between the compressible flow of liquid crystals with low Mach number and the incompressible flow of liquid crystals is studied in a bounded domain. In particular, t…

cs.CV2026

Video = World + Event Stream

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24

The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…

#video streaming#multimodal interaction#real-time AI#audio-visual modeling
physics.acc-ph2026

Fully coherent short wavelength free-electron laser driven by a single sub-microjoule seed

Lanpeng Ni, Zheng Qi, Xingtao Wang +33

High-repetition-rate, fully coherent extreme-ultraviolet (EUV) and X-ray free-electron lasers (FELs) are essential for advanced time-resolved ultrafast spectroscopies. While extern…

math.AP2024

Non-uniqueness for continuous solutions to 1D hyperbolic systems

Robin Ming Chen, Alexis F. Vasseur, Cheng Yu

In this paper, we show that a geometrical condition on systems of conservation laws leads to non-uniqueness in the class of 1D continuous functions. This demonstrates th…

cs.SD2023

Improving Speech Enhancement Performance by Leveraging Contextual Broad Phonetic Class Information

Yen-Ju Lu, Chia-Yu Chang, Cheng Yu +4

Previous studies have confirmed that by augmenting acoustic features with the place/manner of articulatory features, the speech enhancement (SE) process can be guided to consider t…

math.AP2012

Global weak solutions to the Navier-Stokes-Vlasov equations

Cheng Yu

In this paper, the system of particles coupled with fluid is considered. The particles are described by a Vlasov equation, and the fluid is governed by a forced Navier-Stokes equat…

cs.CV2023

FaceChain: A Playground for Human-centric Artificial Intelligence Generated Content

Yang Liu, Cheng Yu, Lei Shang +17

Recent advancement in personalized image generation have unveiled the intriguing capability of pre-trained text-to-image models on learning identity information from a collection o…

cs.CV2026

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

Sashuai Zhou, Qiang Zhou, Junpeng Ma +9

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most…

math.AP2018

The energy equality for the Navier-Stokes equations in bounded domains

Cheng Yu

In this paper, we provide a sufficient condition of the energy equality for the incompressible Navier-Stokes equations in bounded domains.

stat.ME2025

Two-way Matrix Autoregressive Model with Thresholds

Cheng Yu, Dong Li, Xinyu Zhang +1

Recently, matrix-valued time series data have attracted significant attention in the literature with the recognition of threshold nonlinearity representing a significant advance. H…

cs.SD2021

MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh +4

The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Ob…

math.AP2026

Inertial Limit of global weak solutions for Compressible Navier--Stokes

Cheng Yu

We investigate the inertial limit of the compressible Navier--Stokes system posed on the -dimensional torus, and allowing for regions of vacuum. Considering global-in-time finit…

cs.CV2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

Handong Li, Zikang Liu, Longteng Guo +10

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception throug…

physics.acc-ph2025

Enabling Continuous THz Band Coverage via Precise Electron Beam Tailoring in Free-electron Lasers

Yin Kang, Tong Li, Zhen Wang +28

High-power, continuously tunable narrowband terahertz (THz) sources are essential for advancing nonlinear optics, THz-driven material dynamics, and ultrafast spectroscopy. Conventi…

eess.AS2021

Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing

Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh +9

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our st…

cs.SD2026

FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing

Yuxuan Jiang, Mingyang Han, Yusheng Dai +12

Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge. However, existing methods struggle to bal…

eess.AS2021

Attention-based multi-task learning for speech-enhancement and speaker-identification in multi-speaker dialogue scenario

Chiang-Jen Peng, Yun-Ju Chan, Cheng Yu +3

Multi-task learning (MTL) and attention mechanism have been proven to effectively extract robust acoustic features for various speech-related tasks in noisy environments. In this s…

math.AP2020

Global existence of weak solutions to the Navier-Stokes equations with temperature-depending viscosity coefficient

Cheng Yu, Bijun Zuo

In this paper, the initial-boundary value problem to the three-dimensional inhomogeneous, incompressible and heat-conducting Navier-Stokes equations with temperature-depending visc…

stat.ME2025

Large covariance matrix estimation with factor-assisted variable clustering

Dong Li, Xinghao Qiao, Cheng Yu

This paper studies the covariance matrix estimation for high-dimensional time series within a new framework that combines low-rank factor and latent variable-specific cluster struc…

math.AP2023

Learning operators for identifying weak solutions to the Navier-Stokes equations

Dixi Wang, Cheng Yu

This paper focuses on investigating the learning operators for identifying weak solutions to the Navier-Stokes equations. Our objective is to establish a connection between the ini…

cs.CV2026

Unified Thinker: A General Reasoning Modular Core for Image Generation

Sashuai Zhou, Qiang Zhou, Jijin Hu +9

Despite impressive progress in high-fidelity image synthesis, generative models still struggle with logic-intensive instruction following, exposing a persistent reasoning--executio…

math.AP2017

Existence of global weak solutions for the Navier-Stokes-Vlasov-Boltzmann equations

Lei Yao, Cheng Yu

A moderately thick spray can be described by a coupled system of equations consisting of the incompressible Navier-Stokes equations and the Vlasov-Boltzmann equation. We investigat…

physics.acc-ph2026

Theoretical and experimental studies of energy modulation to demodulation in seeded free-electron lasers

Hanxiang Yang, Nanshun Huang, Zipeng Liu +9

Laser manipulation plays a critical role in precisely tailoring relativistic electron beams through energy modulation, enabling the generation of coherent, intense, and ultrashort…

math.AP2012

Global well-posedness for the two dimensional Navier-Stokes-Vlasov Equations

Cheng Yu

The global well-posedness for the incompressible Navier-Stokes-Vlasov equations in two spatial dimensions is established by a priori estimates, the characteristic method and the se…

cond-mat.supr-con2023

Peak Effects Induced by Particle Irradiations in 2H-NbSe2

Wenjie Li, Sunseng Pyon, Akiyoshi Yagi +5

Various peak effects in 2H-NbSe2 single crystals induced by particle irradiations were studied. 3 MeV proton irradiation magnified the peak effect induced by order-disorder transit…

eess.AS2020

HLT-NUS Submission for NIST 2019 Multimedia Speaker Recognition Evaluation

Rohan Kumar Das, Ruijie Tao, Jichen Yang +3

This work describes the speaker verification system developed by Human Language Technology Laboratory, National University of Singapore (HLT-NUS) for 2019 NIST Multimedia Speaker R…

cs.SD2025

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu +5

While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based mo…

cs.SD2021

MetricGAN-U: Unsupervised speech enhancement/ dereverberation based only on noisy/ reverberated speech

Szu-Wei Fu, Cheng Yu, Kuo-Hsuan Hung +2

Most of the deep learning-based speech enhancement models are learned in a supervised manner, which implies that pairs of noisy and clean speech are required during training. Conse…

cs.SD2022

Perceptual Contrast Stretching on Target Feature for Speech Enhancement

Rong Chao, Cheng Yu, Szu-Wei Fu +2

Speech enhancement (SE) performance has improved considerably owing to the use of deep learning models as a base function. Herein, we propose a perceptual contrast stretching (PCS)…

cs.SD2022

Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech

Yang Li, Cheng Yu, Guangzhi Sun +6

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE…

cs.CV2024

Improving generative adversarial network inversion via fine-tuning GAN encoders

Cheng Yu, Wenmin Wang, Roberto Bugiolacchi

Generative adversarial networks (GANs) can synthesize high-quality (HQ) images, and GAN inversion is a technique that discovers how to invert given images back to latent space. Whi…

math.AP2016

The weak solution to a Boltzmann type equation and its energy conservation

Cheng Yu

In this paper, we study the initial value problem of a Boltzmann type equation with a nonlinear degenerate damping. We prove the existence of global weak solutions with large initi…

cs.CY2026

How Should AI Safety Benchmarks Benchmark Safety?

Cheng Yu, Severin Engelmann, Ruoxuan Cao +2

AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210…