papers

Publications (35)

cs.SD2024

Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation

Chien-Chun Wang, Li-Wei Chen, Hung-Shin Lee +2

Cross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch…

q-fin.ST2026

Generalized Stock Price Prediction for Multiple Stocks Combined with News Fusion

Pei-Jun Liao, Hung-Shin Lee, Yao-Fei Cheng +3

Predicting stock prices presents challenges in financial forecasting. While traditional approaches such as ARIMA and RNNs are prevalent, recent developments in Large Language Model…

cs.SD2024

VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka

Li-Wei Chen, Hung-Shin Lee, Chen-Chi Chang

This paper introduces VoxHakka, a text-to-speech (TTS) system designed for Taiwanese Hakka, a critically under-resourced language spoken in Taiwan. Leveraging the YourTTS framework…

cs.CV2024

BackFlip: The Impact of Local and Global Data Augmentations on Artistic Image Aesthetic Assessment

Ombretta Strafforello, Gonzalo Muradas Odriozola, Fatemeh Behrad +4

Assessing the aesthetic quality of artistic images presents unique challenges due to the subjective nature of aesthetics and the complex visual characteristics inherent to artworks…

cs.CL2023

Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings

Ta-Chung Chi, Ting-Han Fan, Li-Wei Chen +2

The use of positional embeddings in transformer language models is widely accepted. However, recent research has called into question the necessity of such embeddings. We further e…

eess.AS2022

A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units

Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody…

cs.CL2025

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

Li-Wei Chen, Takuya Higuchi, Zakaria Aldeneh +2

The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized f…

cs.CV2025

LAPIS: A novel dataset for personalized image aesthetic assessment

Anne-Sofie Maerten, Li-Wei Chen, Stefanie De Winter +2

We present the Leuven Art Personalized Image Set (LAPIS), a novel dataset for personalized image aesthetic assessment (PIAA). It is the first dataset with images of artworks that i…

cs.CV2025

On the Role of Individual Differences in Current Approaches to Computational Image Aesthetics

Li-Wei Chen, Ombretta Strafforello, Anne-Sofie Maerten +2

Image aesthetic assessment (IAA) evaluates image aesthetics, a task complicated by image diversity and user subjectivity. Current approaches address this in two stages: Generic IAA…

eess.AS2025

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

Li-Wei Chen, Takuya Higuchi, He Bai +6

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use…

eess.AS2025

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

Zakaria Aldeneh, Takuya Higuchi, Jee-weon Jung +6

Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to b…

cs.CL2025

Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages

Yao-Fei Cheng, Li-Wei Chen, Hung-Shin Lee +1

This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis an…

physics.flu-dyn2022

Towards high-accuracy deep learning inference of compressible turbulent flows over aerofoils

Li-Wei Chen, Nils Thuerey

The present study investigates the accurate inference of Reynolds-averaged Navier-Stokes solutions for the compressible flow over aerofoils in two dimensions with a deep neural net…

physics.app-ph2025

Mixed-Mode In-Memory Computing: Towards High-Performance Logic Processing In A Memristive Crossbar Array

Nan Du, Ilia Polian, Christopher Bengel +9

In-memory computing is a promising alternative to traditional computer designs, as it helps overcome performance limits caused by the separation of memory and processing units. How…

eess.AS2026

TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition

Cheng-Yeh Yang, Chien-Chun Wang, Li-Wei Chen +3

Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. Whil…

eess.AS2023

A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous Speech

Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

Recent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness. The diversity of human speech, however, often goes beyond the co…

cs.SD2025

Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition

Chien-Chun Wang, Li-Wei Chen, Cheng-Kang Chou +3

While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mi…

cs.LG2023

Learning Similarity Metrics for Volumetric Simulations with Multiscale CNNs

Georg Kohl, Li-Wei Chen, Nils Thuerey

Simulations that produce three-dimensional data are ubiquitous in science, ranging from fluid flows to plasma physics. We propose a similarity model based on entropy, which allows…

eess.AS2023

Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition

Li-Wei Chen, Alexander Rudnicky

While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using diff…

cs.LG2025

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective

Hyung Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko +5

We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing dist…

cs.CV2023

Spectral Analysis for Semantic Segmentation with Applications on Feature Truncation and Weak Annotation

Li-Wei Chen, Wei-Chen Chiu, Chin-Tien Wu

It is well known that semantic segmentation neural networks (SSNNs) produce dense segmentation maps to resolve the objects' boundaries while restrict the prediction on down-sampled…

eess.AS2022

Fine-grained style control in Transformer-based Text-to-speech Synthesis

Li-Wei Chen, Alexander Rudnicky

In this paper, we present a novel architecture to realize fine-grained style control on the transformer-based text-to-speech synthesis (TransformerTTS). Specifically, we model the…

physics.flu-dyn2021

Numerical investigation of minimum drag profiles in laminar flow using deep learning surrogates

Li-Wei Chen, Berkay Alp Cakal, Xiangyu Hu +1

Efficiently predicting the flowfield and load in aerodynamic shape optimisation remains a highly challenging and relevant task. Deep learning methods have been of particular intere…

eess.AS2019

Generative Adversarial Networks for Unpaired Voice Transformation on Impaired Speech

Li-Wei Chen, Hung-Yi Lee, Yu Tsao

This paper focuses on using voice conversion (VC) to improve the speech intelligibility of surgical patients who have had parts of their articulators removed. Due to the difficulty…

cs.CR2023

Power-balanced Memristive Cryptographic Implementation Against Side Channel Attacks

Ziang Chen, Li-Wei Chen, Xianyue Zhao +4

Memristors, as emerging nano-devices, offer promising performance and exhibit rich electrical dynamic behavior. Having already found success in applications such as neuromorphic an…

cs.CL2023

The North System for Formosa Speech Recognition Challenge 2023

Li-Wei Chen, Kai-Chen Cheng, Hung-Shin Lee

This report provides a concise overview of the proposed North system, which aims to achieve automatic word/syllable recognition for Taiwanese Hakka (Sixian). The report outlines th…

cs.SD2022

Speech Representation Learning Combining Conformer CPC with Deep Cluster for the ZeroSpeech Challenge 2021

Takashi Maekaku, Xuankai Chang, Yuya Fujita +3

We present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-la…

eess.IV2026

SCENE: Semantic-aware Codec Enhancement with Neural Embeddings

Han-Yu Lin, Li-Wei Chen, Hung-Shin Lee

Compression artifacts from standard video codecs often degrade perceptual quality. We propose a lightweight, semantic-aware pre-processing framework that enhances perceptual fideli…

physics.flu-dyn2024

Deep learning-based predictive modelling of transonic flow over an aerofoil

Li-Wei Chen, Nils Thuerey

Effectively predicting transonic unsteady flow over an aerofoil poses inherent challenges. In this study, we harness the power of deep neural network (DNN) models using the attenti…

cs.LG2024

Benchmarking Autoregressive Conditional Diffusion Models for Turbulent Flow Simulation

Georg Kohl, Li-Wei Chen, Nils Thuerey

Simulating turbulent flows is crucial for a wide range of applications, and machine learning-based solvers are gaining increasing relevance. However, achieving temporal stability w…

physics.flu-dyn2022

Learned Turbulence Modelling with Differentiable Fluid Solvers: Physics-based Loss-functions and Optimisation Horizons

Björn List, Li-Wei Chen, Nils Thuerey

In this paper, we train turbulence models based on convolutional neural networks. These learned turbulence models improve under-resolved low resolution solutions to the incompressi…

cs.SD2025

Revealing the Role of Audio Channels in ASR Performance Degradation

Kuan-Tang Huang, Li-Wei Chen, Hung-Shin Lee +2

Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the inpu…

physics.comp-ph2024

Differentiability in Unrolled Training of Neural Physics Simulators on Transient Dynamics

Bjoern List, Li-Wei Chen, Kartik Bali +1

Unrolling training trajectories over time strongly influences the inference accuracy of neural network-augmented physics simulators. We analyze this in three variants of training n…

cs.SD2023

A Training and Inference Strategy Using Noisy and Enhanced Speech as Target for Speech Enhancement without Clean Speech

Li-Wei Chen, Yao-Fei Cheng, Hung-Shin Lee +2

The lack of clean speech is a practical challenge to the development of speech enhancement systems, which means that there is an inevitable mismatch between their training criterio…

eess.AS2023

Speaker-Independent Acoustic-to-Articulatory Speech Inversion

Peter Wu, Li-Wei Chen, Cheol Jun Cho +4

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interp…