papers

Publications (296)

math.NA2025

Inverse source problem with a posteriori interior measurements for space-time fractional diffusion equations

Kai Yu, Zhiyuan Li, Yikan Liu

This paper investigates an inverse source problem for space-time fractional diffusion equations from a posteriori interior measurements. The uniqueness result is established by the…

cs.CL2018

On Modular Training of Neural Acoustics-to-Word Model for LVCSR

Zhehuai Chen, Qi Liu, Hao Li +1

End-to-end (E2E) automatic speech recognition (ASR) systems directly map acoustics to words using a unified model. Previous works mostly focus on E2E training a single model which…

cs.CL2020

CREDIT: Coarse-to-Fine Sequence Generation for Dialogue State Tracking

Zhi Chen, Lu Chen, Zihan Xu +3

In dialogue systems, a dialogue state tracker aims to accurately find a compact representation of the current dialogue status, based on the entire dialogue history. While previous…

eess.AS2026

Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

Hankun Wang, Haoran Wang, Yiwei Guo +3

Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coh…

cs.CL2022

META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Liangtai Sun, Xingyu Chen, Lu Chen +3

Task-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation. Current TOD…

cs.SD2023

Beyond the Status Quo: A Contemporary Survey of Advances and Challenges in Audio Captioning

Xuenan Xu, Zeyu Xie, Mengyue Wu +1

Automated audio captioning (AAC), a task that mimics human perception as well as innovatively links audio processing and natural language processing, has overseen much progress ove…

eess.AS2020

End-to-end spoofing detection with raw waveform CLDNNs

Heinrich Dinkel, Nanxin Chen, Yanmin Qian +1

Albeit recent progress in speaker verification generates powerful models, malicious attacks in the form of spoofed speech, are generally not coped with. Recent results in ASVSpoof2…

cs.CV2024

VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization

Tao Liu, Ziyang Ma, Qi Chen +4

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across d…

cs.CL2021

WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Xingyu Chen, Zihan Zhao, Lu Chen +5

Web search is an essential way for humans to obtain information, but it's still a great challenge for machines to understand the contents of web pages. In this paper, we introduce…

cs.AI2026

Towards High-Level Semantic Intelligence

Xiujie Song, Gefei Yang, Yining You +6

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear traje…

cs.CV2025

DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder

Chenpeng Du, Qi Chen, Tianyu He +5

While recent research has made significant progress in speech-driven talking face generation, the quality of the generated video still lags behind that of real recordings. One reas…

cs.CL2023

CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical Dataset

Hanchong Zhang, Jieyu Li, Lu Chen +5

The cross-domain text-to-SQL task aims to build a system that can parse user questions into SQL on complete unseen databases, and the single-domain text-to-SQL task evaluates the p…

quant-ph2022

Manifold Learning for Dimensionality Reduction: Quantum Isomap algorithm

WeiJun Feng, GongDe Guo, Kai Yu +2

Isomap algorithm is a representative manifold learning algorithm. The algorithm simplifies the data analysis process and is widely used in neuroimaging, spectral analysis and other…

cs.CL2024

Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model's Personality

Yiming Ai, Zhiwei He, Ziyin Zhang +5

In this study, we delve into the validity of conventional personality questionnaires in capturing the human-like personality traits of Large Language Models (LLMs). Our objective i…

cs.CL2025

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity

Da Ma, Lu Chen, Situo Zhang +8

The rapid expansion of context window sizes in Large Language Models~(LLMs) has enabled them to tackle increasingly complex tasks involving lengthy documents. However, this progres…

cs.SD2025

ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models

Bohan Li, Wenbin Huang, Yuhang Qiu +7

Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted int…

cs.CL2019

A Hierarchical Decoding Model For Spoken Language Understanding From Unaligned Data

Zijian Zhao, Su Zhu, Kai Yu

Spoken language understanding (SLU) systems can be trained on two types of labelled data: aligned or unaligned. Unaligned data do not require word by word annotation and is easier…

cs.AI2026

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation

Ziyue Yang, Da Ma, Hanqi Li +8

As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limit…

cs.CL2023

Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation

Zheng Liang, Zheshu Song, Ziyang Ma +3

Recently, end-to-end (E2E) automatic speech recognition (ASR) models have made great strides and exhibit excellent performance in general speech recognition. However, there remain…

cs.CL2025

Developing ChemDFM as a large language foundation model for chemistry

Zihan Zhao, Da Ma, Lu Chen +11

Artificial intelligence (AI) has played an increasingly important role in chemical research. However, most models currently used in chemistry are specialist models that require tra…

cs.SD2024

StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations

Sen Liu, Yiwei Guo, Xie Chen +1

While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of arti…

math.AP2026

Uniqueness for an inverse coefficient problem of a weakly coupled parabolic system

Caixuan Ren, Kai Yu, Zhiyuan Li

This paper considers the weakly coupled parabolic system with the homogeneous Neumann boundary condition, where \(P(x)\) is a \(2\times2\) sym…

cs.IT2024

In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models

Matteo Zecchin, Kai Yu, Osvaldo Simeone

Large pre-trained sequence models, such as transformer-based architectures, have been recently shown to have the capacity to carry out in-context learning (ICL). In ICL, a decision…

stat.ME2026

Using Importance Sampling to Estimate -values in All-Subset Meta-Analysis, with Applications to Single-Cell eQTL Mapping

Samuel Anyaso-Samuel, Thong Luong, Fei Qin +4

Pooling genome-wide association studies of multiple related traits can substantially increase power for detecting genetic variants with pleiotropic effects. ASSET, which exhaustive…

cs.CV2023

Reliable Federated Disentangling Network for Non-IID Domain Feature

Meng Wang, Kai Yu, Chun-Mei Feng +6

Federated learning (FL), as an effective decentralized distributed learning approach, enables multiple institutions to jointly train a model without sharing their local data. Howev…

cs.LG2026

Artificial Intelligence-Assistant Cardiotocography: Unified Model for Signal Reconstruction, Fetal Heart Rate Analysis, and Variability Assessment

Xiaohua Wang, Kai Yu, XuXiao Liang +2

The monitoring of fetal heart rate (FHR) and the assessment of its variability are crucial for preventing fetal compromise and adverse outcomes. However, traditional methods encoun…

cs.CL2019

Semantic Parsing with Dual Learning

Ruisheng Cao, Su Zhu, Chen Liu +2

Semantic parsing converts natural language queries into structured logical forms. The paucity of annotated training samples is a fundamental challenge in this field. In this work,…

eess.AS2025

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

Yifan Yang, Zheshu Song, Jianheng Zhuo +13

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, w…

cs.CL2020

Jointly Encoding Word Confusion Network and Dialogue Context with BERT for Spoken Language Understanding

Chen Liu, Su Zhu, Zijian Zhao +3

Spoken Language Understanding (SLU) converts hypotheses from automatic speech recognizer (ASR) into structured semantic representations. ASR recognition errors can severely degener…

eess.AS2025

What Does the Speaker Embedding Encode?

Shuai Wang, Yanmin Qian, Kai Yu

Developing a good speaker embedding has received tremendous interest in the speech community, with representations such as i-vector and d-vector demonstrating remarkable performanc…

cs.CL2025

Reducing Tool Hallucination via Reliability Alignment

Hongshen Xu, Zichen Zhu, Lei Pan +6

Large Language Models (LLMs) have expanded their capabilities beyond language generation to interact with external tools, enabling automation and real-world applications. However,…

eess.IV2025

Enhancing Diagnostic Accuracy in Rare and Common Fundus Diseases with a Knowledge-Rich Vision-Language Model

Meng Wang, Tian Lin, Aidi Lin +46

Previous foundation models for fundus images were pre-trained with limited disease categories and knowledge base. Here we introduce a knowledge-rich vision-language model (RetiZero…

eess.AS2024

VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature

Chenpeng Du, Yiwei Guo, Xie Chen +1

The mainstream neural text-to-speech(TTS) pipeline is a cascade system, including an acoustic model(AM) that predicts acoustic feature from the input transcript and a vocoder that…

cs.CV2025

Phased One-Step Adversarial Equilibrium for Video Diffusion Models

Jiaxiang Cheng, Bing Ma, Xuhua Ren +7

Video diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapte…

cs.LG2026

LightningRL: Breaking the Accuracy-Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning

Yanzhe Hu, Yijie Jin, Pengfei Liu +2

Diffusion Large Language Models (dLLMs) have emerged as a promising paradigm for parallel token generation, with block-wise variants garnering significant research interest. Despit…

cs.CL2025

MULTI: Multimodal Understanding Leaderboard with Text and Images

Zichen Zhu, Yang Xu, Lu Chen +11

The rapid development of multimodal large language models (MLLMs) raises the question of how they compare to human performance. While existing datasets often feature synthetic or o…

eess.AS2025

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding

Yu Xi, Haoyu Li, Xiaoyu Gu +2

Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explor…

cs.CL2026

Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

Ziyang Ma, Ruiyang Xu, Zhenghao Xing +9

Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), ca…

cs.CL2026

DiSRouter: Distributed Self-Routing for LLM Selections

Hang Zheng, Hongshen Xu, Yongkai Lin +3

The proliferation of Large Language Models (LLMs) has created a diverse ecosystem of models with highly varying performance and costs, necessitating effective query routing to bala…

cs.SD2024

DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation

Baihan Li, Zeyu Xie, Xuenan Xu +5

Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the l…

cs.LG2020

Text-based depression detection on sparse data

Heinrich Dinkel, Mengyue Wu, Kai Yu

Previous text-based depression detection is commonly based on large user-generated data. Sparse scenarios like clinical conversations are less investigated. This work proposes a te…

cs.CV2026

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +7

Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existi…

eess.AS2026

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

Wenxi Chen, Dongya Jia, Yushen Chen +11

Recently, diffusion models operating on VAE latents or mel-spectrograms have become the dominant paradigm for zero-shot TTS. Although these compressed representations improve gener…

cs.SD2023

DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech

Sen Liu, Yiwei Guo, Chenpeng Du +2

Although high-fidelity speech can be obtained for intralingual speech synthesis, cross-lingual text-to-speech (CTTS) is still far from satisfactory as it is difficult to accurately…

eess.AS2026

CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate

Hankun Wang, Yiwei Guo, Chongtian Shao +2

Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number o…

cs.AI2026

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

Zijian Wang, Hanqi Li, Ziyue Yang +17

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside…

eess.AS2024

VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching

Yiwei Guo, Chenpeng Du, Ziyang Ma +2

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms th…

q-bio.GN2025

MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging

Siyuan Li, Kai Yu, Anna Wang +7

Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit.…

cs.SD2022

BER: Balanced Error Rate For Speaker Diarization

Tao Liu, Kai Yu

DER is the primary metric to evaluate diarization performance while facing a dilemma: the errors in short utterances or segments tend to be overwhelmed by longer ones. Short segmen…

cs.AI2025

TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks

Hanwen Xu, Xuyao Huang, Yuzhe Liu +2

Large language model (LLM) agents have exhibited strong problem-solving competence across domains like research and coding. Yet, it remains underexplored whether LLM agents can tac…

stat.ME2025

Retrospective score tests versus prospective score tests for genetic association with case-control data

Yukun Liu, Pengfei Li, Lei Song +2

Since the seminal work by Prentice and Pyke (1979), the prospective logistic likelihood has become the standard method of analysis for retrospectively collected case-control data,…

cs.SD2021

Towards duration robust weakly supervised sound event detection

Heinrich Dinkel, Mengyue Wu, Kai Yu

Sound event detection (SED) is the task of tagging the absence or presence of audio events and their corresponding interval within a given audio clip. While SED can be done using s…

cs.CL2018

Sequence Discriminative Training for Deep Learning based Acoustic Keyword Spotting

Zhehuai Chen, Yanmin Qian, Kai Yu

Speech recognition is a sequence prediction problem. Besides employing various deep learning approaches for framelevel classification, sequence-level discriminative training has be…

cs.SD2024

Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis

Chenpeng Du, Kai Yu

Generating natural speech with a diverse and smooth prosody pattern is a challenging task. Although random sampling with phone-level prosody distribution has been investigated to g…

cs.CL2019

Data Augmentation with Atomic Templates for Spoken Language Understanding

Zijian Zhao, Su Zhu, Kai Yu

Spoken Language Understanding (SLU) converts user utterances into structured semantic representations. Data sparsity is one of the main obstacles of SLU due to the high cost of hum…

cs.CL2026

Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

Yushen Chen, Junzhe Liu, Yujie Tu +6

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, sc…

quant-ph2025

Quantum Graph Convolutional Networks Based on Spectral Methods

Zi Ye, Kai Yu, Song Lin

Graph Convolutional Networks (GCNs) are specialized neural networks for feature extraction from graph-structured data. In contrast to traditional convolutional networks, GCNs offer…

cs.SD2025

X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System

Zhanxun Liu, Yifan Duan, Mengmeng Wang +15

We present X-Talk, an open-source framework that champions a decoupled, modular design for LLM-driven speech-to-speech (S2S) systems. While the dominant trend favors end-to-end (E2…

eess.AS2025

VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining

Jianheng Zhuo, Yifan Yang, Yiwen Shao +4

Automatic speech recognition (ASR) has made remarkable progress but heavily relies on large-scale labeled data, which is scarce for low-resource languages like Vietnamese. While ex…

eess.AS2025

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Chenpeng Du, Yiwei Guo, Hankun Wang +6

Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation give…

cs.LG2023

Towards Instance-adaptive Inference for Federated Learning

Chun-Mei Feng, Kai Yu, Nian Liu +3

Federated learning (FL) is a distributed learning paradigm that enables multiple clients to learn a powerful global model by aggregating local training. However, the performance of…

cs.SE2025

Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis

Xueying Du, Kai Yu, Chong Wang +6

Static bug analyzers play a crucial role in ensuring software quality. However, existing analyzers for bug detection in large codebases often suffer from high false positive rates.…

cs.CL2025

When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models

Yingming Zheng, Hanqi Li, Kai Yu +1

Large language models (LLMs) have achieved impressive performance across natural language processing (NLP) tasks. As real-world applications increasingly demand longer context wind…

math.NA2026

Simultaneous recovery of multiple parameters in nonlocal diffusion equations from internal measurements

Kai Yu, Zhiyuan Li, Yikan Liu

This paper is devoted to simultaneously recovering multiple parameters from internal measurements for nonlocal diffusion equations. The uniqueness of the inverse problem is establi…

cs.LG2026

PaperGuide: Making Small Language-Model Paper-Reading Agents More Efficient

Zijian Wang, Tiancheng Huang, Hanqi Li +3

The accelerating growth of the scientific literature makes it increasingly difficult for researchers to track new advances through manual reading alone. Recent progress in large la…

cs.MA2026

Distributed Agent System: Fault-Tolerant Collaboration Among Embodied Agents

Kai Yu, Lu Chen, Hanqi Li

AI engineering is shifting from passive text generation by large language models (LLMs) to agent-driven task execution, creating new reliability challenges for long-horizon tasks u…

cs.LG2012

Collaborative Ensemble Learning: Combining Collaborative and Content-Based Information Filtering via Hierarchical Bayes

Kai Yu, Anton Schwaighofer, Volker Tresp +2

Collaborative filtering (CF) and content-based filtering (CBF) have widely been used in information filtering applications. Both approaches have their strengths and weaknesses whic…

cs.CL2021

Few-Shot NLU with Vector Projection Distance and Abstract Triangular CRF

Su Zhu, Lu Chen, Ruisheng Cao +3

Data sparsity problem is a key challenge of Natural Language Understanding (NLU), especially for a new target domain. By training an NLU model in source domains and applying the mo…

eess.AS2024

Text-aware Speech Separation for Multi-talker Keyword Spotting

Haoyu Li, Baochen Yang, Yu Xi +4

For noisy environments, ensuring the robustness of keyword spotting (KWS) systems is essential. While much research has focused on noisy KWS, less attention has been paid to multi-…

cs.CL2023

ASTormer: An AST Structure-aware Transformer Decoder for Text-to-SQL

Ruisheng Cao, Hanchong Zhang, Hongshen Xu +4

Text-to-SQL aims to generate an executable SQL program given the user utterance and the corresponding database schema. To ensure the well-formedness of output SQLs, one prominent a…

cs.CL2026

HeartAgent: An Autonomous Agent System for Explainable Differential Diagnosis in Cardiology

Shuang Zhou, Kai Yu, Song Wang +11

Heart diseases remain a leading cause of morbidity and mortality worldwide, necessitating accurate and trustworthy differential diagnosis. However, existing artificial intelligence…

eess.AS2024

SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs

Wenxi Chen, Ziyang Ma, Xiquan Li +5

Automated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs)…

cs.AI2024

AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures

Situo Zhang, Hankun Wang, Da Ma +4

Speculative Decoding (SD) is a popular lossless technique for accelerating the inference of Large Language Models (LLMs). We show that the decoding speed of SD frameworks with stat…

cs.CL2016

On Training Bi-directional Neural Network Language Model with Noise Contrastive Estimation

Tianxing He, Yu Zhang, Jasha Droppo +1

We propose to train bi-directional neural network language model(NNLM) with noise contrastive estimation(NCE). Experiments are conducted on a rescore task on the PTB data set. It i…

eess.AS2026

Dual-LoRA: Parameter-Efficient Adversarial Disentanglement for Cross-Lingual Speaker Verification

Qituan Shangguan, Junhao Du, Kunyang Peng +5

Cross-lingual speaker verification suffers from severe language-speaker entanglement. This causes systematic degradation in the hardest scenario: correctly accepting utterances fro…

cs.SD2026

SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

Ziyang Ma, Guanrou Yang, Wenxi Chen +19

The recent surge in open-source Multimodal Large Language Models (MLLM) frameworks, such as LLaVA, provides a convenient kickoff for artificial intelligence developers and research…

eess.AS2025

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability

Yuxiang Zhao, Yunchong Xiao, Yushen Chen +4

Recent advances in Text-To-Speech (TTS) technology have enabled synthetic speech to mimic human voices with remarkable realism, raising significant security concerns. This undersco…

cs.CL2022

TIE: Topological Information Enhanced Structural Reading Comprehension on Web Pages

Zihan Zhao, Lu Chen, Ruisheng Cao +3

Recently, the structural reading comprehension (SRC) task on web pages has attracted increasing research interests. Although previous SRC work has leveraged extra information such…

cs.LG2025

MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation

Yang Han, Pengyu Wang, Kai Yu +2

Mass spectrometry (MS) plays a critical role in molecular identification, significantly advancing scientific discovery. However, structure elucidation from MS data remains challeng…

cs.LG2025

Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations

Da Ma, Gonghu Shang, Zhi Chen +6

Instruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains chal…

physics.chem-ph2024

From Generalist to Specialist: A Survey of Large Language Models for Chemistry

Yang Han, Ziping Wan, Lu Chen +2

Large Language Models (LLMs) have significantly transformed our daily life and established a new paradigm in natural language processing (NLP). However, the predominant pretraining…

eess.AS2023

EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance

Yiwei Guo, Chenpeng Du, Xie Chen +1

Although current neural text-to-speech (TTS) models are able to generate high-quality speech, intensity controllable emotional TTS is still a challenging task. Most existing method…

eess.AS2020

Future Vector Enhanced LSTM Language Model for LVCSR

Qi Liu, Yanmin Qian, Kai Yu

Language models (LM) play an important role in large vocabulary continuous speech recognition (LVCSR). However, traditional language models only predict next single word with given…

cs.CL2022

DFM: Dialogue Foundation Model for Universal Large-Scale Dialogue-Oriented Task Learning

Zhi Chen, Jijia Bao, Lu Chen +10

Building a universal conversational agent has been a long-standing goal of the dialogue research community. Most previous works only focus on a small set of dialogue tasks. In this…

cs.CL2023

ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought

Hanchong Zhang, Ruisheng Cao, Lu Chen +2

Recently Large Language Models (LLMs) have been proven to have strong abilities in various domains and tasks. We study the problem of prompt designing in the text-to-SQL task and a…

eess.AS2020

Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model

Qi Liu, Zhehuai Chen, Hao Li +3

End-to-end (E2E) systems have played a more and more important role in automatic speech recognition (ASR) and achieved great performance. However, E2E systems recognize output word…

quant-ph2024

Quantum Convolutional Neural Network with Flexible Stride

Kai Yu, Song Lin, Bin-Bin Cai

Convolutional neural network is a crucial tool for machine learning, especially in the field of computer vision. Its unique structure and characteristics provide significant advant…

eess.AS2024

TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer

Yu Xi, Hao Li, Baochen Yang +3

Designing an efficient keyword spotting (KWS) system that delivers exceptional performance on resource-constrained edge devices has long been a subject of significant attention. Ex…

cs.CL2024

IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation

Senyu Han, Lu Chen, Li-Min Lin +2

Large language models have demonstrated their capabilities in storyline creation and human-like character role-playing. Current language model agents mainly focus on reasonable beh…

physics.flu-dyn2024

Discrepancy in Oil Displacement Mechanisms at the Equivalent Interfacial Tensions: Differentiating Contributions from Surfactant and Nanoparticles on Interfacial Activities

Suparit Tangparitkul, Thakheru Akamine, David Harbottle +2

This study examines discrepancies in oil displacement mechanisms at equivalent interfacial tensions, focusing on the distinct contributions of surfactants and nanoparticles. It was…

cs.CV2024

DiffusionGAN3D: Boosting Text-guided 3D Generation and Domain Adaptation by Combining 3D GANs and Diffusion Priors

Biwen Lei, Kai Yu, Mengyang Feng +2

Text-guided domain adaptation and generation of 3D-aware portraits find many applications in various fields. However, due to the lack of training data and the challenges in handlin…

cs.SD2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Ziyang Ma, Yinghao Ma, Yanqiao Zhu +31

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,00…

cs.CL2020

Prior Knowledge Driven Label Embedding for Slot Filling in Natural Language Understanding

Su Zhu, Zijian Zhao, Rao Ma +1

Traditional slot filling in natural language understanding (NLU) predicts a one-hot vector for each word. This form of label representation lacks semantic correlation modelling, wh…

cs.CL2022

Climate and Weather: Inspecting Depression Detection via Emotion Recognition

Wen Wu, Mengyue Wu, Kai Yu

Automatic depression detection has attracted increasing amount of attention but remains a challenging task. Psychological research suggests that depressive mood is closely related…

cs.CL2022

OPAL: Ontology-Aware Pretrained Language Model for End-to-End Task-Oriented Dialogue

Zhi Chen, Yuncong Liu, Lu Chen +3

This paper presents an ontology-aware pretrained language model (OPAL) for end-to-end task-oriented dialogue (TOD). Unlike chit-chat dialogue models, task-oriented dialogue models…

cs.CL2021

LET: Linguistic Knowledge Enhanced Graph Transformer for Chinese Short Text Matching

Boer Lyu, Lu Chen, Su Zhu +1

Chinese short text matching is a fundamental task in natural language processing. Existing approaches usually take Chinese characters or words as input tokens. They have two limita…

cs.SD2026

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

Wenbin Huang, Yuhang Qiu, Bohan Li +5

Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream…

cs.SD2026

MMAE: A Massive Multitask Audio Editing Benchmark

Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…

quant-ph2025

Knowledge Distillation for Variational Quantum Convolutional Neural Networks on Heterogeneous Data

Kai Yu, Binbin Cai, Song Lin

Distributed quantum machine learning faces significant challenges due to heterogeneous client data and variations in local model structures, which hinder global model aggregation.…

cs.AI2026

Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework

Xiaohua Wang, Chao Han, Kai Yu +2

The rapid evolution of Large Language Model (LLM) agents has produced diverse interaction paradigms, yet few production systems integrate multiple paradigms within a unified archit…