Publications (88)
Distortion Recovery: A Two-Stage Method for Guitar Effect Removal
Ying-Shuo Lee, Yueh-Po Peng, Jui-Te Wu +3
Removing audio effects from electric guitar recordings makes it easier for post-production and sound editing. An audio distortion recovery model not only improves the clarity of th…
Density-guided Translator Boosts Synthetic-to-Real Unsupervised Domain Adaptive Segmentation of 3D Point Clouds
Zhimin Yuan, Wankang Zeng, Yanfei Su +4
3D synthetic-to-real unsupervised domain adaptive segmentation is crucial to annotating new domains. Self-training is a competitive approach for this task, but its performance is l…
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
Xingjian Diao, Weiyi Wu, Keyi Kong +5
Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous…
GMA3D: Local-Global Attention Learning to Estimate Occluded Motions of Scene Flow
Zhiyang Lu, Ming Cheng
Scene flow represents the motion information of each point in the 3D point clouds. It is a vital downstream method applied to many tasks, such as motion segmentation and object tra…
On Evaluating and Comparing Open Domain Dialog Systems
Anu Venkatesh, Chandra Khatri, Ashwin Ram +10
Conversational agents are exploding in popularity. However, much work remains in the area of non goal-oriented conversations, despite significant growth in research interest over r…
Joint Beamforming and Computation Offloading for Multi-user Mobile-Edge Computing
Changfeng Ding, Jun-Bo Wang, Ming Cheng +3
Mobile edge computing (MEC) is considered as an efficient method to relieve the computation burden of mobile devices. In order to reduce the energy consumption and time delay of mo…
Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis
Hongyu Sun, Qiuhong Ke, Ming Cheng +4
This paper proposes a general solution to enable point cloud recognition models to handle distribution shifts at test time. Unlike prior methods, which rely heavily on training dat…
OPDN: Omnidirectional Position-aware Deformable Network for Omnidirectional Image Super-Resolution
Xiaopeng Sun, Weiqi Li, Zhenyu Zhang +8
360° omnidirectional images have gained research attention due to their immersive and interactive experience, particularly in AR/VR applications. However, they suffer from lower a…
Transformer for seismic image super-resolution
Shiqi Dong, Xintong Dong, Kaiyuan Zheng +3
Seismic images obtained by stacking or migration are usually characterized as low signal-to-noise ratio (SNR), low dominant frequency and sparse sampling both in depth (or time) an…
VeTraSS: Vehicle Trajectory Similarity Search Through Graph Modeling and Representation Learning
Ming Cheng, Bowen Zhang, Ziyu Wang +4
Trajectory similarity search plays an essential role in autonomous driving, as it enables vehicles to analyze the information and characteristics of different trajectories to make…
Learning Musical Representations for Music Performance Question Answering
Xingjian Diao, Chunhui Zhang, Tingxuan Wu +4
Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals th…
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
Ming Cheng, Jiaying Gong, Chenhan Yuan +3
Existing text simplification or paraphrase datasets mainly focus on sentence-level text generation in a general domain. These datasets are typically developed without using domain…
CVKD-UDA: Cross-View Knowledge Distillation for 3D Unsupervised Domain Adaptive Segmentation
Zhimin Yuan, Ming Cheng, Shangshu Yu +4
3D unsupervised domain adaptive (UDA) segmentation mitigates the high cost of manual annotations of the new domain data. Self-training has emerged as the dominant approach in this…
Conversational AI: The Science Behind the Alexa Prize
Ashwin Ram, Rohit Prasad, Chandra Khatri +15
Conversational agents are exploding in popularity. However, much work remains in the area of social conversation as well as free-form conversation over a broad range of domains and…
Solitary wave of the Schrodinger lattice system with nonlinear hopping
Ming Cheng
This paper is concerned with the nonlinear Schrodinger lattice with nonlinear hopping. Via variation approach and the Nehari manifold argument, we obtain two types of solution: pe…
A Process-Aware Demand Response Evaluation Framework for Hydrogen-Integrated Zero-Carbon Steel Plants Coupled with Methanol Production
Qiang Ji, Lin Cheng, Yue Zhou +4
High penetration of renewables (RES) and the retirement of thermal units aggravate flexibility scarcity in power systems. Hydrogen-based low-carbon steel production systems possess…
Spin-orbit-enabled Fermi-surface splitting in noncollinear antiferromagnetic SmBi
Long Zhang, Ming Cheng, Jingyu Li +10
Spin-split electronic structures in compensated antiferromagnets are commonly sought in the nonrelativistic limit, where magnetic order lifts spin degeneracy without spin-orbit cou…
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
Haoxu Wang, Ming Cheng, Qiang Fu +1
In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWW…
Bandit Inspired Beam Searching Scheme for mmWave High-Speed Train Communications
Jun-Bo Wang, Ming Cheng, Jin-Yuan Wang +4
High-speed trains (HSTs) are being widely deployed around the world. To meet the high-rate data transmission requirements on HSTs, millimeter wave (mmWave) HST communications have…
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
Ze Li, Ming Cheng, Ming Li
Large-scale self-supervised Pre-Trained Models (PTMs) have shown significant improvements in the speaker verification (SV) task by providing rich feature representations. In this p…
Expanding solutions near unstable Lane-Emden stars
Ming Cheng, Xing Cheng, Zhiwu Lin
We consider the compressible Euler-Poisson equations for polytropes with and the white dwarf stars. For $γ=\frac{4}{3}…
CrossGP: Cross-Day Glucose Prediction Excluding Physiological Information
Ziyi Zhou, Ming Cheng, Yanjun Cui +2
The increasing number of diabetic patients is a serious issue in society today, which has significant negative impacts on people's health and the country's financial expenditures.…
Outage Constrained Robust Secure Beamforming in Cognitive Satellite-Aerial Networks
Bai Zhao, Min Lin, Ming Cheng +2
This paper proposes a robust beamforming scheme to enhance the physical layer security (PLS) of multicast transmission in a cognitive satellite and aerial network (CSAN) operating…
Anisotropy of linear magnetoresistance in Kagome metal ZrVSn
Yifan Deng, Ming Cheng, Lanxin Liu +9
The Kagome lattice has attracted extensive attention due to the diverse magnetic properties and non-trivial electronic states generated by its unique atomic arrangement, which prov…
SAIC: Integration of Speech Anonymization and Identity Classification
Ming Cheng, Xingjian Diao, Shitong Cheng +1
Speech anonymization and de-identification have garnered significant attention recently, especially in the healthcare area including telehealth consultations, patient voiceprint ma…
The DKU-DukeECE Diarization System for the VoxCeleb Speaker Recognition Challenge 2022
Weiqing Wang, Xiaoyi Qin, Ming Cheng +3
This paper discribes the DKU-DukeECE submission to the 4th track of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our system contains a fused voice activity detectio…
The evolution of network controllability in growing networks
Rui Zhang, Xiaomeng Wang, Ming Cheng +1
The study of network structural controllability focuses on the minimum number of driver nodes needed to control a whole network. Despite intensive studies on this topic, most of th…
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation
Jianbo Dai, Jianqiao Lu, Yunlong Feng +6
Recent advancements in large language models (LLMs) have greatly improved code generation, specifically at the function level. For instance, GPT-4o has achieved a 91.0\% pass rate…
VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
Ming Cheng, Tong Wu, Jiazhen Hu +2
Attribute Value Extraction (AVE) is important for structuring product information in e-commerce. However, existing AVE datasets are primarily limited to text-to-text or image-to-te…
Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
Li Li, Ming Cheng, Juan Liu +1
This paper proposes a Spatially-Augmented Sequence-to-Sequence Neural Diarization (SA-S2SND) framework, which integrates direction-of-arrival (DOA) cues estimated by SRP-DNN into t…
FT2TF: First-Person Statement Text-To-Talking Face Generation
Xingjian Diao, Ming Cheng, Wayner Barrios +1
Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatar…
Advancing the State of the Art in Open Domain Dialog Systems through the Alexa Prize
Chandra Khatri, Behnam Hedayatnia, Anu Venkatesh +19
Building open domain conversational systems that allow users to have engaging conversations on topics of their choice is a challenging task. Alexa Prize was launched in 2016 to tac…
Learned Quality Enhancement via Multi-Frame Priors for HEVC Compliant Low-Delay Applications
Ming Lu, Ming Cheng, Yiling Xu +3
Networked video applications, e.g., video conferencing, often suffer from poor visual quality due to unexpected network fluctuation and limited bandwidth. In this paper, we have de…
A Dual Camera System for High Spatiotemporal Resolution Video Acquisition
Ming Cheng, Zhan Ma, M. Salman Asif +4
This paper presents a dual camera system for high spatiotemporal resolution (HSTR) video acquisition, where one camera shoots a video with high spatial resolution and low frame rat…
Intelligent Autofocus
Chengyu Wang, Qian Huang, Ming Cheng +2
We demonstrate that deep learning methods can determine the best focus position from 1-2 image samples, enabling 5-10x faster focus than traditional search-based methods. In contra…
Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
Ming Cheng, Yuke Lin, Ming Li
This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequen…
Turning point principle for stability of viscous gaseous stars
Ming Cheng, Zhiwu Lin, Yucong Wang
We consider stability of non-rotating viscous gaseous stars modeled by the Navier-Stokes-Poisson system. Under general assumptions on the equations of states, we proved that the nu…
Ultrafast μeV-Precision Bandgap Engineering in Low-Dimensional Topological Insulators
Peng Tan, Yuantao Chen, Yuqi Zhang +14
Precise and ultrafast control of electronic band structures is a central challenge for advancing quantum functional materials and devices. Conventional approaches--such as chemical…
LO-Net: Deep Real-time Lidar Odometry
Qing Li, Shaoyang Chen, Cheng Wang +4
We present a novel deep convolutional network pipeline, LO-Net, for real-time lidar odometry estimation. Unlike most existing lidar odometry (LO) estimations that go through indivi…
DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models
Sicheng Yang, Zhiyong Wu, Minglei Li +5
The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to th…
DLA-Net: Learning Dual Local Attention Features for Semantic Segmentation of Large-Scale Building Facade Point Clouds
Yanfei Su, Weiquan Liu, Zhimin Yuan +4
Semantic segmentation of building facade is significant in various applications, such as urban building reconstruction and damage assessment. As there is a lack of 3D point clouds…
Walking Further: Semantic-aware Multimodal Gait Recognition Under Long-Range Conditions
Zhiyang Lu, Wen Jiang, Tianren Wu +4
Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range,…
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
Xingjian Diao, Chunhui Zhang, Weiyi Wu +5
Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models fa…
Geometry-aware Depth-guided Representation Learning for Structure-preserving Low-light Image Enhancement
Fang Gao, Jiongkai Qin, Jiabao Wang +5
Low-light degradation reduces image visibility and weakens structural cues that are important for visual representation and scene understanding. Existing low-light image enhancemen…
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
Ming Cheng, Ming Li
Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that…
Multi-Graph Fusion Networks for Urban Region Embedding
Shangbin Wu, Xu Yan, Xiaoliang Fan +5
Learning the embeddings for urban regions from human mobility data can reveal the functionality of regions, and then enables the correlated but distinct tasks such as crime predict…
Hybrid Transformer and CNN Attention Network for Stereo Image Super-resolution
Ming Cheng, Haoyu Ma, Qiufang Ma +7
Multi-stage strategies are frequently employed in image restoration tasks. While transformer-based methods have exhibited high efficiency in single-image super-resolution tasks, th…
RWF-2000: An Open Large Scale Video Database for Violence Detection
Ming Cheng, Kunjing Cai, Ming Li
In recent years, surveillance cameras are widely deployed in public places, and the general crime rate has been reduced significantly due to these ubiquitous devices. Usually, thes…
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
Li Li, Ming Cheng, Weixin Zhu +3
Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but al…
Design and Research of a Self-Propelled Pipeline Robot Based on Force Analysis and Dynamic Simulation
Yan Gao, Jiliang Wang, Ming Cheng +1
In pipeline inspection, traditional tethered inspection robots are severely constrained by cable length and weight, which greatly limit their travel range and accessibility. To add…
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
Zhiyang Lu, Ming Cheng
Cross-modal 2D-3D gait recognition is impeded by inherent domain discrepancies between 2D silhouette and 3D LiDAR range-view representations. While prior methods align only final e…
Viia-hand: a Reach-and-grasp Restoration System Integrating Voice interaction, Computer vision and Auditory feedback for Blind Amputees
Chunhao Peng, Dapeng Yang, Ming Cheng +3
Visual feedback plays a crucial role in the process of amputation patients completing grasping in the field of prosthesis control. However, for blind and visually impaired (BVI) am…
STARFlow: Spatial Temporal Feature Re-embedding with Attentive Learning for Real-world Scene Flow
Zhiyang Lu, Qinghan Chen, Ming Cheng
Scene flow prediction is a crucial underlying task in understanding dynamic scenes as it offers fundamental motion information. However, contemporary scene flow methods encounter t…
VoxBlink: A Large Scale Speaker Verification Dataset on Camera
Yuke Lin, Xiaoyi Qin, Guoqing Zhao +4
In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual dat…
A material decomposition method for dual-energy CT via dual interactive Wasserstein generative adversarial networks
Zaifeng Shi, Huilong Li, Qingjie Cao +2
Dual-energy computed tomography has great potential in material characterization and identification, whereas the reconstructed material-specific images always suffer from magnified…
Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
Ming Cheng, Fei Su, Cancan Li +2
This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequen…
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
Yuke Lin, Ming Cheng, Ze Li +1
We present the DKU system for Task 2 of the MLC-SLM Challenge, which aims to perform multi-speaker automatic speech recognition directly from raw audio without Oracle speaker label…
H2-Stereo: High-Speed, High-Resolution Stereoscopic Video System
Ming Cheng, Yiling Xu, Wang Shen +4
High-speed, high-resolution stereoscopic (H2-Stereo) video allows us to perceive dynamic 3D content at fine granularity. The acquisition of H2-Stereo video, however, remains challe…
Partial Procedural Geometric Model Fitting for Point Clouds
Zongliang Zhang, Jonathan Li, Yulan Guo +3
Geometric model fitting is a fundamental task in computer graphics and computer vision. However, most geometric model fitting methods are unable to fit an arbitrary geometric model…
Field test of mode-pairing quantum key distribution
Hao-Tao Zhu, Yizhi Huang, Wen-Xin Pan +10
Quantum key distribution is a cornerstone of quantum technology, offering information-theoretical secure keys for remote parties. With many quantum communication networks establish…
The DKU-MSXF Diarization System for the VoxCeleb Speaker Recognition Challenge 2023
Ming Cheng, Weiqing Wang, Xiaoyi Qin +4
This paper describes the DKU-MSXF submission to track 4 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). Our system pipeline contains voice activity detection, clust…
Text-guided Feature Disentanglement for Cross-modal Gait Recognition
Zhiyang Lu, Ming Cheng
Gait recognition is a biometric technique that identifies individuals based on their walking patterns, offering advantages in long-range, non-intrusive scenarios. However, real-wor…
An Efficient Solver for Cumulative Density Function-based Solutions of Uncertain Kinematic Wave Models
Ming Cheng, Yi Qin, Akil Narayan +3
We develop a numerical framework to implement the cumulative density function (CDF) method for obtaining the probability distribution of the system state described by a kinematic w…
Seismic Interpolation Transformer for Consecutively Missing Data: A Case Study in DAS-VSP Data
Ming Cheng, Jun Lin, Xintong Dong +2
Distributed optical fiber acoustic sensing (DAS) is a rapidly-developed seismic acquisition technology with advantages of low cost, high resolution, high sensitivity, and small int…
NewsNet-SDF: Stochastic Discount Factor Estimation with Pretrained Language Model News Embeddings via Adversarial Networks
Shunyao Wang, Ming Cheng, Christina Dan Wang
Stochastic Discount Factor (SDF) models provide a unified framework for asset pricing and risk assessment, yet traditional formulations struggle to incorporate unstructured textual…
GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian Splatting
Changkun Liu, Shuai Chen, Yash Bhalgat +5
We leverage 3D Gaussian Splatting (3DGS) as a scene representation and propose a novel test-time camera pose refinement (CPR) framework, GS-CPR. This framework enhances the localiz…
GluMarker: A Novel Predictive Modeling of Glycemic Control Through Digital Biomarkers
Ziyi Zhou, Ming Cheng, Xingjian Diao +2
The escalating prevalence of diabetes globally underscores the need for diabetes management. Recent research highlights the growing focus on digital biomarkers in diabetes manageme…
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
Yuke Lin, Ming Cheng, Ze Li +2
Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and…
LACE: Exploring Turn-Taking and Parallel Interaction Modes in Human-AI Co-Creation for Iterative Image Generation
YenKai Huang, Zheng Ning, Ming Cheng
This paper introduces LACE, a co-creative system enabling professional artists to leverage generative AI through controlled prompting and iterative refinement within Photoshop. Add…
AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoencoder
Xingjian Diao, Ming Cheng, Shitong Cheng
Learning high-quality video representation has shown significant applications in computer vision and remains challenging. Previous work based on mask autoencoders such as ImageMAE…
Practical quantum access network over a 10 Gbit/s Ethernet passive optical network
Bi-Xiao Wang, Shi-Biao Tang, Yingqiu Mao +5
Quantum key distribution (QKD) provides an information-theoretically secure method to share keys between legitimate users. To achieve large-scale deployment of QKD, it should be ea…
VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Yuke Lin, Ming Cheng, Fulin Zhang +3
In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This…
Efflex: Efficient and Flexible Pipeline for Spatio-Temporal Trajectory Graph Modeling and Representation Learning
Ming Cheng, Ziyi Zhou, Bowen Zhang +7
In the landscape of spatio-temporal data analytics, effective trajectory representation learning is paramount. To bridge the gap of learning accurate representations with efficient…
SSRFlow: Semantic-aware Fusion with Spatial Temporal Re-embedding for Real-world Scene Flow
Zhiyang Lu, Qinghan Chen, Zhimin Yuan +1
Scene flow, which provides the 3D motion field of the first frame from two consecutive point clouds, is vital for dynamic scene perception. However, contemporary scene flow methods…
RF-Net: An End-to-End Image Matching Network based on Receptive Field
Xuelun Shen, Cheng Wang, Xin Li +5
This paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable mat…
Visual Zero-Shot E-Commerce Product Attribute Value Extraction
Jiaying Gong, Ming Cheng, Hongda Shen +3
Existing zero-shot product attribute value (aspect) extraction approaches in e-Commerce industry rely on uni-modal or multi-modal models, where the sellers are asked to provide det…
Uncertainty-Aware Motion Planning for Autonomous Driving in Mixed Traffic Environment
Ming Cheng, Hao Chen, Ziyi Yang +2
In mixed-traffic environments where autonomous and human-driven vehicles may co-exist, motion planning for autonomous vehicles requires anticipating the future behaviors of surroun…
Learning Sparsity for Effective and Efficient Music Performance Question Answering
Xingjian Diao, Tianzhen Yang, Chunhui Zhang +3
Music performances, characterized by dense and continuous audio as well as seamless audio-visual integration, present unique challenges for multimodal scene understanding and reaso…
Local Density of States and Angle-Resolved Photoemission Spectral Function of an Inhomogeneous D-wave Superconductor
Ming Cheng, W. P. Su
Nanoscale inhomogeneity seems to be a central feature of the d-wave superconductivity in the cuprates. Such a feature can strongly affect the local density of states (LDOS) and the…
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
Wenhao You, Xingjian Diao, Wenjun Huang +9
While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Au…
Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
Ming Cheng, Jiaying Gong, Hoda Eldardiry
Lay paraphrasing aims to make scientific information accessible to audiences without technical backgrounds. However, most existing studies focus on a single domain, such as biomedi…
Masked Cross-image Encoding for Few-shot Segmentation
Wenbo Xu, Huaxi Huang, Ming Cheng +3
Few-shot segmentation (FSS) is a dense prediction task that aims to infer the pixel-wise labels of unseen classes using only a limited number of annotated images. The key challenge…
DASGIL: Domain Adaptation for Semantic and Geometric-aware Image-based Localization
Hanjiang Hu, Zhijian Qiao, Ming Cheng +2
Long-Term visual localization under changing environments is a challenging problem in autonomous driving and mobile robotics due to season, illumination variance, etc. Image retrie…
Toward Short-Term Glucose Prediction Solely Based on CGM Time Series
Ming Cheng, Xingjian Diao, Ziyi Zhou +3
The global diabetes epidemic highlights the importance of maintaining good glycemic control. Glucose prediction is a fundamental aspect of diabetes management, facilitating real-ti…
Smart Cameras
David J. Brady, Minghao Hu, Chengyu Wang +6
We review camera architecture in the age of artificial intelligence. Modern cameras use physical components and software to capture, compress and display image data. Over the past…
Target-Speaker Voice Activity Detection via Sequence-to-Sequence Prediction
Ming Cheng, Weiqing Wang, Yucong Zhang +2
Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence…
The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis
Haoxu Wang, Ming Cheng, Qiang Fu +1
This paper further explores our previous wake word spotting system ranked 2-nd in Track 1 of the MISP Challenge 2021. First, we investigate a robust unimodal approach based on 3D a…
A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling
Shiyu Ji, Farnoosh Hashemi, Joice Chen +9
Rhetorical strategies are central to persuasive communication, from political discourse and marketing to legal argumentation. However, analysis of rhetorical strategies has been li…