papers

Publications (17)

cs.SD2022

Progressive Learning for Stabilizing Label Selection in Speech Separation with Mapping-based Method

Chenyang Gao, Yue Gu, Ivan Marsic

Speech separation has been studied in time domain because of lower latency and higher performance compared to time-frequency domain. The masking-based method has been mostly used i…

cs.CV2021

VidTr: Video Transformer Without Convolutions

Yanyi Zhang, Xinyu Li, Chunhui Liu +6

We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatio-temporal infor…

cs.CV2021

Multi-Label Activity Recognition using Activity-specific Features and Activity Correlations

Yanyi Zhang, Xinyu Li, Ivan Marsic

Multi-label activity recognition is designed for recognizing multiple activities that are performed simultaneously or sequentially in each video. Most recent activity recognition n…

cs.AI2022

Exploring Runtime Decision Support for Trauma Resuscitation

Keyi Li, Sen Yang, Travis M. Sullivan +2

AI-based recommender systems have been successfully applied in many domains (e.g., e-commerce, feeds ranking). Medical experts believe that incorporating such methods into a clinic…

cs.CV2024

MaskMatch: Boosting Semi-Supervised Learning Through Mask Autoencoder-Driven Feature Learning

Wenjin Zhang, Keyi Li, Sen Yang +4

Conventional methods in semi-supervised learning (SSL) often face challenges related to limited data utilization, mainly due to their reliance on threshold-based techniques for sel…

cs.SD2023

Improving Label Assignments Learning by Dynamic Sample Dropout Combined with Layer-wise Optimization in Speech Separation

Chenyang Gao, Yue Gu, Ivan Marsic

In supervised speech separation, permutation invariant training (PIT) is widely used to handle label ambiguity by selecting the best permutation to update the model. Despite its su…

cs.LG2022

Generating Privacy-Preserving Process Data with Deep Generative Models

Keyi Li, Sen Yang, Travis M. Sullivan +2

Process data with confidential information cannot be shared directly in public, which hinders the research in process data mining and analytics. Data encryption methods have been s…

cs.CV2017

Online People Tracking and Identification with RFID and Kinect

Xinyu Li, Yanyi Zhang, Ivan Marsic +1

We introduce a novel, accurate and practical system for real-time people tracking and identification. We used a Kinect V2 sensor for tracking that generates a body skeleton for up…

cs.CV2022

TubeR: Tubelet Transformer for Video Action Detection

Jiaojiao Zhao, Yanyi Zhang, Xinyu Li +10

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed ac…

cs.LG2017

Progress Estimation and Phase Detection for Sequential Processes

Xinyu Li, Yanyi Zhang, Jianyu Zhang +7

Process modeling and understanding are fundamental for advanced human-computer interfaces and automation systems. Most recent research has focused on activity recognition, but litt…

cs.CV2017

Concurrent Activity Recognition with Multimodal CNN-LSTM Structure

Xinyu Li, Yanyi Zhang, Jianyu Zhang +4

We introduce a system that recognizes concurrent activities from real-world data captured by multiple sensors of different types. The recognition is achieved in two steps. First, w…

eess.AS2019

RHR-Net: A Residual Hourglass Recurrent Neural Network for Speech Enhancement

Jalal Abdulbaqi, Yue Gu, Ivan Marsic

Most current speech enhancement models use spectrogram features that require an expensive transformation and result in phase information loss. Previous work has overcome these issu…

cs.DS2017

Process-oriented Iterative Multiple Alignment for Medical Process Mining

Shuhong Chen, Sen Yang, Moliang Zhou +2

Adapted from biological sequence alignment, trace alignment is a process mining technique used to visualize and analyze workflow data. Any analysis done with this method, however,…

cs.CL2018

Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment

Yue Gu, Kangning Yang, Shiyu Fu +3

Multimodal affective computing, learning to recognize and interpret human affects and subjective information from multiple data sources, is still challenging because: (i) it is har…

cs.OH2017

Evaluation of Trace Alignment Quality and its Application in Medical Process Mining

Moliang Zhou, Sen Yang, Shuyu Lv +5

Trace alignment algorithms have been used in process mining for discovering the consensus treatment procedures and process deviations. Different alignment algorithms, however, may…

cs.CV2018

Tri-axial Self-Attention for Concurrent Activity Recognition

Yanyi Zhang, Xinyu Li, Kaixiang Huang +3

We present a system for concurrent activity recognition. To extract features associated with different activities, we propose a feature-to-activity attention that maps the extracte…

cs.CL2018

Deep Multimodal Learning for Emotion Recognition in Spoken Language

Yue Gu, Shuhong Chen, Ivan Marsic

In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics.…