papers

Publications (9)

cs.SD2021

Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study

Dawei Liang, Yangyang Shi, Yun Wang +6

Detection of common events and scenes from audio is useful for extracting and understanding human contexts in daily life. Prior studies have shown that leveraging knowledge from a…

cs.SD2022

Automated detection of foreground speech with wearable sensing in everyday home environments: A transfer learning approach

Dawei Liang, Zifan Xu, Yinuo Chen +3

Acoustic sensing has proved effective as a foundation for numerous applications in health and human behavior analysis. In this work, we focus on the problem of detecting in-person…

cs.CV2020

Cross-modal supervised learning for better acoustic representations

Shaoyong Jia, Xin Shu, Yang Yang +3

Obtaining large-scale human-labeled datasets to train acoustic representation models is a very challenging task. On the contrary, we can easily collect data with machine-generated…

cs.CV2025

FG-CLIP: Fine-Grained Visual and Textual Alignment

Chunyu Xie, Bin Wang, Fanjing Kong +5

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding du…

cs.LG2026

Detecting In-Person Conversations in Noisy Real-World Environments with Smartwatch Audio and Motion Sensing

Alice Zhang, Callihan Bertley, Dawei Liang +1

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-ver…

cs.HC2018

Audio-Based Activities of Daily Living (ADL) Recognition with Large-Scale Acoustic Embeddings from Online Videos

Dawei Liang, Edison Thomaz

Over the years, activity sensing and recognition has been shown to play a key enabling role in a wide range of applications, from sustainability and human-computer interaction to h…

cs.SD2022

Dynamic Speech Endpoint Detection with Regression Targets

Dawei Liang, Hang Su, Tarun Singh +5

Interactive voice assistants have been widely used as input interfaces in various scenarios, e.g. on smart homes devices, wearables and on AR devices. Detecting the end of a speech…

cs.CV2026

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

Chunyu Xie, Bin Wang, Fanjing Kong +5

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, parti…

cs.CV2025

RzenEmbed: Towards Comprehensive Multimodal Retrieval

Weijian Jian, Yajun Zhang, Dawei Liang +4

The rapid advancement of Multimodal Large Language Models (MLLMs) has extended CLIP-based frameworks to produce powerful, universal embeddings for retrieval tasks. However, existin…