collaborators

5 papers

cs.CV2026

DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

Koki Nagano, Hongyu Liu, Seonwook Park +9

We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this f…

cs.CV2026

LongCat-Video-Avatar 1.5 Technical Report

Meituan LongCat Team, Xunliang Cai, Meng Cheng +10

Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upgraded open-source framework…

cs.SD2026

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

Yuxiang Wang, Hongyu Liu, Yijiang Xu +9

As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speak…

cs.SD2025

DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition

HongYu Liu, Junxin Li, Changxi Guo +5

Recognizing speaker intent in long audio dialogues among speakers has a wide range of applications, but is a non-trivial AI task due to complex inter-dependencies in speaker uttera…

cs.SD2025

MSMT-FN: Multi-segment Multi-task Fusion Network for Marketing Audio Classification

HongYu Liu, Ruijie Wan, Yueju Han +4

Audio classification plays an essential role in sentiment analysis and emotion recognition, especially for analyzing customer attitudes in marketing phone calls. Efficiently catego…