papers

Publications (66)

cs.RO2025

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

Teli Ma, Jia Zheng, Zifan Wang +3

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of ac…

cs.LG2024

CKGConv: General Graph Convolution with Continuous Kernels

Liheng Ma, Soumyasundar Pal, Yitian Zhang +3

The existing definitions of graph convolution, either from spatial or spectral perspectives, are inflexible and not unified. Defining a general convolution operator in the graph do…

cs.SD2025

M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper

Jiaming Zhou, Shiwan Zhao, Jiabei He +6

State-of-the-art models like OpenAI's Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizin…

cs.CV2024

ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Jiaming Zhou, Junwei Liang, Kun-Yu Lin +2

Zero-shot action recognition (ZSAR) aims to learn an alignment model between videos and class descriptions of seen actions that is transferable to unseen actions. The text queries…

cs.SD2025

DIFFA: Large Language Diffusion Models Can Listen and Understand

Jiaming Zhou, Hongjie Chen, Shiwan Zhao +9

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged…

cs.CV2023

GeoDeformer: Geometric Deformable Transformer for Action Recognition

Jinhui Ye, Jiaming Zhou, Hui Xiong +1

Vision transformers have recently emerged as an effective alternative to convolutional networks for action recognition. However, vision transformers still struggle with geometric v…

cs.SD2024

AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection

Rong Gong, Hongfei Xue, Lezhi Wang +11

The rapid advancements in speech technologies over the past two decades have led to human-level performance in tasks like automatic speech recognition (ASR) for fluent speech. Howe…

cs.CL2023

MADI: Inter-domain Matching and Intra-domain Discrimination for Cross-domain Speech Recognition

Jiaming Zhou, Shiwan Zhao, Ning Jiang +2

End-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) ai…

cs.SD2025

WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations

Hui Wang, Jiaming Zhou, Jiabei He +2

Elderly speech poses unique challenges for automatic processing due to age-related changes such as slower articulation and vocal tremors. Existing Chinese datasets are mostly recor…

cs.CL2025

Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores

Jiaming Zhou, Shiwan Zhao, Hui Wang +4

The kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, pres…

cs.CV2023

Diversifying Spatial-Temporal Perception for Video Domain Generalization

Kun-Yu Lin, Jia-Run Du, Yipeng Gao +2

Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain g…

cs.CV2024

Human-Centric Transformer for Domain Adaptive Action Recognition

Kun-Yu Lin, Jiaming Zhou, Wei-Shi Zheng

We study the domain adaptation task for action recognition, namely domain adaptive action recognition, which aims to effectively transfer action recognition power from a label-suff…

cs.SD2026

SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation

Hui Wang, Jinghua Zhao, Yifan Yang +9

Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scala…

cs.RO2024

Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation

Teli Ma, Jiaming Zhou, Zifan Wang +2

Developing robots capable of executing various manipulation tasks, guided by natural language instructions and visual observations of intricate real-world environments, remains a s…

cs.SD2025

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

Jiaming Zhou, Shiyao Wang, Shiwan Zhao +10

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However,…

cs.SD2025

A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition

Shiyao Wang, Jiaming Zhou, Shiwan Zhao +1

Dysarthric speech recognition (DSR) enhances the accessibility of smart devices for dysarthric speakers with limited mobility. Previously, DSR research was constrained by the fact…

cs.CL2026

Extracting and Following Paths for Robust Relational Reasoning with Large Language Models

Ge Zhang, Mohammad Ali Alomrani, Hongjian Gu +7

Large language models (LLMs) possess vast semantic knowledge but often struggle with complex reasoning tasks, particularly in relational reasoning problems such as kinship or spati…

cs.SD2025

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Haoqin Sun, Jingguang Tian, Jiaming Zhou +8

The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the…

cs.CL2025

CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition

Jiaming Zhou, Yujie Guo, Shiwan Zhao +9

Code-switching (CS), the alternation between two or more languages within a single conversation, presents significant challenges for automatic speech recognition (ASR) systems. Exi…

cs.CV2025

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

Jiaming Zhou, Teli Ma, Kun-Yu Lin +3

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited s…

cs.RO2025

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Ke Ye, Jiaming Zhou, Yuanfeng Qiu +4

Generalizing to long-horizon manipulation tasks in a zero-shot setting remains a central challenge in robotics. Current multimodal foundation based approaches, despite their capabi…

cs.CL2026

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning

Shiwan Zhao, Zhihu Wang, Xuyang Zhao +10

Post-training has become central to turning pretrained large language models (LLMs) into aligned, capable, and deployable systems. Recent progress spans supervised fine-tuning (SFT…

cs.SD2026

AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation

Hui Wang, Jinghua Zhao, Junyang Cheng +5

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only li…

cs.SD2024

Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition

Haoqin Sun, Shiwan Zhao, Xiangyu Kong +4

Recognizing emotions from speech is a daunting task due to the subtlety and ambiguity of expressions. Traditional speech emotion recognition (SER) systems, which typically rely on…

cs.RO2026

Native Video-Action Pretraining for Generalizable Robot Control

Qihang Zhang, Lin Li, Luyao Zhang +26

The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…

#robot control#video-action models#foundation models#causal pretraining
cs.CL2026

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning

Yihong Wu, Liheng Ma, Muzhi Li +7

Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to interface with search engines for co…

cs.CV2026

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Gangwei Xu, Qihang Zhang, Jiaming Zhou +4

Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence and limited co…

eess.AS2024

Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge

Hongfei Xue, Rong Gong, Mingchen Shao +10

The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recogni…

cs.CL2024

Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs

Aobo Kong, Shiwan Zhao, Hao Chen +6

Recent advancements in LLMs have showcased their remarkable role-playing capabilities, able to accurately simulate the dialogue styles and cognitive processes of various roles base…

cs.RO2025

GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Teli Ma, Zifan Wang, Jiaming Zhou +2

Inferring affordable (i.e., graspable) parts of arbitrary objects based on human specifications is essential for robots advancing toward open-vocabulary manipulation. Current grasp…

eess.AS2026

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning

Hui Wang, Yifan Yang, Zeyue Tian +8

Audio generation and audio-to-text understanding remain largely separate, with diffusion models dominating high-fidelity synthesis and autoregressive (AR) language models driving c…

physics.geo-ph2023

Phase Field Characterization of Rock Fractures in Brazilian Splitting Test Specimens Containing Voids and Inclusions

Shuwei Zhou, Xiaoying Zhuang, Jiaming Zhou +1

The Brazilian splitting test is a widely used testing procedure for characterizing the tensile strength of natural rock or rock-like material due to the fact. However, the results…

cs.CL2025

FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Hui Wang, Shujie Liu, Lingwei Meng +9

To advance continuous-valued token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow mat…

cs.CV2024

Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition

Jiaming Zhou, Hanjun Li, Kun-Yu Lin +1

Developing end-to-end action recognition models on long videos is fundamental and crucial for long-video action understanding. Due to the unaffordable cost of end-to-end training o…

cs.LG2024

PostRainBench: A comprehensive benchmark and a new model for precipitation forecasting

Yujin Tang, Jiaming Zhou, Xiang Pan +2

Accurate precipitation forecasting is a vital challenge of societal importance. Though data-driven approaches have emerged as a widely used solution, solely relying on data-driven…

cs.SD2026

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Junyang Chen, Yuhang Jia, Hui Wang +2

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing.…

cs.SD2025

MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation

Cheng Liu, Hui Wang, Jinghua Zhao +6

The technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primaril…

cs.SD2026

DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding

Jiaming Zhou, Xuxin Cheng, Shiwan Zhao +5

Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains cost…

cs.SD2024

Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation

Shiyao Wang, Shiwan Zhao, Jiaming Zhou +2

Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to…

cs.MM2025

EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

Haoqin Sun, Xuechen Wang, Jinghua Zhao +9

In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for…

cs.CV2026

XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition

Kun-Yu Lin, Henghui Ding, Jia-Run Du +6

Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective vi…

cs.CL2024

Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data

Jiaming Zhou, Abbas Ghaddar, Ge Zhang +7

Despite recent advances in training and prompting strategies for Large Language Models (LLMs), these models continue to face challenges with complex logical reasoning tasks that in…

cs.RO2025

Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments

Zifan Wang, Teli Ma, Yufei Jia +5

Agile locomotion in complex 3D environments requires robust spatial awareness to safely avoid diverse obstacles such as aerial clutter, uneven terrain, and dynamic agents. Depth-ba…

cs.SD2024

PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge

Shiyao Wang, Jiaming Zhou, Shiwan Zhao +1

For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extra…

cs.RO2025

Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization

Jiaming Zhou, Ke Ye, Jiayi Liu +6

The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However…

cs.CV2025

FVAR: Visual Autoregressive Modeling via Next Focus Prediction

Xiaofan Li, Chenming Wu, Yanpeng Sun +6

Visual autoregressive models achieve remarkable generation quality through next-scale predictions across multi-scale token pyramids. However, the conventional method uses uniform s…

cs.SD2025

MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning

Haoqin Sun, Chenyang Lyu, Xiangyu Kong +9

Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discr…

cs.LG2024

Uncertainty-Aware Mean Opinion Score Prediction

Hui Wang, Shiwan Zhao, Jiaming Zhou +4

Mean Opinion Score (MOS) prediction has made significant progress in specific domains. However, the unstable performance of MOS prediction models across diverse samples presents on…

cs.RO2025

End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy

Zifan Wang, Xun Yang, Jianzhuang Zhao +5

The deployment of humanoid robots in unstructured, human-centric environments requires navigation capabilities that extend beyond simple locomotion to include robust perception, pr…

cs.LG2025

Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling

Derek Li, Jiaming Zhou, Leo Maxime Brunswic +8

The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-…

cs.SD2025

Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios

Shiyao Wang, Shiwan Zhao, Jiaming Zhou +1

Dysarthric speech recognition (DSR) research has witnessed remarkable progress in recent years, evolving from the basic understanding of individual words to the intricate comprehen…

cs.SD2024

CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition

Tian-Hao Zhang, Dinghao Zhou, Guiping Zhong +2

RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and th…

cs.SD2025

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

Hui Wang, Cheng Liu, Junyang Chen +7

Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generaliza…

cs.CL2025

RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis

Enzhi Wang, Qicheng Li, Shiwan Zhao +6

In recent years, large language models (LLMs) have achieved remarkable advancements in multimodal processing, including end-to-end speech-based language models that enable natural…

cs.SD2025

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

Hui Wang, Yifan Yang, Shujie Liu +7

Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time appli…

cs.SD2026

H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR

Yujie Guo, Jiaming Zhou, Yuhang Jia +2

Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. Wh…

cs.SD2026

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

Junyang Chen, Yuhang Jia, Hui Wang +3

Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consis…

cs.MM2025

Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides

Jinghua Zhao, Yuhang Jia, Shiyao Wang +3

Incorporating visual modalities to assist Automatic Speech Recognition (ASR) tasks has led to significant improvements. However, existing Audio-Visual Speech Recognition (AVSR) dat…

cs.SD2024

kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels

Jiaming Zhou, Shiwan Zhao, Yaqi Liu +3

The success of retrieval-augmented language models in various natural language processing (NLP) tasks has been constrained in automatic speech recognition (ASR) applications due to…

cs.CL2025

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

Yang Chen, Hui Wang, Shiyao Wang +7

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal…

cs.CL2026

Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue

Yuhang Jia, Pei Liu, Haoqin Sun +6

End-to-end Spoken Language Models (SLMs) hold great potential for paralinguistic perception, and numerous studies have aimed to enhance their capabilities, particularly for empathe…

eess.AS2024

Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

Xuechen Wang, Shiwan Zhao, Haoqin Sun +3

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective m…

cs.LG2025

Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning

Shiwan Zhao, Xuyang Zhao, Jiaming Zhou +3

Supervised fine-tuning (SFT) of large language models can be viewed as an off-policy learning problem, where expert demonstrations come from a fixed behavior policy while training…

cs.MM2024

Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework

Haoqin Sun, Shiwan Zhao, Shaokai Li +7

Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this…

cs.CV2026

EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations

Jiayi Liu, Jiaming Zhou, Ke Ye +3

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless…

cs.SD2026

GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR

Yujie Guo, Jiaming Zhou, Yuhang Jia +2

End-to-end multi-talker automatic speech recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech. A critical bottleneck is that speaker-speci…