activity
20212026
most citedSinger Identification for Metaverse with Timbral and Middle-Level Perceptual Features

3 citations · 7 across the 25 of their papers we have counts for

collaborators

25 papers

cs.RO2026

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

Jianzong Wang, Botao Zhao, Yayun He +2

Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face limitations such as extensive tr…

cs.CV2026

From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration

Jiaqi Shi, Yuechan Li, Xulong Zhang +2

High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Existing acceleration strategi…

cs.SD2026

Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition

Qingran Yang, Botao Zhao, Zuheng Kang +7

The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Kno…

cs.CV2026

MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control

Renjie Lu, Xulong Zhang, Xiaoyang Qu +2

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation…

cs.RO2026

CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control

Jiaqi Shi, Xulong Zhang, Xiaoyang Qu +1

Recent advances in Vision-Language-Action (VLA) models have shown promise for robot control, but their dependence on action supervision limits scalability and generalization. To ad…

cs.SD2025

CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation

Ziqi Liang, Xulong Zhang, Chang Liu +3

Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However,…