7 papers
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
Renjie Lu, Xulong Zhang, Xiaoyang Qu +2
Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge that lie…
Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
Jianzong Wang, Botao Zhao, Yayun He +2
Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face limitations such as extensive tr…
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
Jiaqi Shi, Yuechan Li, Xulong Zhang +2
High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Existing acceleration strategi…
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
Qingran Yang, Botao Zhao, Zuheng Kang +7
The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Kno…
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
Renjie Lu, Xulong Zhang, Xiaoyang Qu +2
Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation…
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
Jiaqi Shi, Xulong Zhang, Xiaoyang Qu +1
Recent advances in Vision-Language-Action (VLA) models have shown promise for robot control, but their dependence on action supervision limits scalability and generalization. To ad…