activity
20242026
collaborators

6 papers

cs.SD2026

Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training

Jielin Qiu, Ming Zhu, Wenting Zhao +11

Audio-native large language models (audio-LLMs) commonly use Whisper as their audio encoder. However, Whisper was trained exclusively on speech data, producing weak representations…

cs.RO2026

Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation

Yiming Wu, Huan Wang, Zhenghao Chen +2

The growing complexity of visuomotor policies poses significant challenges for deployment with heterogeneous robotic hardware constraints. However, most existing model-efficient ap…

cs.RO2025

On-Device Diffusion Transformer Policy for Efficient Robot Manipulation

Yiming Wu, Huan Wang, Zhenghao Chen +2

Diffusion Policies have significantly advanced robotic manipulation tasks via imitation learning, but their application on resource-constrained mobile platforms remains challenging…

cs.CV2025

MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs

Hui Sun, Shiyin Lu, Huanyu Wang +5

Video large language models (Video-LLMs) have made significant progress in understanding videos. However, processing multiple frames leads to lengthy visual token sequences, presen…

cs.CV2024

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models

Yiming Wu, Zhenghao Chen, Huan Wang +1

The high computational cost and slow inference time are major obstacles to deploying Video Diffusion Models (VDMs). To overcome this, we introduce a new Video Diffusion Model Compr…

cs.CV2024

Frame-Voyager: Learning to Query Frames for Video Large Language Models

Sicheng Yu, Chengkai Jin, Huanyu Wang +9

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it…