activity
20172025
most citedLearning to Run with Actor-Critic Ensemble

17 citations · 21 across the 4 of their papers we have counts for

collaborators

8 papers

cs.SD2025

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Ailin Huang, Bingxin Li, Bruce Wang +73

Large Audio-Language Models (LALMs) have significantly advanced intelligent human-computer interaction, yet their reliance on text-based outputs limits their ability to generate na…

cs.CL20251 cited

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Ailin Huang, Boyong Wu, Bruce Wang +142

Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such…

cs.CV2024

ARCON: Advancing Auto-Regressive Continuation for Driving Videos

Ruibo Ming, Jingwei Wu, Zhewei Huang +4

Recent advancements in auto-regressive large language models (LLMs) have led to their application in video generation. This paper explores the use of Large Vision Models (LVMs) for…

cs.CV2024

A Survey on Future Frame Synthesis: Bridging Deterministic and Generative Approaches

Ruibo Ming, Zhewei Huang, Jingwei Wu +5

Future Frame Synthesis (FFS), the task of generating subsequent video frames from context, represents a core challenge in machine intelligence and a cornerstone for developing pred…

cs.CV20233 cited

Scale-Adaptive Feature Aggregation for Efficient Space-Time Video Super-Resolution

Zhewei Huang, Ailin Huang, Xiaotao Hu +3

The Space-Time Video Super-Resolution (STVSR) task aims to enhance the visual quality of videos, by simultaneously performing video frame interpolation (VFI) and video super-resolu…

cs.CV2019

Learning to Paint With Model-based Deep Reinforcement Learning

Zhewei Huang, Wen Heng, Shuchang Zhou

We show how to teach machines to paint like human painters, who can use a small number of strokes to create fantastic paintings. By employing a neural renderer in model-based Deep…