17 citations · 21 across the 4 of their papers we have counts for
8 papers
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Ailin Huang, Bingxin Li, Bruce Wang +73
Large Audio-Language Models (LALMs) have significantly advanced intelligent human-computer interaction, yet their reliance on text-based outputs limits their ability to generate na…
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Ailin Huang, Boyong Wu, Bruce Wang +142
Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such…
ARCON: Advancing Auto-Regressive Continuation for Driving Videos
Ruibo Ming, Jingwei Wu, Zhewei Huang +4
Recent advancements in auto-regressive large language models (LLMs) have led to their application in video generation. This paper explores the use of Large Vision Models (LVMs) for…
A Survey on Future Frame Synthesis: Bridging Deterministic and Generative Approaches
Ruibo Ming, Zhewei Huang, Jingwei Wu +5
Future Frame Synthesis (FFS), the task of generating subsequent video frames from context, represents a core challenge in machine intelligence and a cornerstone for developing pred…
Scale-Adaptive Feature Aggregation for Efficient Space-Time Video Super-Resolution
Zhewei Huang, Ailin Huang, Xiaotao Hu +3
The Space-Time Video Super-Resolution (STVSR) task aims to enhance the visual quality of videos, by simultaneously performing video frame interpolation (VFI) and video super-resolu…
Learning to Paint With Model-based Deep Reinforcement Learning
Zhewei Huang, Wen Heng, Shuchang Zhou
We show how to teach machines to paint like human painters, who can use a small number of strokes to create fantastic paintings. By employing a neural renderer in model-based Deep…