9 papers
Learning Native Continuation for Action Chunking Flow Policies
Yufeng Liu, Hang Yu, Juntu Zhao +9
Action chunking enables Vision Language Action (VLA) models to run in real time, but naive chunked execution often exhibits discontinuities at chunk boundaries. Real-Time Chunking…
Diffusion-Based Data Augmentation for Image Recognition: A Systematic Analysis and Evaluation
Zekun Li, Yinghuan Shi, Yang Gao +1
Diffusion-based data augmentation (DiffDA) has emerged as a promising approach to improving classification performance under data scarcity. However, existing works vary significant…
Spatially anchored Tactile Awareness for Robust Dexterous Manipulation
Jialei Huang, Yang Ye, Yuanqing Gong +3
Dexterous manipulation requires precise geometric reasoning, yet existing visuo-tactile learning methods struggle with sub-millimeter precision tasks that are routine for tradition…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
Ziqi Dai, Yiting Chen, Jiacheng Xu +8
The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, sc…
Intern-S1: A Scientific Multimodal Foundation Model
Lei Bai, Zhongrui Cai, Yuhang Cao +173
In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that…