collaborators

9 papers

cs.CL2026

X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance

Rime Wen, Zehan Liu, Shawn Qin +4

Streaming text-to-speech is essential for low-latency spoken dialogue systems, yet many systems wait for sentence-level text and are therefore only pseudo-streaming. True token-lev…

cs.CL2026

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Kaiqi Fu, Rime Wen, Altman Lin +4

Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, a…

cs.RO2026

HOST:Robots Acquire Manipulation Skills in Seconds from a Single Human Video

Guangyan Chen, Meiling Wang, Te Cui +9

The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-…

cs.RO2026

ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation

Yu Sun, Meng Cao, Yang Ping +24

Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constra…

cs.CV2026

X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining

Miracle Kang, Lights Shi, Lucy Liang +10

Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions prim…

cs.DC2026

DMuon: Efficient Distributed Muon Training with Near-Adam Overhead

Vincent Chen, Starrick Liu, Regis Cheng +8

Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads. The matrix-awar…